Goto

Collaborating Authors

 Country


A note on adjusting $R^2$ for using with cross-validation

arXiv.org Machine Learning

We show how to adjust the coefficient of determination ($R^2$) when used for measuring predictive accuracy via leave-one-out cross-validation.


Clustering on the Edge: Learning Structure in Graphs

arXiv.org Machine Learning

With the recent popularity of graphical clustering methods, there has been an increased focus on the information between samples. We show how learning cluster structure using edge features naturally and simultaneously determines the most likely number of clusters and addresses data scale issues. These results are particularly useful in instances where (a) there are a large number of clusters and (b) we have some labeled edges. Applications in this domain include image segmentation, community discovery and entity resolution. Our model is an extension of the planted partition model and our solution uses results of correlation clustering, which achieves a partition O(log(n))-close to the log-likelihood of the true clustering.


The IBM Speaker Recognition System: Recent Advances and Error Analysis

arXiv.org Machine Learning

We present the recent advances along with an error analysis of the IBM speaker recognition system for conversational speech. Some of the key advancements that contribute to our system include: a nearest-neighbor discriminant analysis (NDA) approach (as opposed to LDA) for intersession variability compensation in the i-vector space, the application of speaker and channel-adapted features derived from an automatic speech recognition (ASR) system for speaker recognition, and the use of a DNN acoustic model with a very large number of output units ( 10k senones) to compute the frame-level soft alignments required in the i-vector estimation process. We evaluate these techniques on the NIST 2010 SRE extended core conditions (C1-C9), as well as the 10sec-10sec condition. To our knowledge, results achieved by our system represent the best performances published to date on these conditions. For example, on the extended tel-tel condition (C5) the system achieves an EER of 0.59%. To garner further understanding of the remaining errors (on C5), we examine the recordings associated with the low scoring target trials, where various issues are identified for the problematic recordings/trials. Interestingly, it is observed that correcting the pathological recordings not only improves the scores for the target trials but also for the nontarget trials.


Provable Bayesian Inference via Particle Mirror Descent

arXiv.org Machine Learning

Bayesian methods are appealing in their flexibility in modeling complex data and ability in capturing uncertainty in parameters. However, when Bayes' rule does not result in tractable closed-form, most approximate inference algorithms lack either scalability or rigorous guarantees. To tackle this challenge, we propose a simple yet provable algorithm, \emph{Particle Mirror Descent} (PMD), to iteratively approximate the posterior density. PMD is inspired by stochastic functional mirror descent where one descends in the density space using a small batch of data points at each iteration, and by particle filtering where one uses samples to approximate a function. We prove result of the first kind that, with $m$ particles, PMD provides a posterior density estimator that converges in terms of $KL$-divergence to the true posterior in rate $O(1/\sqrt{m})$. We demonstrate competitive empirical performances of PMD compared to several approximate inference algorithms in mixture models, logistic regression, sparse Gaussian processes and latent Dirichlet allocation on large scale datasets.


Information Recovery from Pairwise Measurements

arXiv.org Machine Learning

This paper is concerned with jointly recovering $n$ node-variables $\left\{ x_{i}\right\}_{1\leq i\leq n}$ from a collection of pairwise difference measurements. Imagine we acquire a few observations taking the form of $x_{i}-x_{j}$; the observation pattern is represented by a measurement graph $\mathcal{G}$ with an edge set $\mathcal{E}$ such that $x_{i}-x_{j}$ is observed if and only if $(i,j)\in\mathcal{E}$. To account for noisy measurements in a general manner, we model the data acquisition process by a set of channels with given input/output transition measures. Employing information-theoretic tools applied to channel decoding problems, we develop a \emph{unified} framework to characterize the fundamental recovery criterion, which accommodates general graph structures, alphabet sizes, and channel transition measures. In particular, our results isolate a family of \emph{minimum} \emph{channel divergence measures} to characterize the degree of measurement corruption, which together with the size of the minimum cut of $\mathcal{G}$ dictates the feasibility of exact information recovery. For various homogeneous graphs, the recovery condition depends almost only on the edge sparsity of the measurement graph irrespective of other graphical metrics; alternatively, the minimum sample complexity required for these graphs scales like \[ \text{minimum sample complexity }\asymp\frac{n\log n}{\mathsf{Hel}_{1/2}^{\min}} \] for certain information metric $\mathsf{Hel}_{1/2}^{\min}$ defined in the main text, as long as the alphabet size is not super-polynomial in $n$. We apply our general theory to three concrete applications, including the stochastic block model, the outlier model, and the haplotype assembly problem. Our theory leads to order-wise tight recovery conditions for all these scenarios.


Multilingual Twitter Sentiment Classification: The Role of Human Annotators

arXiv.org Artificial Intelligence

What are the limits of automated Twitter sentiment classification? We analyze a large set of manually labeled tweets in different languages, use them as training data, and construct automated classification models. It turns out that the quality of classification models depends much more on the quality and size of training data than on the type of the model trained. Experimental results indicate that there is no statistically significant difference between the performance of the top classification models. We quantify the quality of training data by applying various annotator agreement measures, and identify the weakest points of different datasets. We show that the model performance approaches the inter-annotator agreement when the size of the training set is sufficiently large. However, it is crucial to regularly monitor the self- and inter-annotator agreements since this improves the training datasets and consequently the model performance. Finally, we show that there is strong evidence that humans perceive the sentiment classes (negative, neutral, and positive) as ordered.


Machine Learning Turns Unstructured Data into Valuable Health Information - DATAVERSITY

#artificialintelligence

Kevin Murnane reports in Forbes, "Machine learning's ability to produce actionable results from unstructured data is clearly demonstrated in a study published in the April 2016 volume of The Journal of Biomedical Informatics that used machine learning algorithms to identify cancer diagnoses from free text pathology reports. The results provide strong evidence that machine learning techniques can bypass a central bottleneck that interferes with turning unstructured data into useful information." He continues, "Medical practitioners routinely create clinical reports on their patients that contain a wealth of potentially useful and valuable information. Many benefits could be realized if the information in these reports could be compiled and easily accessed to produce actionable results. For example, the dangerous levels of lead in Flint, Michigan's water supply might have been discovered sooner if individual doctors' reports noting unusually high lead levels in children's blood had been gathered in an accessible format at an earlier date."


It's official: Google self-driving tech will debut in a future Chrysler Pacifica Hybrid

#artificialintelligence

Stop the presses on this one: Google is going to start working with Chrysler on their new Pacifica mini-van. The plan is to test a fleet of the 2017 model for now and have the technology integrated into the vehicle, not just as an aftermarket proof of concept. There will be a future version of the Chrysler Pacifica that has the self-driving tech. This is the first time, according to Chrysler, that Google has integrated sensors and software into a passenger car for consumer use in a future model, not just as a test in a current passenger car. One of the most surprising moves is that both engineering teams will co-locate in Michigan to work together and design the technology for eventual consumer use.


Rich and powerful warn robots are coming for your jobs

#artificialintelligence

"Most of the benefits we see from automation is about higher quality and fewer errors, but in many cases it does reduce labor," Michael Chui, a partner at the McKinsey Global Institute, said on Tuesday during a panel on "Is Any Job Truly Safe?" The four-day annual conference, which began on Sunday, has 3,500 invite-only participants exploring "The Future of Human Kind." Technology has not only done away with low-wage, low-skill jobs, some of the more than 700 speakers said. They cited robots operating trucks in some Australian mines; corporate litigation software replacing employees with advanced degrees who used to sift through thousands of documents prior to trials; and on Wall Street, the automation of jobs previously done by bankers with MBAs or PhDs. "Anyone whose job is moving data from one spreadsheet to another ..., that's what is going to get automated," said Daniel Nadler, chief executive of Kensho, a financial services analytics company partly owned by Goldman Sachs Group Inc. "Goldman Sachs will be in here in 10 years, JPMorgan will be here. They're just going to be much more efficient in terms of operating leverage and headcount," he added.


Rich and powerful warn robots are coming for your jobs

#artificialintelligence

Some of the most powerful people in the world have gathered this week to discuss the most pressing issues affecting humanity. And the overwhelming conclusion is that the robots are coming. At the Milken Institute's Global Conference in California, at least four panels focused ontechnology taking over markets to mining, and most importantly,jobs. Some of the most powerful people in the world have gathered this week to discuss the most pressing issues affecting humanity, and the overwhelming conclusion is the robots are coming. At the Milken Institute's Global Conference in California, four panels focused on technology taking over markets and jobs (stock image) 'Most of the benefits we see from automation is about higherquality and fewer errors, but in many cases it does reducelabor,' Michael Chui, a partner at the McKinsey GlobalInstitute, said on Tuesday during a panel on'Is Any Job TrulySafe?'