Goto

Collaborating Authors

 Genre


Unified Convergence Analysis of Stochastic Momentum Methods for Convex and Non-convex Optimization

arXiv.org Machine Learning

Recently, {\it stochastic momentum} methods have been widely adopted in training deep neural networks. However, their convergence analysis is still underexplored at the moment, in particular for non-convex optimization. This paper fills the gap between practice and theory by developing a basic convergence analysis of two stochastic momentum methods, namely stochastic heavy-ball method and the stochastic variant of Nesterov's accelerated gradient method. We hope that the basic convergence results developed in this paper can serve the reference to the convergence of stochastic momentum methods and also serve the baselines for comparison in future development of stochastic momentum methods. The novelty of convergence analysis presented in this paper is a unified framework, revealing more insights about the similarities and differences between different stochastic momentum methods and stochastic gradient method. The unified framework exhibits a continuous change from the gradient method to Nesterov's accelerated gradient method and finally the heavy-ball method incurred by a free parameter, which can help explain a similar change observed in the testing error convergence behavior for deep learning. Furthermore, our empirical results for optimizing deep neural networks demonstrate that the stochastic variant of Nesterov's accelerated gradient method achieves a good tradeoff (between speed of convergence in training error and robustness of convergence in testing error) among the three stochastic methods.


Classical Statistics and Statistical Learning in Imaging Neuroscience

arXiv.org Machine Learning

Neuroimaging research has predominantly drawn conclusions based on classical statistics, including null-hypothesis testing, t-tests, and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity, including cross-validation, pattern classification, and sparsity-inducing regression. These two methodological families used for neuroimaging data analysis can be viewed as two extremes of a continuum. Yet, they originated from different historical contexts, build on different theories, rest on different assumptions, evaluate different outcome metrics, and permit different conclusions. This paper portrays commonalities and differences between classical statistics and statistical learning with their relation to neuroimaging research. The conceptual implications are illustrated in three common analysis scenarios. It is thus tried to resolve possible confusion between classical hypothesis testing and data-guided model estimation by discussing their ramifications for the neuroimaging access to neurobiology.


Ethnicity sensitive author disambiguation using semi-supervised learning

arXiv.org Machine Learning

Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital libraries are increasingly offering tools for authors to manually curate their publications and claim those that are theirs. Indirectly, these tools allow for the inexpensive collection of large annotated training data, which can be further leveraged to build a complementary automated disambiguation system capable of inferring patterns for identifying publications written by the same person. Building on more than 1 million publicly released crowdsourced annotations, we propose an automated author disambiguation solution exploiting this data (i) to learn an accurate classifier for identifying coreferring authors and (ii) to guide the clustering of scientific publications by distinct authors in a semi-supervised way. To the best of our knowledge, our analysis is the first to be carried out on data of this size and coverage. With respect to the state of the art, we validate the general pipeline used in most existing solutions, and improve by: (i) proposing phonetic-based blocking strategies, thereby increasing recall; and (ii) adding strong ethnicity-sensitive features for learning a linkage function, thereby tailoring disambiguation to non-Western author names whenever necessary.


The Hidden Convexity of Spectral Clustering

arXiv.org Machine Learning

Partitioning a dataset into classes based on a similarity between data points, known as cluster analysis, is one of the most basic and practically important problems in data analysis and machine learning. It has a vast array of applications from speech recognition to image analysis to bioinformatics and to data compression. There is an extensive literature on the subject, including a number of different methodologies as well as their various practical and theoretical aspects [11]. In recent years spectral clustering--a class of methods based on the eigenvectors of a certain matrix, typically the graph Laplacian constructed from data--has become a widely used method for cluster analysis. This is due to the simplicity of the algorithm, a number of desirable properties it exhibits and its amenability to theoretical analysis. In its simplest form, spectral bi-partitioning is an attractively straightforward algorithm based on thresholding the second bottom eigenvector of the Laplacian matrix of a graph. However, the more practically significant problem of multiway spectral clustering is considerably more complex. While hierarchical methods based on a sequence of binary splits have been used, the most common approaches use k-means or weighted k-means clustering in the spectral space or related iterative procedures [17, 15, 2, 25].


Operational Machine Learning -- Madrid Workshop

#artificialintelligence

It provides an agnostic introduction to operational ML with open source and cloud platforms. It is the first ML workshop to go all the way from data preparation to the integration of predictive models in real-world applications and their deployment in production. Participants will learn to use Python open source libraries scikit-learn, Pandas and SKLL, and cloud platforms Microsoft Azure ML, Amazon ML, BigML and Indico (along with their APIs).


Cognex (CGNX) Robert J. Willett on Q1 2016 Results - Earnings Call Transcript

#artificialintelligence

Currently at this time, all participants are in a listen-only mode. Later, we will conduct the question-and-answer session and instructions will follow at that time. Also, as a reminder, this conference call is being recorded. I would now like to turn the call over to your host to Richard Morin. Thank you, and good evening, everyone. Earlier today, we issued a news release announcing Cognex's earnings for the first quarter of 2016, and we've also filed our quarterly report on Form 10-Q. For those of you who have not yet seen these materials, both are available on our website at www.cognex.com. They contain highly detailed information about our financial results. During tonight's call, we may use a non-GAAP financial measure, if we believe it is useful to investors, or if we believe it will help investors better understand our results or business trends. For your reference, you can see a reconciliation of certain items from GAAP to non-GAAP in Exhibit 2 of the earnings release. I'd like to emphasize that any forward-looking statements we made in the earnings release or any that we may make during this call are based upon information that we believe to be true as of today. Things often change and actual results may differ materially from those projected or anticipated. You should refer to the company's SEC filings, including our most recent Form 10-K, for a detailed list of these risk factors. Now, I'll turn the call over to Cognex's Chairman, Dr. Bob Shillman.


Temporal Clustering of Time Series via Threshold Autoregressive Models: Application to Commodity Prices

arXiv.org Machine Learning

This study aimed to find temporal clusters for several commodity prices using the threshold nonlinear autoregressive model. It is expected that the process of determining the commodity groups that are time-dependent will advance the current knowledge about the dynamics of co-moving and coherent prices, and can serve as a basis for multivariate time series analyses. The clustering of commodity prices was examined using the proposed clustering approach based on time series models to incorporate the time varying properties of price series into the clustering scheme. Accordingly, the primary aim in this study was grouping time series according to the similarity between their Data Generating Mechanisms (DGMs) rather than comparing pattern similarities in the time series traces. The approximation to the DGM of each series was accomplished using threshold autoregressive models, which are recognized for their ability to represent nonlinear features in time series, such as abrupt changes, time-irreversibility and regime-shifting behavior. Through the use of the proposed approach, one can determine and monitor the set of co-moving time series variables across the time dimension. Furthermore, generating a time varying commodity price index and sub-indexes can become possible. Consequently, we conducted a simulation study to assess the effectiveness of the proposed clustering approach and the results are presented for both the simulated and real data sets. Keywords: Clustering Nonlinear Time Series Models, Regime Switching, Spectral 1. Introduction The movement of commodity prices and the associated dynamics are interrelated with economics and directly affect many industries.


Decentralized Dynamic Discriminative Dictionary Learning

arXiv.org Machine Learning

We develop a framework to solve machine learning problems in cases where latent geometric structure in the feature space may be exploited. We consider cases where the number of training examples is either very large, or signals are sequentially observed by a platform operating in real-time such as an autonomous robot. In the former case, since the sample size is large-scale, processing a few training examples at a time is necessary due to computational cost. However, doing so at a centralized location may be impractical, which motivates the use of learning techniques that may be done collaboratively by a network of interconnected computing servers. In the later case, an autonomous robot with no priors on its operating environment only has access to information based on the path it has traversed, which may omit regions of the feature space crucial for tasks such as learning-based control. By communicating with other robots in a network, individuals may learn over a broader domain associated with that which has been explored by the whole network, and thus more effectively solve autonomous learning tasks.


Efficient Distributed Estimation of Inverse Covariance Matrices

arXiv.org Machine Learning

ABSTRACT In distributed systems, communication is a major concern due to issues such as its vulnerability or efficiency. In this paper, we are interested in estimating sparse inverse covariance matrices when samples are distributed into different machines. We address communication efficiency by proposing a method where, in a single round of communication, each machine transfers a small subset of the entries of the inverse covariance matrix. We show that, with this efficient distributed method, the error rates can be comparable with estimation in a non-distributed setting, and correct model selection is still possible. Practical performance is shown through simulations.


An evaluation of randomized machine learning methods for redundant data: Predicting short and medium-term suicide risk from administrative records and risk assessments

arXiv.org Machine Learning

Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording processes. We investigate the efficacy of three most effective randomized machine learning techniques - random forests, gradient boosting machines, and deep neural nets with dropout - in predicting suicide risk. Using a cohort of mental health patients from a regional Australian hospital, we compare the predictive performance with popular traditional approaches - clinician judgments based on a checklist, sparse logistic regression and decision trees. The randomized methods demonstrated robustness against data redundancies and superior predictive performance on AUC and F-measure. Keywords: Suicide risk, Electronic medical record, Predictive models, Randomized machine learning, Deep learning 1. Introduction Every year, about 2000 Australians die by suicide causing huge trauma to families, friends, workplaces and communities[1].