Goto

Collaborating Authors

 Genre


AIDE: Fast and Communication Efficient Distributed Optimization

arXiv.org Machine Learning

In this paper, we present two new communication-efficient methods for distributed minimization of an average of functions. The first algorithm is an inexact variant of the DANE algorithm that allows any local algorithm to return an approximate solution to a local subproblem. We show that such a strategy does not affect the theoretical guarantees of DANE significantly. In fact, our approach can be viewed as a robustification strategy since the method is substantially better behaved than DANE on data partition arising in practice. It is well known that DANE algorithm does not match the communication complexity lower bounds. To bridge this gap, we propose an accelerated variant of the first method, called AIDE, that not only matches the communication lower bounds but can also be implemented using a purely first-order oracle. Our empirical results show that AIDE is superior to other communication efficient algorithms in settings that naturally arise in machine learning applications.


Kullback-Leibler Penalized Sparse Discriminant Analysis for Event-Related Potential Classification

arXiv.org Machine Learning

A brain computer interface (BCI) is a system that measures brain activity and converts it into an artificial output which is able to replace, restore or improve any normal output (neuromuscular or hormonal) used by a person to communicate and control his/her external or internal environment. Thus, BCI can significantly improve the quality of life of people with severe neuromuscular disabilities [35]. Communication between the brain of a person and the outside world can be appropriately established by means of a BCI system based on eventrelated potentials (ERPs), which are manifestations of neural activity as a consequence of certain infrequent or relevant stimuli. The main reason for using ERP-based BCI are: it is noninvasive, it requires minimal user training and it is quite robust (in the sense that it can be use by more than 90 % of people) [34]. One of the main components of such ERPs is the P300 wave, which is a positive deflection occurring in the scalp-recorded EEG approximately 300 ms after the stimulus has been applied. The P300 wave is unconsciously generated and its latency and amplitude vary between different EEG records of the same person, and even more, between EEG records of different persons [18].


Variance Reduction for Faster Non-Convex Optimization

arXiv.org Machine Learning

We consider the fundamental problem in non-convex optimization of efficiently reaching a stationary point. In contrast to the convex case, in the long history of this basic problem, the only known theoretical results on first-order non-convex optimization remain to be full gradient descent that converges in $O(1/\varepsilon)$ iterations for smooth objectives, and stochastic gradient descent that converges in $O(1/\varepsilon^2)$ iterations for objectives that are sum of smooth functions. We provide the first improvement in this line of research. Our result is based on the variance reduction trick recently introduced to convex optimization, as well as a brand new analysis of variance reduction that is suitable for non-convex optimization. For objectives that are sum of smooth functions, our first-order minibatch stochastic method converges with an $O(1/\varepsilon)$ rate, and is faster than full gradient descent by $\Omega(n^{1/3})$. We demonstrate the effectiveness of our methods on empirical risk minimizations with non-convex loss functions and training neural nets.


An Oracle Inequality for Quasi-Bayesian Non-Negative Matrix Factorization

arXiv.org Machine Learning

The aim of this paper is to provide some theoretical understanding of Bayesian non-negative matrix factorization methods. We derive an oracle inequality for a quasi-Bayesian estimator. This result holds for a very general class of prior distributions and shows how the prior affects the rate of convergence. We illustrate our theoretical results with a short numerical study along with a discussion on existing implementations.


Bridging AIC and BIC: a new criterion for autoregression

arXiv.org Machine Learning

We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information criterion. When the data is generated from a finite order autoregression, the Bayesian information criterion is known to be consistent, and so is the new criterion. When the true order is infinity or suitably high with respect to the sample size, the Akaike information criterion is known to be efficient in the sense that its prediction performance is asymptotically equivalent to the best offered by the candidate models; in this case, the new criterion behaves in a similar manner. Different from the two classical criteria, the proposed criterion adaptively achieves either consistency or efficiency depending on the underlying true model. In practice where the observed time series is given without any prior information about the model specification, the proposed order selection criterion is more flexible and robust compared with classical approaches. Numerical results are presented demonstrating the adaptivity of the proposed technique when applied to various datasets.


Multi-View Fuzzy Clustering with Minimax Optimization for Effective Clustering of Data from Multiple Sources

arXiv.org Machine Learning

Multi-view data clustering refers to categorizing a data set by making good use of related information from multiple representations of the data. It becomes important nowadays because more and more data can be collected in a variety of ways, in different settings and from different sources, so each data set can be represented by different sets of features to form different views of it. Many approaches have been proposed to improve clustering performance by exploring and integrating heterogeneous information underlying different views. In this paper, we propose a new multi-view fuzzy clustering approach called MinimaxFCM by using minimax optimization based on well-known Fuzzy c means. In MinimaxFCM the consensus clustering results are generated based on minimax optimization in which the maximum disagreements of different weighted views are minimized. Moreover, the weight of each view can be learned automatically in the clustering process. In addition, there is only one parameter to be set besides the fuzzifier. The detailed problem formulation, updating rules derivation, and the in-depth analysis of the proposed MinimaxFCM are provided here. Experimental studies on nine multi-view data sets including real world image and document data sets have been conducted. We observed that MinimaxFCM outperforms related multi-view clustering approaches in terms of clustering accuracy, demonstrating the great potential of MinimaxFCM for multi-view data analysis.


Smart machines, platforms and human-centric tech tipped as top trends

#artificialintelligence

Smart dust, 4D printing and brain-computer interfaces are some of the most attention-grabbing technologies featuring in Gartner's latest Hype Cycle for Emerging Technologies. The forecast reveals three major technology trends set to be the highest priority for organizations looking to compete in the digital world. Emerging technologies are revolutionizing the concepts of how platforms are defined and used, Gartner says. The report notes, "The shift from technical infrastructure to ecosystem-enabling platforms is laying the foundations for entirely new business models that are forming the bridge between humans and technology. Within these dynamic ecosystems, organizations must proactively understand and redefine their strategy to create platform-based business models, and to exploit internal and external algorithms in order to generate value."


Magnetic Appoints Data and Machine Learning Veteran Paul Phillips as Chief Data Officer

#artificialintelligence

A data entrepreneur, Phillips comes to Magnetic from leading data analytics provider, Causata, which Phillips founded and led. The company was acquired by NICE Systems Ltd. Phillips also founded Touch Clarity, which specialized in personalization and machine learning. The company was acquired by Omniture, eventually becoming a part of Adobe. "Magnetic is unique in having access to both data from the world of advertising and the world of CRM. The future of marketing will not only demand that we understand what people are in-market for right now, but what every touch may mean to the expected lifetime value of a customer. Magnetic's data assets and technology platform position us well to deliver on that promise," says Phillips.


Girl Geeks Toronto

#artificialintelligence

"AI ...surely will be a trend at least on the size of big data. It almost certainly will be a trend on the size of mobile. It might be a trend on the size of the internet. And maybe, just maybe, it'll be a trend on the size of software; that the software before machine intelligence and after will be two worlds that are very different from each other." It's undeniable that artificial intelligence (AI) is one of tech's hottest topics and a trend that is permeating every part of our world.


Meet Alrobot, the remote-controlled robotic tank being used by the Iraqi army to fight ISIS

Daily Mail - Science & tech

The Iraqi army is testing its latest weapon in the fight against ISIS, a remote-controlled battle robot. The unmanned vehicle has a heavy machine gun turret for taking out targets picked out by the controller. Called Alrobot, the bullet-proof vehicle is being tested in the Iraqi desert as part of the army's efforts to retake the city of Mosul from ISIS. Iraq is preparing a remote controlled attack vehicle being which experts say could be used in the fight to retake the city of Mosul from Isis. Reports indicate the remote controlled vehicle, called'Alrobot', was designed by two brothers in Baghdad.