Goto

Collaborating Authors

 Genre


Stochastic Heavy Ball

arXiv.org Machine Learning

This paper deals with a natural stochastic optimization procedure derived from the so-called Heavy-ball method differential equation, which was introduced by Polyak in the 1960s with his seminal contribution [Pol64]. The Heavy-ball method is a second-order dynamics that was investigated to minimize convex functions f . The family of second-order methods recently received a large amount of attention, until the famous contribution of Nesterov [Nes83], leading to the explosion of large-scale optimization problems. This work provides an in-depth description of the stochastic heavy-ball method, which is an adaptation of the deterministic one when only unbiased evalutions of the gradient are available and used throughout the iterations of the algorithm. We first describe some almost sure convergence results in the case of general non-convex coercive functions f . We then examine the situation of convex and strongly convex potentials and derive some non-asymptotic results about the stochastic heavy-ball method. We end our study with limit theorems on several rescaled algorithms.


Semi-Supervised Learning with Generative Adversarial Networks

arXiv.org Machine Learning

We extend Generative Adversarial Networks (GANs) to the semi-supervised context by forcing the discriminator network to output class labels. We train a generative model G and a discriminator D on a dataset with inputs belonging to one of N classes. At training time, D is made to predict which of N 1 classes the input belongs to, where an extra class is added to correspond to the outputs of G. We show that this method can be used to create a more data-efficient classifier and that it allows for generating higher quality samples than a regular GAN.


A Theory of Local Learning, the Learning Channel, and the Optimality of Backpropagation

arXiv.org Machine Learning

In a physical neural system, where storage and processing are intimately intertwined, the rules for adjusting the synaptic weights can only depend on variables that are available locally, such as the activity of the pre- and post-synaptic neurons, resulting in local learning rules. A systematic framework for studying the space of local learning rules is obtained by first specifying the nature of the local variables, and then the functional form that ties them together into each learning rule. Such a framework enables also the systematic discovery of new learning rules and exploration of relationships between learning rules and group symmetries. We study polynomial local learning rules stratified by their degree and analyze their behavior and capabilities in both linear and non-linear units and networks. Stacking local learning rules in deep feedforward networks leads to deep local learning. While deep local learning can learn interesting representations, it cannot learn complex input-output functions, even when targets are available for the top layer. Learning complex input-output functions requires local deep learning where target information is communicated to the deep layers through a backward learning channel. The nature of the communicated information about the targets and the structure of the learning channel partition the space of learning algorithms. We estimate the learning channel capacity associated with several algorithms and show that backpropagation outperforms them by simultaneously maximizing the information rate and minimizing the computational cost, even in recurrent networks. The theory clarifies the concept of Hebbian learning, establishes the power and limitations of local learning rules, introduces the learning channel which enables a formal analysis of the optimality of backpropagation, and explains the sparsity of the space of learning rules discovered so far.


Hybrid clustering-classification neural network in the medical diagnostics of reactive arthritis

arXiv.org Machine Learning

Self-organizing maps (SOM) and neural networks of learning vector quantization (LVQ) have seen extensive use for solving different problems in Data Mining domain (clustering, classification, fault detection and compression of information etc.). This type of neural networks was proposed by T. Kohonen [1, 2] and represents, in fact, a single-layer feedforward architecture, which provides an operator for mapping of input space into the output space. Operation-wise SOM and LVQ are quite similar to each neuron is fed input signal (sample) producing output, which is used during competition stage to determine winning neuron - usually the one with maximum output signal value. Vector of synaptic weights for winning neuron is the one closest to the input sample in terms of the metric chosen (which is Euclidian metric in most cases). Next is neurons adjustment phase.


Convex Formulation for Kernel PCA and its Use in Semi-Supervised Learning

arXiv.org Machine Learning

In this paper, Kernel PCA is reinterpreted as the solution to a convex optimization problem. Actually, there is a constrained convex problem for each principal component, so that the constraints guarantee that the principal component is indeed a solution, and not a mere saddle point. Although these insights do not imply any algorithmic improvement, they can be used to further understand the method, formulate possible extensions and properly address them. As an example, a new convex optimization problem for semi-supervised classification is proposed, which seems particularly well-suited whenever the number of known labels is small. Our formulation resembles a Least Squares SVM problem with a regularization parameter multiplied by a negative sign, combined with a variational principle for Kernel PCA. Our primal optimization principle for semi-supervised learning is solved in terms of the Lagrange multipliers. Numerical experiments in several classification tasks illustrate the performance of the proposed model in problems with only a few labeled data.


UTA-poly and UTA-splines: additive value functions with polynomial marginals

arXiv.org Artificial Intelligence

Additive utility function models are widely used in multiple criteria decision analysis. In such models, a numerical value is associated to each alternative involved in the decision problem. It is computed by aggregating the scores of the alternative on the different criteria of the decision problem. The score of an alternative is determined by a marginal value function that evolves monotonically as a function of the performance of the alternative on this criterion. Determining the shape of the marginals is not easy for a decision maker. It is easier for him/her to make statements such as "alternativea is preferred tob". In order to help the decision maker, UTA disaggregation procedures use linear programming to approximate the marginals by piecewise linear functions based only on such statements. In this paper, we propose to infer polynomials and splines instead of piecewise linear functions for the marginals. In this aim, we use semidefinite programming instead of linear programming. We illustrate this new elicitation method and present some experimental results. Introduction The theory of value functions aims at assigning a number to each alternative in such a way that the decision maker's preference order on the alternatives is the same as the order on the numbers associated with the alternatives. The number or value associated to an alternative is a monotone function of its evaluations on the various relevant criteria. For preferences satisfying some additional properties (includingpreferential independence), the value of an alternative can be obtained as the sum of marginal value functions each depending only on a single criterion [20, Chapter 6]. These functions usually are monotone, i.e., marginal value functions either increase or decrease with the assessment of the alternative on the associated criterion. Many questioning protocols have been proposed aiming to elicit an additive value function [20, 9] through interactions with the decision maker (DM). These direct elicitation methods are time-consuming and require a substantial cognitive effort from the DM. Therefore, in certain cases, an indirect approach may prove fruitful. The latter consists inlearning an additive value model (or a set of such models) from a set of declared or observed preferences. Learning approaches have been proposed not only for inferring an additive value function that is used to rank all other alternatives.


Why Intel Is Tweaking Xeon Phi For Deep Learning

#artificialintelligence

If there is anything that chip giant Intel has learned over the past two decades as it has gradually climbed to dominance in processing in the datacenter, it is ironically that one size most definitely does not fit all. As the tight co-design of hardware and software continues in all parts of the IT industry, we can expect fine-grained customization for very precise – and lucrative – workloads, like data analytics and machine learning, just to name two of the hottest areas today. Software will run most efficiently on hardware that is tuned for it, although we are used to thinking of that process in a mirror image, where programmers tweak their code to take advantage of the forward-looking features a chip maker conceives of four or five years before they are etched into its transistors and delivered as a product. The competition is fierce these days, and Intel has to move fast if it is to keep its compute hegemony in the datacenter. That is why at the Intel Developer Forum in San Francisco the company put a new path on the Knights family of many-core processors that will see the company deliver a version of this chip specifically tuned for machine learning workloads.


How Artificial Intelligence Is Helping Enhance Human Capabilities

#artificialintelligence

In the past half decade, artificial intelligence and machine learning have made significant leaps into the mainstream and into our daily lives. According to research firm Markets and Markets, the artificial intelligence market is set to grow to 5.05 billion by 2020 thanks to the increased applicability of various AI technologies into everything from finance to healthcare to retail. Today, doctors can diagnose Sepsis with an AI algorithm, for instance, and researchers can track endangered species through AI-enhanced photo capture systems. Clearly, these new self-learning and ever-improving technologies have limitless potential in a number of innovative industries. The U.S. Chamber of Commerce's Technology Engagement Center (C_TEC) recently hosted a panel discussion during its TecNation 2016 event that focused on where we stand with Artificial Intelligence and how it will affect our lives and unlock our potential in the long run.


Top 10 Data Science and Machine Learning Podcasts - Dataconomy

#artificialintelligence

In order to protect the world's iconic marine wildlife including whales, sea turtles and sharks, we first have to understand their biology. However, this is often easier said than done. In this 45-minute interactive lesson, students will be taken around the world to learn about new and exciting ways that marine scientists are uncovering the lives of these elusive creatures. Track humpback whales as they feed in Alaska, and come along for the ride as video cameras are deployed on sea turtles in Western Australia. The lesson uses photos and videos from a variety of active research projects, begins with historical context about human impacts on marine wildlife populations and ends with a discussion of what students can do in their lives to help learn about and protect our oceans.


AppTek Expands its Talk2Me Brand to Include Speech-to-Speech Translati

#artificialintelligence

AppTek, a leader of automatic speech recognition (ASR), machine learning and artificial intelligence, today announces the newest mobile application addition to its Talk2Me product suite: Talk2Me Mobile. The speech-to-speech translation app, now available for free on iOS and Android markets, is used for instantaneous bi-directional Spanish/English and Arabic/English translations. In addition to live speech-to-speech translation, the app supports the ability to record, transcribe and archive content for future offline use. Whether for customer service calls, interviews, conference calls, web conferences or peer-to-peer conversations, Talk2Me converts audio and video assets into searchable, actionable data to help organizations draw comprehensive insights and transform the way people and businesses operate. "Our advances in speech technology, machine learning and artificial intelligence are lowering the barrier for people and businesses to effectively navigate, retrieve and transact across multiple languages," stated AppTek's CEO, Adam Sutherland.