Genre
Stagewise Learning for Sparse Clustering of Discretely-Valued Data
Zhao, Vincent, Zucker, Steven W.
We study the model-based sparse clustering problem for discrete data using a mixture model of product distributions [9, 7]. This model has application in many fields, including computational neurosciences, crowdsourcing and bioinformatics, and is interesting because it differs technically from the problem for continuous data, where the well-known Gaussian mixture model has been applied successfully. A fundamental difficulty is that, in high-dimensional datasets, some features can be noisy, redundant or generally uninformative for clustering, and these can push clustering algorithms toward inappropriate or uninteresting results. If these uninformative or noise data points could be eliminated then, we argue, the results should be much more satisfying. This is precisely our goal: to find an informative set of data points and to use these to drive the clustering.
Improved SVRG for Non-Strongly-Convex or Sum-of-Non-Convex Objectives
Many classical algorithms are found until several years later to outlive the confines in which they were conceived, and continue to be relevant in unforeseen settings. In this paper, we show that SVRG is one such method: being originally designed for strongly convex objectives, it is also very robust in non-strongly convex or sum-of-non-convex settings. More precisely, we provide new analysis to improve the state-of-the-art running times in both settings by either applying SVRG or its novel variant. Since non-strongly convex objectives include important examples such as Lasso or logistic regression, and sum-of-non-convex objectives include famous examples such as stochastic PCA and is even believed to be related to training deep neural nets, our results also imply better performances in these applications.
Particle Metropolis-adjusted Langevin algorithms
Nemeth, Christopher, Sherlock, Chris, Fearnhead, Paul
Markov chain Monte Carlo algorithms are a popular and well-studied methodology that can be used to draw samples from posterior distributions. Over the past few years these algorithms have been extended to tackle problems where the model likelihood is intractable (Beaumont, 2003). Andrieu and Roberts (2009) showed that within the Metropolis-Hastings algorithm, if the likelihood is replaced with an unbiased estimate, then the sampler still targets the correct stationary distribution. Andrieu et al. (2010) extended this work further to create a class of 1 Markov chain algorithms that use sequential Monte Carlo methods, also known as particle filters. Current implementations of pseudo-marginal and particle Markov chain Monte Carlo use random-walk proposals to update the parameters (e.g., Golightly and Wilkinson, 2011; Knape and de Valpine, 2012) and shall be referred to herein as particle random-walk Metropolis algorithms. Random walk-based algorithms propose a new value from some symmetric density centred on the current value.
Learning may need only a few bits of synaptic precision
Baldassi, Carlo, Gerace, Federica, Lucibello, Carlo, Saglietti, Luca, Zecchina, Riccardo
Learning in neural networks poses peculiar challenges when using discretized rather then continuous synaptic states. The choice of discrete synapses is motivated by biological reasoning and experiments, and possibly by hardware implementation considerations as well. In this paper we extend a previous large deviations analysis which unveiled the existence of peculiar dense regions in the space of synaptic states which accounts for the possibility of learning efficiently in networks with binary synapses. We extend the analysis to synapses with multiple states and generally more plausible biological features. The results clearly indicate that the overall qualitative picture is unchanged with respect to the binary case, and very robust to variation of the details of the model. We also provide quantitative results which suggest that the advantages of increasing the synaptic precision (i.e.~the number of internal synaptic states) rapidly vanish after the first few bits, and therefore that, for practical applications, only few bits may be needed for near-optimal performance, consistently with recent biological findings. Finally, we demonstrate how the theoretical analysis can be exploited to design efficient algorithmic search strategies.
EEF: Exponentially Embedded Families with Class-Specific Features for Classification
Tang, Bo, Kay, Steven, He, Haibo, Baggenstoss, Paul M.
Classification is one of fundamental problems in the fields of machine learning and signal processing. The commonly used classifier assigns a sample or a signal to the class with maximum posterior probability, which usually requires probability density function (PDF) estimation in an either model-driven or data-driven manner [1] [2] [3]. For high-dimensional data sets, it is necessary to perform feature reduction to estimate the PDFs robustly in a lowdimensional feature subspace. However, feature reduction may lose pertinent information for discrimination. For example, data samples from different classes that could be well separated in the raw data space may be overlapped in the feature subspace, causing classification errors. The PDF reconstruction approach provides a solution to address this information loss issue in feature reduction by reconstructing the PDF on raw data and making classification in raw data space, which could improve classification performance. Several approaches have been developed along this track.
On the Expressive Power of Deep Learning: A Tensor Analysis
Cohen, Nadav, Sharir, Or, Shashua, Amnon
It has long been conjectured that hypotheses spaces suitable for data that is compositional in nature, such as text or images, may be more efficiently represented with deep hierarchical networks than with shallow ones. Despite the vast empirical evidence supporting this belief, theoretical justifications to date are limited. In particular, they do not account for the locality, sharing and pooling constructs of convolutional networks, the most successful deep learning architecture to date. In this work we derive a deep network architecture based on arithmetic circuits that inherently employs locality, sharing and pooling. An equivalence between the networks and hierarchical tensor factorizations is established. We show that a shallow network corresponds to CP (rank-1) decomposition, whereas a deep network corresponds to Hierarchical Tucker decomposition. Using tools from measure theory and matrix algebra, we prove that besides a negligible set, all functions that can be implemented by a deep network of polynomial size, require exponential size in order to be realized (or even approximated) by a shallow network. Since log-space computation transforms our networks into SimNets, the result applies directly to a deep learning architecture demonstrating promising empirical performance. The construction and theory developed in this paper shed new light on various practices and ideas employed by the deep learning community.
On the Sensitivity of the Lasso to the Number of Predictor Variables
Flynn, Cheryl J., Hurvich, Clifford M., Simonoff, Jeffrey S.
The Lasso is a computationally efficient regression regularization procedure that can produce sparse estimators when the number of predictors (p) is large. Oracle inequalities provide probability loss bounds for the Lasso estimator at a deterministic choice of the regularization parameter. These bounds tend to zero if p is appropriately controlled, and are thus commonly cited as theoretical justification for the Lasso and its ability to handle high-dimensional settings. Unfortunately, in practice the regularization parameter is not selected to be a deterministic quantity, but is instead chosen using a random, data-dependent procedure. To address this shortcoming of previous theoretical work, we study the loss of the Lasso estimator when tuned optimally for prediction. Assuming orthonormal predictors and a sparse true model, we prove that the probability that the best possible predictive performance of the Lasso deteriorates as p increases is positive and can be arbitrarily close to one given a sufficiently high signal to noise ratio and sufficiently large p. We further demonstrate empirically that the amount of deterioration in performance can be far worse than the oracle inequalities suggest and provide a real data example where deterioration is observed.
Replicating brain activity while we snooze could help to improve our memories
Anyone who has had a bad night's sleep will know that the impact it can have on your ability to think clearly and remember things the following day. But it appears that we may be able to improve our retrieval of memories, by simulating how our brains work during sleep. New research has discovered the brain circuit that controls how certain memories are consolidated in the brain overnight. Researchers from the RIKEN Brain Science Institute in Japan have now shown how manipulating a specific brain circuit can prevent or enhance the memories you retain. Masanori Murayama, who led the study, said: 'Our findings on sleep deprivation are particularly interesting from a clinical perspective.
Deep learning applied to drug discovery and repurposing
IMAGE: This image shows deep neural networks for drug discovery. Deep learning, frequently referred to as artificial intelligence, a branch of machine learning utilizing multiple layers of neurons to model high-level abstractions in data, has outperformed humans in tasks including image, text and voice recognition, autonomous driving and others, and is now being applied to drug discovery and biomarker development. In a study published in Molecular Pharmaceutics, a prestigious journal published by the American Chemical Society, scientists from Insilico Medicine in collaboration with Datalytic Solutions and Mind Research Network trained deep neural networks to predict the therapeutic use of large number of multiple drugs using gene expression data obtained from high-throughput experiments on human cell lines. Deep neural networks outperformed other machine learning techniques and did not result in significant drop in performance as the number of classes increased. When the networks got confused and guessed the therapeutic use of the drugs incorrectly, the drugs often had dual use, indicating the possibility of using DNNs for drug repurposing.
Power BI & Azure ML Better Together
There has been a lot of interest in the analytics community in visualizing the output of an Azure Machine Learning model inside Power BI. To add to the challenge, it would also be great to operationalize Azure ML models through the Power BI service. Imagine if you could have Power BI regularly bring in the latest output of your fraud model or the sentiment for recent Tweets about your products. The following tutorial will outline a proposed approach for doing just that. For the purpose of this tutorial we will assume your data is sitting inside an Azure SQL database.