Genre
Learning an Astronomical Catalog of the Visible Universe through Scalable Bayesian Inference
Regier, Jeffrey, Pamnany, Kiran, Giordano, Ryan, Thomas, Rollin, Schlegel, David, McAuliffe, Jon, Prabhat, null
Celeste is a procedure for inferring astronomical catalogs that attains state-of-the-art scientific results. To date, Celeste has been scaled to at most hundreds of megabytes of astronomical images: Bayesian posterior inference is notoriously demanding computationally. In this paper, we report on a scalable, parallel version of Celeste, suitable for learning catalogs from modern large-scale astronomical datasets. Our algorithmic innovations include a fast numerical optimization routine for Bayesian posterior inference and a statistically efficient scheme for decomposing astronomical optimization problems into subproblems. Our scalable implementation is written entirely in Julia, a new high-level dynamic programming language designed for scientific and numerical computing. We use Julia's high-level constructs for shared and distributed memory parallelism, and demonstrate effective load balancing and efficient scaling on up to 8192 Xeon cores on the NERSC Cori supercomputer.
Disentangling factors of variation in deep representations using adversarial training
Mathieu, Michael, Zhao, Junbo, Sprechmann, Pablo, Ramesh, Aditya, LeCun, Yann
We introduce a conditional generative model for learning to disentangle the hidden factors of variation within a set of labeled observations, and separate them into complementary codes. One code summarizes the specified factors of variation associated with the labels. The other summarizes the remaining unspecified variability. During training, the only available source of supervision comes from our ability to distinguish among different observations belonging to the same class. Examples of such observations include images of a set of labeled objects captured at different viewpoints, or recordings of set of speakers dictating multiple phrases. In both instances, the intra-class diversity is the source of the unspecified factors of variation: each object is observed at multiple viewpoints, and each speaker dictates multiple phrases. Learning to disentangle the specified factors from the unspecified ones becomes easier when strong supervision is possible. Suppose that during training, we have access to pairs of images, where each pair shows two different objects captured from the same viewpoint. This source of alignment allows us to solve our task using existing methods. However, labels for the unspecified factors are usually unavailable in realistic scenarios where data acquisition is not strictly controlled. We address the problem of disentanglement in this more general setting by combining deep convolutional autoencoders with a form of adversarial training. Both factors of variation are implicitly captured in the organization of the learned embedding space, and can be used for solving single-image analogies. Experimental results on synthetic and real datasets show that the proposed method is capable of generalizing to unseen classes and intra-class variabilities.
Distributed Estimation and Learning over Heterogeneous Networks
Rahimian, M. Amin, Jadbabaie, Ali
We consider several estimation and learning problems that networked agents face when making decisions given their uncertainty about an unknown variable. Our methods are designed to efficiently deal with heterogeneity in both size and quality of the observed data, as well as heterogeneity over time (intermittence). The goal of the studied aggregation schemes is to efficiently combine the observed data that is spread over time and across several network nodes, accounting for all the network heterogeneities. Moreover, we require no form of coordination beyond the local neighborhood of every network agent or sensor node. The three problems that we consider are (i) maximum likelihood estimation of the unknown given initial data sets, (ii) learning the true model parameter from streams of data that the agents receive intermittently over time, and (iii) minimum variance estimation of a complete sufficient statistic from several data points that the networked agents collect over time. In each case we rely on an aggregation scheme to combine the observations of all agents; moreover, when the agents receive streams of data over time, we modify the update rules to accommodate the most recent observations. In every case, we demonstrate the efficiency of our algorithms by proving convergence to the globally efficient estimators given the observations of all agents. We supplement these results by investigating the rate of convergence and providing finite-time performance guarantees.
Policy Search with High-Dimensional Context Variables
Tangkaratt, Voot, van Hoof, Herke, Parisi, Simone, Neumann, Gerhard, Peters, Jan, Sugiyama, Masashi
Direct contextual policy search methods learn to improve policy parameters and simultaneously generalize these parameters to different context or task variables. However, learning from high-dimensional context variables, such as camera images, is still a prominent problem in many real-world tasks. A naive application of unsupervised dimensionality reduction methods to the context variables, such as principal component analysis, is insufficient as task-relevant input may be ignored. In this paper, we propose a contextual policy search method in the model-based relative entropy stochastic search framework with integrated dimensionality reduction. We learn a model of the reward that is locally quadratic in both the policy parameters and the context variables. Furthermore, we perform supervised linear dimensionality reduction on the context variables by nuclear norm regularization. The experimental results show that the proposed method outperforms naive dimensionality reduction via principal component analysis and a state-of-the-art contextual policy search method.
Feature Selection with the R Package MXM: Discovering Statistically-Equivalent Feature Subsets
Lagani, Vincenzo, Athineou, Giorgos, Farcomeni, Alessio, Tsagris, Michail, Tsamardinos, Ioannis
The statistically equivalent signature (SES) algorithm is a method for feature selection inspired by the principles of constrained-based learning of Bayesian Networks. Most of the currently available feature-selection methods return only a single subset of features, supposedly the one with the highest predictive power. We argue that in several domains multiple subsets can achieve close to maximal predictive accuracy, and that arbitrarily providing only one has several drawbacks. The SES method attempts to identify multiple, predictive feature subsets whose performances are statistically equivalent. Under that respect SES subsumes and extends previous feature selection algorithms, like the max-min parent children algorithm. SES is implemented in an homonym function included in the R package MXM, standing for mens ex machina, meaning 'mind from the machine' in Latin. The MXM implementation of SES handles several data-analysis tasks, namely classification, regression and survival analysis. In this paper we present the SES algorithm, its implementation, and provide examples of use of the SES function in R. Furthermore, we analyze three publicly available data sets to illustrate the equivalence of the signatures retrieved by SES and to contrast SES against the state-of-the-art feature selection method LASSO. Our results provide initial evidence that the two methods perform comparably well in terms of predictive accuracy and that multiple, equally predictive signatures are actually present in real world data.
Low Data Drug Discovery with One-shot Learning
Altae-Tran, Han, Ramsundar, Bharath, Pappu, Aneesh S., Pande, Vijay
Recent advances in machine learning have made significant contributions to drug discovery. Deep neural networks in particular have been demonstrated to provide significant boosts in predictive power when inferring the properties and activities of small-molecule compounds. However, the applicability of these techniques has been limited by the requirement for large amounts of training data. In this work, we demonstrate how one-shot learning can be used to significantly lower the amounts of data required to make meaningful predictions in drug discovery applications. We introduce a new architecture, the residual LSTM embedding, that, when combined with graph convolutional neural networks, significantly improves the ability to learn meaningful distance metrics over small-molecules. We open source all models introduced in this work as part of DeepChem, an open-source framework for deep-learning in drug discovery.
Learning to Reason With Adaptive Computation
Neumann, Mark, Stenetorp, Pontus, Riedel, Sebastian
Multi-hop inference is necessary for machine learning systems to successfully solve tasks such as Recognising Textual Entailment and Machine Reading. In this work, we demonstrate the effectiveness of adaptive computation for learning the number of inference steps required for examples of different complexity and that learning the correct number of inference steps is difficult. We introduce the first model involving Adaptive Computation Time which provides a small performance benefit on top of a similar model without an adaptive component as well as enabling considerable insight into the reasoning process of the model.
New wireless device helps paralyzed monkeys regain use of their legs
GENEVA – A new device has allowed two monkeys to regain use of their paralyzed legs by transmitting brain signals wirelessly, bypassing their spinal cord lesions, a study released Wednesday by the journal Nature said. The implantable device, called a neuroprosthetic interface, was developed by an international team led by researchers at the Federal Polytechnic School of Lausanne (EPFL) and may soon be tested as a remedy for paralysis in humans. "For the first time, I can imagine a completely paralyzed patient able to move their legs through this brain-spine interface," Jocelyne Bloch, a neurosurgeon at the Lausanne University Hospital, said in a press release from EPFL. The interface conceived at EPFL is a multicomponent brain-spine connector, which decodes signals from the part of the motor cortex responsible for leg movements. It then relays those signals in real time to the lumbar region of the spinal cord that activates leg muscles to walk.
British Artificial Intelligence will soon run Clinical Trials
BenevolentAI is a London start-up that specializes in artificial intelligence (AI); its BenevolentBio division, formerly Stratified Medical, applies AI to human health and biotech. Its baby is a technology that could speed up late-stage development of drugs and provide richer clinical data. Now, it will test it using clinical stage drug candidates licensed from Janssen. Although there are no details on the particulars of the agreement, BenevolentAI is confident that it can accelerate clinical development and begin Phase IIb trials in mid-2017. If everything works out well, the company will have exclusive rights to develop, manufacture and commercialize these candidates.
A Machine Learning Approach to Identifying the Thought Markers of Suicidal Subjects: A Prospective Multicenter Trial - Pestian - 2016 - Suicide and Life-Threatening Behavior - Wiley Online Library
Efforts to understand suicide risks can be roughly clustered into traits or states. Trait analyses focus on stable characteristics rooted in and measured using biological processes (Costanza et al., 2014; Le-Niculescu et al., 2013), whereas state analyses measure dynamic characteristics like verbal and nonverbal communication, termed "thought markers" (Pestian et al., 2015). Machine learning and natural language processing have successfully identified differences in retrospective suicide notes, newsgroups, and social media (Gomez, 2014; Huang, Goh, & Liew, 2007; Matykiewicz, Duch, & Pestian, 2009). Jashinsky et al. (2015) used multiple annotators to identify the risk of suicide from the keywords and phrases (interrater reliability .79) in geographically based tweets. Thompson, Poulin, and Bryan (2014) and Desmet (2014) used text-based signals to identify suicide risk that ranged from 60% to 90%.