Goto

Collaborating Authors

 Europe


Rover Descent: Learning to optimize by learning to navigate on prototypical loss surfaces

arXiv.org Machine Learning

Learning to optimize - the idea that we can learn from data algorithms that optimize a numerical criterion - has recently been at the heart of a growing number of research efforts. One of the most challenging issues within this approach is to learn a policy that is able to optimize over classes of functions that are fairly different from the ones that it was trained on. We propose a novel way of framing learning to optimize as a problem of learning a good navigation policy on a partially observable loss surface. To this end, we develop Rover Descent, a solution that allows us to learn a fairly broad optimization policy from training on a small set of prototypical two-dimensional surfaces that encompasses the classically hard cases such as valleys, plateaus, cliffs and saddles and by using strictly zero-order information. We show that, without having access to gradient or curvature information, we achieve state-of-the-art convergence speed on optimization problems not presented at training time such as the Rosenbrock function and other hard cases in two dimensions. We extend our framework to optimize over high dimensional landscapes, while still handling only two-dimensional local landscape information and show good preliminary results.


On the Statistical Challenges of Echo State Networks and Some Potential Remedies

arXiv.org Machine Learning

Echo state networks are powerful recurrent neural networks. However, they are often unstable and shaky, making the process of finding an good ESN for a specific dataset quite hard. Obtaining a superb accuracy by using the Echo State Network is a challenging task. We create, develop and implement a family of predictably optimal robust and stable ensemble of Echo State Networks via regularizing the training and perturbing the input. Furthermore, several distributions of weights have been tried based on the shape to see if the shape of the distribution has the impact for reducing the error. We found ESN can track in short term for most dataset, but it collapses in the long run. Short-term tracking with large size reservoir enables ESN to perform strikingly with superior prediction. Based on this scenario, we go a further step to aggregate many of ESNs into an ensemble to lower the variance and stabilize the system by stochastic replications and bootstrapping of input data.


AutoPrognosis: Automated Clinical Prognostic Modeling via Bayesian Optimization with Structured Kernel Learning

arXiv.org Machine Learning

Clinical prognostic models derived from largescale healthcare data can inform critical diagnostic and therapeutic decisions. To enable off-theshelf usage of machine learning (ML) in prognostic research, we developed AUTOPROGNOSIS: a system for automating the design of predictive modeling pipelines tailored for clinical prognosis. AUTOPROGNOSIS optimizes ensembles of pipeline configurations efficiently using a novel batched Bayesian optimization (BO) algorithm that learns a low-dimensional decomposition of the pipelines high-dimensional hyperparameter space in concurrence with the BO procedure. This is achieved by modeling the pipelines performances as a black-box function with a Gaussian process prior, and modeling the similarities between the pipelines baseline algorithms via a sparse additive kernel with a Dirichlet prior. Meta-learning is used to warmstart BO with external data from similar patient cohorts by calibrating the priors using an algorithm that mimics the empirical Bayes method. The system automatically explains its predictions by presenting the clinicians with logical association rules that link patients features to predicted risk strata. We demonstrate the utility of AUTOPROGNOSIS using 10 major patient cohorts representing various aspects of cardiovascular patient care.


High-Quality Prediction Intervals for Deep Learning: A Distribution-Free, Ensembled Approach

arXiv.org Machine Learning

Deep neural networks are a powerful technique for learning complex functions from data. However, their appeal in real-world applications can be hindered by an inability to quantify the uncertainty of predictions. In this paper, the generation of prediction intervals (PI) for quantifying uncertainty in regression tasks is considered. It is axiomatic that high-quality PIs should be as narrow as possible, whilst capturing a specified portion of data. In this paper we derive a loss function directly from this high-quality principle that requires no distributional assumption. We show how its form derives from a likelihood principle, that it can be used with gradient descent, and that in ensembled form, model uncertainty is accounted for. This remedies limitations of a popular model developed on the same high-quality principle. Experiments are conducted on ten regression benchmark datasets. The proposed quality-driven (QD) method is shown to outperform current state-of-the-art uncertainty quantification methods, reducing average PI width by around 10%.


i-RevNet: Deep Invertible Networks

arXiv.org Machine Learning

It is widely believed that the success of deep convolutional networks is based on progressively discarding uninformative variability about the input with respect to the problem at hand. This is supported empirically by the difficulty of recovering images from their hidden representations, in most commonly used network architectures. In this paper we show via a one-to-one mapping that this loss of information is not a necessary condition to learn representations that generalize well on complicated problems, such as ImageNet. Via a cascade of homeomorphic layers, we build thei -RevNet, a network that can be fully inverted up to the final projection onto the classes, i.e. no information is discarded. Building an invertible architecture is difficult, for one, because the local inversion is ill-conditioned, we overcome this by providing an explicit inverse. An analysis of i-RevNets learned representations suggests an alternative explanation for the success of deep networks by a progressive contraction and linear separation with depth. To shed light on the nature of the model learned by thei -RevNet we reconstruct linear interpolations between natural image representations. A CNN may be very effective in classifying images of all sorts (He et al., 2016; Krizhevsky et al., 2012), but the cascade of linear and nonlinear operators reveals little about the contribution of the internal representation to the classification. The learning process is characterized by a steady reduction of large amounts of uninformative variability in the images while simultaneously revealing the essence of the visual class.


Information Theory: A Tutorial Introduction

arXiv.org Machine Learning

In 1948, Claude Shannon published a paper called A Mathematical Theory of Communication[1]. This paper heralded a transformation in our understanding of information. Before Shannon's paper, information had been viewed as a kind of poorly defined miasmic fluid. But after Shannon's paper, it became apparent that information is a well-defined and, above all, measurable quantity. Indeed, as noted by Shannon, A basic idea in information theory is that information can be treated very much like a physical quantity, such as mass or energy.


Differentiable Dynamic Programming for Structured Prediction and Attention

arXiv.org Machine Learning

Dynamic programming (DP) solves a variety of structured combinatorial problems by iteratively breaking them down into smaller subproblems. In spite of their versatility, DP algorithms are usually non-differentiable, which hampers their use as a layer in neural networks trained by backpropagation. To address this issue, we propose to smooth the max operator in the dynamic programming recursion, using a strongly convex regularizer. This allows to relax both the optimal value and solution of the original combinatorial problem, and turns a broad class of DP algorithms into differentiable operators. Theoretically, we provide a new probabilistic perspective on backpropagating through these DP operators, and relate them to inference in graphical models. We derive two particular instantiations of our framework, a smoothed Viterbi algorithm for sequence prediction and a smoothed DTW algorithm for time-series alignment. We showcase these instantiations on two structured prediction tasks and on structured and sparse attention for neural machine translation.


Logic Programming Applications: What Are the Abstractions and Implementations?

arXiv.org Artificial Intelligence

This article presents an overview of applications of logic programming, classifying them based on the abstractions and implementations of logic languages that support the applications. The three key abstractions are join, recursion, and constraint. Their essential implementations are for-loops, fixed points, and backtracking, respectively. The corresponding kinds of applications are database queries, inductive analysis, and combinatorial search, respectively. We also discuss language extensions and programming paradigms, summarize example application problems by application areas, and touch on example systems that support variants of the abstractions with different implementations.


BBC micro:bit beat an iPhone using Siri in a race

Daily Mail - Science & tech

Apple's mighty iPhone had a chunk bitten out of its reputation, after losing in a computer power test to a £12.99 ($18) IT education aid. In a David versus Goliath battle, the BBC's Micro:bit beat the Siri-enabled smartphone in a race between eight computers from the last 75 years. Each device was given 15 seconds to generate as many numbers as possible from the Fibonacci sequence, where each number is the sum of the previous two. The low-spec teaching tool, given to schoolchildren for free, came out on top while the iPhone languished second to last. Apple's mighty iPhone had a chunk bitten out of its reputation, after losing in a computer power test to a £12.99 ($18) IT education aid.


Mike Olson on Cloudera, Machine Learning, and the Adoption of the Cloud

#artificialintelligence

At the Cloudera Sessions event in Munich, Germany, Paige Roberts of Syncsort sat down with Mike Olson, Chief Strategy Officer of Cloudera. In this first of a three-part interview, Mike Olson goes into what's new at Cloudera, how machine learning is evolving, and the adoption of the Cloud in organizations. I'm a co-founder and am chief strategy officer of the company and I'm excited to speak with you, Paige. Well, I spent some time in the sessions here in Munich today talking about what we're seeing in the adoption of machine learning and some of these advanced analytic techniques. That's really exciting, like the use cases that are getting built, using these new analytic techniques…it's pretty awesome. I mean in healthcare, diagnosing disease better than ever before, delivering better treatments.