Goto

Collaborating Authors

 Europe


Drone used to search for escaped Borth lynx

BBC News

A heat-seeking drone is being used to hunt a lynx which went missing from a zoo more than a week ago. Lilleth the Eurasian lynx escaped from her enclosure at Borth Wild Animal Kingdom near Aberystwyth. The drone has a specialist night scope and thermal cameras which zoo staff searching for her hope will help pinpoint her location. So far Lilleth has evaded police helicopters, tracking devices and traps. Staff said the lynx's brother Tyrion, who also lives at the zoo, has been pining for her every night and calling out to her.


Boltzmann Exploration Done Right

arXiv.org Machine Learning

Boltzmann exploration is a classic strategy for sequential decision-making under uncertainty, and is one of the most standard tools in Reinforcement Learning (RL). Despite its widespread use, there is virtually no theoretical understanding about the limitations or the actual benefits of this exploration scheme. Does it drive exploration in a meaningful way? Is it prone to misidentifying the optimal actions or spending too much time exploring the suboptimal ones? What is the right tuning for the learning rate? In this paper, we address several of these questions in the classic setup of stochastic multi-armed bandits. One of our main results is showing that the Boltzmann exploration strategy with any monotone learning-rate sequence will induce suboptimal behavior. As a remedy, we offer a simple non-monotone schedule that guarantees near-optimal performance, albeit only when given prior access to key problem parameters that are typically not available in practical situations (like the time horizon $T$ and the suboptimality gap $\Delta$). More importantly, we propose a novel variant that uses different learning rates for different arms, and achieves a distribution-dependent regret bound of order $\frac{K\log^2 T}{\Delta}$ and a distribution-independent bound of order $\sqrt{KT}\log K$ without requiring such prior knowledge. To demonstrate the flexibility of our technique, we also propose a variant that guarantees the same performance bounds even if the rewards are heavy-tailed.


UCB Exploration via Q-Ensembles

arXiv.org Machine Learning

We show how an ensemble of $Q^*$-functions can be leveraged for more effective exploration in deep reinforcement learning. We build on well established algorithms from the bandit setting, and adapt them to the $Q$-learning setting. We propose an exploration strategy based on upper-confidence bounds (UCB). Our experiments show significant gains on the Atari benchmark.


A Tutorial on Canonical Correlation Methods

arXiv.org Machine Learning

Canonical correlation analysis is a family of multivariate statistical methods for the analysis of paired sets of variables. Since its proposition, canonical correlation analysis has for instance been extended to extract relations between two sets of variables when the sample size is insufficient in relation to the data dimensionality, when the relations have been considered to be non-linear, and when the dimensionality is too large for human interpretation. This tutorial explains the theory of canonical correlation analysis including its regularised, kernel, and sparse variants. Additionally, the deep and Bayesian CCA extensions are briefly reviewed. Together with the numerical examples, this overview provides a coherent compendium on the applicability of the variants of canonical correlation analysis. By bringing together techniques for solving the optimisation problems, evaluating the statistical significance and generalisability of the canonical correlation model, and interpreting the relations, we hope that this article can serve as a hands-on tool for applying canonical correlation methods in data analysis.


Identification of Gaussian Process State Space Models

arXiv.org Machine Learning

The Gaussian process state space model (GPSSM) is a non-linear dynamical system, where unknown transition and/or measurement mappings are described by GPs. Most research in GPSSMs has focussed on the state estimation problem, i.e., computing a posterior of the latent state given the model. However, the key challenge in GPSSMs has not been satisfactorily addressed yet: system identification, i.e., learning the model. To address this challenge, we impose a structured Gaussian variational posterior distribution over the latent states, which is parameterised by a recognition model in the form of a bi-directional recurrent neural network. Inference with this structure allows us to recover a posterior smoothed over sequences of data. We provide a practical algorithm for efficiently computing a lower bound on the marginal likelihood using the reparameterisation trick. This further allows for the use of arbitrary kernels within the GPSSM. We demonstrate that the learnt GPSSM can efficiently generate plausible future trajectories of the identified system after only observing a small number of episodes from the true system.


Learning the distribution with largest mean: two bandit frameworks

arXiv.org Machine Learning

Over the past few years, the multi-armed bandit model has become increasingly popular in the machine learning community, partly because of applications including online content optimization. This paper reviews two different sequential learning tasks that have been considered in the bandit literature ; they can be formulated as (sequentially) learning which distribution has the highest mean among a set of distributions, with some constraints on the learning process. For both of them (regret minimization and best arm identification) we present recent, asymptotically optimal algorithms. We compare the behaviors of the sampling rule of each algorithm as well as the complexity terms associated to each problem.


Convex Optimization with Nonconvex Oracles

arXiv.org Machine Learning

In machine learning and optimization, one often wants to minimize a convex objective function $F$ but can only evaluate a noisy approximation $\hat{F}$ to it. Even though $F$ is convex, the noise may render $\hat{F}$ nonconvex, making the task of minimizing $F$ intractable in general. As a consequence, several works in theoretical computer science, machine learning and optimization have focused on coming up with polynomial time algorithms to minimize $F$ under conditions on the noise $F(x)-\hat{F}(x)$ such as its uniform-boundedness, or on $F$ such as strong convexity. However, in many applications of interest, these conditions do not hold. Here we show that, if the noise has magnitude $\alpha F(x) + \beta$ for some $\alpha, \beta > 0$, then there is a polynomial time algorithm to find an approximate minimizer of $F$. In particular, our result allows for unbounded noise and generalizes those of Applegate and Kannan, and Zhang, Liang and Charikar, who proved similar results for the bounded noise case, and that of Belloni et al. who assume that the noise grows in a very specific manner and that $F$ is strongly convex. Turning our result on its head, one may also view our algorithm as minimizing a nonconvex function $\hat{F}$ that is promised to be related to a convex function $F$ as above. Our algorithm is a "simulated annealing" modification of the stochastic gradient Langevin Markov chain and gradually decreases the temperature of the chain to approach the global minimizer. Analyzing such an algorithm for the unbounded noise model and a general convex function turns out to be challenging and requires several technical ideas that might be of independent interest in deriving non-asymptotic bounds for other simulated annealing based algorithms.


Facebook Messenger's money transfer tool is heading to the UK

Engadget

Back in 2015, Facebook introduced the ability to send money to friends through Messenger and now it has brought that capability to UK users. It's the first time Facebook has launched the feature outside of the US. A number of companies have begun working peer-to-peer payment abilities into their services. Skype lets users in nearly two dozen countries send cash within its mobile app via PayPal and PayPal has a bot that let's you send money within Slack. In May, the encrypted messaging app Telegram began supporting payments through chatbots, as did Facebook last year.


1000 local events expected during the European Robotics Week 2017

Robohub

The importance of robotics for Europe's regions will be the focus of a week-long celebration of robotics taking place around Europe on 17–27 November 2017. The European Robotics Week 2017 (ERW2017) is expected to include more than 1000 local events for the public -- open days by factories and research laboratories, school visits by robots, talks by experts and robot competitions are just some of the events. Robotics is increasingly important in education. "Since 2011, we have been asking schools throughout all regions of Europe to demonstrate robotics education at all levels," says Reinhard Lafrenz, the Secretary General of euRobotics, the association for robotics researchers and industry which organises ERW2017. "I am delighted that many skilled teachers and enthusiastic local organisers have taken up this challenge and we have seen huge success in participation, with over 1000 events expected to be organised in all regions of Europe this year."


Localis report warns of job losses because of robots

Daily Mail - Science & tech

England's poorest towns are set to be hit hardest by the rise of robots and automation, a new report today warns. Large swathes of the country face potentially devastating job losses because of a failure to equip workers with the skills for industries of the future, the paper says. The report, by the think-tank Localis, warned of a'staggering gulf' in how well-prepared different parts of the country are to adapt to the looming change in jobs. And it called for the government to devolve polices like apprenticeships and skills training to help different areas re-train people in a bid to prevent mass job losses. The report said England's economic prosperity faces a triple threat from automation, poor skills and a potential drop in migrant labour after Brexit. The vast bulk of funding is poured into the'golden triangle' of London, Oxford and Cambridge - leaving them well-equipped to take advantage of the change.