Genre
Free Data Science eBooks - February 2017
Reinforcement learning, one of the most active research areas in artificial intelligence, is a computational approach to learning whereby an agent tries to maximize the total amount of reward it receives when interacting with a complex, uncertain environment. In Reinforcement Learning, Richard Sutton and Andrew Barto provide a clear and simple account of the key ideas and algorithms of reinforcement learning. The only necessary mathematical background is familiarity with elementary concepts of probability. The book is divided into three parts. Part I defines the reinforcement learning problem in terms of Markov decision processes.
Google brings AI to Raspberry Pi
Dear colleagues: Advances in Artificial Intelligence (AI) technology have opened up new markets and new opportunities for progress in critical areas such as health, education, energy, and the environment. In recent years, machines have surpassed humans in the performance of certain specific tasks, such as some aspects of image recognition. Experts forecast that rapid progress in the field of specialized artificial intelligence will continue. Although it is very unlikely that machines will exhibit broadly-applicable intelligence comparable to or exceeding that of humans in the next 20 years, it is to be expected that machines will reach and exceed human performance on more and more tasks. As a contribution toward preparing the United States for a future in which AI plays a growing role, this report surveys the current state of AI, its existing and potential applications, and the questions that are raised for society and public policy by progress in AI. The report also makes recommendations for specific further actions by Federal agencies and other actors. A companion document lays out a strategic plan for Federally-funded research and development in AI. Additionally, in the coming months, the Administration will release a follow-on report exploring in greater depth the effect of AI-driven automation on jobs and the economy.
This visual recognition startup has poached AI talent from Twitter
Clarifai, a startup that creates visual recognition software, has hired four members of Twitter's machine learning team, Cortex, as well as an engineer formerly with Google Brain, the search giant's research group focused on artificial intelligence. Founded in 2013 by computer science PhD Matthew Zeiler after he did an internship with the Google research team, Clarifai licenses customizable software that can automatically organize and filter images. The startup's clients include BuzzFeed, travel site Trivago and consumer-packaged goods giant Unilever. The 40-employee startup (including new hires) has raised $41 million, including $30 million in a Series B round led by Menlo Ventures. The round included contributions from Union Square Ventures, Lux Capital and Qualcomm Ventures.
Towards Better Analysis of Machine Learning Models: A Visual Analytics Perspective
Liu, Shixia, Wang, Xiting, Liu, Mengchen, Zhu, Jun
Interactive model analysis, the process of understanding, diagnosing, and refining a machine learning model with the help of interactive visualization, is very important for users to efficiently solve real-world artificial intelligence and data mining problems. Dramatic advances in big data analytics has led to a wide variety of interactive model analysis tasks. In this paper, we present a comprehensive analysis and interpretation of this rapidly developing area. Specifically, we classify the relevant work into three categories: understanding, diagnosis, and refinement. Each category is exemplified by recent influential work. Possible future research opportunities are also explored and discussed. Keywords: interactive model analysis, interactive visualization, machine learning, understanding, diagnosis, refinement 1. Introduction Machine learning has been successfully applied to a wide variety of fields ranging from information retrieval, data mining, and speech recognition, to computer graphics, visualization, and human-computer interaction. However, most users often treat a machine learning model as a black box because of its incomprehensible functions and unclear working mechanism [1, 2, 3]. Without a clear understanding of how and why a model works, the development of highperformance models typically relies on a time-consuming trial-and-error pro-Fully documented templates are available in the elsarticle package on CTAN.
Probabilistic Sensor Fusion for Ambient Assisted Living
Diethe, Tom, Twomey, Niall, Kull, Meelis, Flach, Peter, Craddock, Ian
There is a widely-accepted need to revise current forms of healthcare provision, with particular interest in sensing systems in the home. Given a multiple-modality sensor platform with heterogeneous network connectivity, as is under development in the Sensor Platform for HEalthcare in Residential Environment (SPHERE) Interdisciplinary Research Collaboration (IRC), we face specific challenges relating to the fusion of the heterogeneous sensor modalities. We introduce Bayesian models for sensor fusion, which aims to address the challenges of fusion of heterogeneous sensor modalities. Using this approach we are able to identify the modalities that have most utility for each particular activity, and simultaneously identify which features within that activity are most relevant for a given activity. We further show how the two separate tasks of location prediction and activity recognition can be fused into a single model, which allows for simultaneous learning an prediction for both tasks. We analyse the performance of this model on data collected in the SPHERE house, and show its utility. We also compare against some benchmark models which do not have the full structure, and show how the proposed model compares favourably to these methods.
Query Efficient Posterior Estimation in Scientific Experiments via Bayesian Active Learning
Kandasamy, Kirthevasan, Schneider, Jeff, Póczos, Barnabás
A common problem in disciplines of applied Statistics research such as Astrostatistics is of estimating the posterior distribution of relevant parameters. Typically, the likelihoods for such models are computed via expensive experiments such as cosmological simulations of the universe. An urgent challenge in these research domains is to develop methods that can estimate the posterior with few likelihood evaluations. In this paper, we study active posterior estimation in a Bayesian setting when the likelihood is expensive to evaluate. Existing techniques for posterior estimation are based on generating samples representative of the posterior. Such methods do not consider efficiency in terms of likelihood evaluations. In order to be query efficient we treat posterior estimation in an active regression framework. We propose two myopic query strategies to choose where to evaluate the likelihood and implement them using Gaussian processes. Via experiments on a series of synthetic and real examples we demonstrate that our approach is significantly more query efficient than existing techniques and other heuristics for posterior estimation.
Energy Prediction using Spatiotemporal Pattern Networks
Jiang, Zhanhong, Liu, Chao, Akintayo, Adedotun, Henze, Gregor, Sarkar, Soumik
This paper presents a novel data-driven technique based on the spatiotemporal pattern network (STPN) for energy/power prediction for complex dynamical systems. Built on symbolic dynamic filtering, the STPN framework is used to capture not only the individual system characteristics but also the pair-wise causal dependencies among different sub-systems. For quantifying the causal dependency, a mutual information based metric is presented. An energy prediction approach is subsequently proposed based on the STPN framework. For validating the proposed scheme, two case studies are presented, one involving wind turbine power prediction (supply side energy) using the Western Wind Integration data set generated by the National Renewable Energy Laboratory (NREL) for identifying the spatiotemporal characteristics, and the other, residential electric energy disaggregation (demand side energy) using the Building America 2010 data set from NREL for exploring the temporal features. In the energy disaggregation context, convex programming techniques beyond the STPN framework are developed and applied to achieve improved disaggregation performance.
Detecting changes in slope with an $L_0$ penalty
Maidstone, Robert, Fearnhead, Paul, Letchford, Adam
Whilst there are many approaches to detecting changes in mean for a univariate time-series, the problem of detecting multiple changes in slope has comparatively been ignored. Part of the reason for this is that detecting changes in slope is much more challenging. For example, simple binary segmentation procedures do not work for this problem, whilst efficient dynamic programming methods that work well for the change in mean problem cannot be directly used for detecting changes in slope. We present a novel dynamic programming approach, CPOP, for finding the "best" continuous piecewise-linear fit to data. We define best based on a criterion that measures fit to data using the residual sum of squares, but penalises complexity based on an $L_0$ penalty on changes in slope. We show that using such a criterion is more reliable at estimating changepoint locations than approaches that penalise complexity using an $L_1$ penalty. Empirically CPOP has good computational properties, and can analyse a time-series with over 10,000 observations and over 100 changes in a few minutes. Our method is used to analyse data on the motion of bacteria, and provides fits to the data that both have substantially smaller residual sum of squares and are more parsimonious than two competing approaches.
Edge-exchangeable graphs and sparsity (NIPS 2016)
Cai, Diana, Campbell, Trevor, Broderick, Tamara
Many popular network models rely on the assumption of (vertex) exchangeability, in which the distribution of the graph is invariant to relabelings of the vertices. However, the Aldous-Hoover theorem guarantees that these graphs are dense or empty with probability one, whereas many real-world graphs are sparse. We present an alternative notion of exchangeability for random graphs, which we call edge exchangeability, in which the distribution of a graph sequence is invariant to the order of the edges. We demonstrate that edge-exchangeable models, unlike models that are traditionally vertex exchangeable, can exhibit sparsity. To do so, we outline a general framework for graph generative models; by contrast to the pioneering work of Caron and Fox (2015), models within our framework are stationary across steps of the graph sequence. In particular, our model grows the graph by instantiating more latent atoms of a single random measure as the dataset size increases, rather than adding new atoms to the measure.
New method for making effective calculations in 'high-dimensional space'
Researchers have developed a new method for making effective calculations in "high-dimensional space" – and proved its worth by using it to solve a 93-dimensional problem. Researchers have developed a new technique for making calculations in "high-dimensional space" – mathematical problems so wide-ranging in their scope, that they seem at first to be beyond the limits of human calculation. In what sounds like the title of a rejected script for an Indiana Jones movie, the method improves on existing approaches to beat a well-known problem known as "The Curse Of Dimensionality". It was devised by a team of researchers at the University of Cambridge. In rough terms, the "Curse", refers to the apparent impossibility of making calculations in situations where the number of variables, attributes, and possible outcomes is so large that it seems futile even to try to comprehend the problem in the first place.