Europe
Ordered Preference Elicitation Strategies for Supporting Multi-Objective Decision Making
Zintgraf, Luisa M, Roijers, Diederik M, Linders, Sjoerd, Jonker, Catholijn M, Nowé, Ann
In multi-objective decision planning and learning, much attention is paid to producing optimal solution sets that contain an optimal policy for every possible user preference profile. We argue that the step that follows, i.e, determining which policy to execute by maximising the user's intrinsic utility function over this (possibly infinite) set, is under-studied. This paper aims to fill this gap. We build on previous work on Gaussian processes and pairwise comparisons for preference modelling, extend it to the multi-objective decision support scenario, and propose new ordered preference elicitation strategies based on ranking and clustering. Our main contribution is an in-depth evaluation of these strategies using computer and human-based experiments. We show that our proposed elicitation strategies outperform the currently used pairwise methods, and found that users prefer ranking most. Our experiments further show that utilising monotonicity information in GPs by using a linear prior mean at the start and virtual comparisons to the nadir and ideal points, increases performance. We demonstrate our decision support framework in a real-world study on traffic regulation, conducted with the city of Amsterdam.
Continual Lifelong Learning with Neural Networks: A Review
Parisi, German I., Kemker, Ronald, Part, Jose L., Kanan, Christopher, Wermter, Stefan
Humans and animals have the ability to continually acquire and fine-tune knowledge throughout their lifespan. This ability is mediated by a rich set of neurocognitive functions that together contribute to the early development and experience-driven specialization of our sensorimotor skills. Consequently, the ability to learn from continuous streams of information is crucial for computational learning systems and autonomous agents (inter)acting in the real world. However, continual lifelong learning remains a long-standing challenge for machine learning and neural network models since the incremental acquisition of new skills from non-stationary data distributions generally leads to catastrophic forgetting or interference. This limitation represents a major drawback also for state-of-the-art deep neural network models that typically learn representations from stationary batches of training data, thus without accounting for situations in which the number of tasks is not known a priori and the information becomes incrementally available over time. In this review, we critically summarize the main challenges linked to continual lifelong learning for artificial learning systems and compare existing neural network approaches that alleviate, to different extents, catastrophic interference. Although significant advances have been made in domain-specific continual lifelong learning with neural networks, extensive research efforts are required for the development of general-purpose artificial intelligence and autonomous agents. We discuss well-established research and recent methodological trends motivated by experimentally observed lifelong learning factors in biological systems. Such factors include principles of neurosynaptic stability-plasticity, critical developmental stages, intrinsically motivated exploration, transfer learning, and crossmodal integration.
A Generative Deep Recurrent Model for Exchangeable Data
Korshunova, Iryna, Degrave, Jonas, Huszár, Ferenc, Gal, Yarin, Gretton, Arthur, Dambre, Joni
We present a novel model architecture which leverages deep learning tools to perform exact Bayesian inference on sets of high dimensional, complex observations. Our model is provably exchangeable, meaning that the joint distribution over observations is invariant under permutation: this property lies at the heart of Bayesian inference. The model does not require variational approximations to train, and new samples can be generated conditional on previous samples, with cost linear in the size of the conditioning set. The advantages of our architecture are demonstrated on learning tasks requiring generalisation from short observed sequences while modelling sequence variability, such as conditional image generation, few-shot learning, set completion, and anomaly detection.
The Gaussian Process Autoregressive Regression Model (GPAR)
Requeima, James, Tebbutt, Will, Bruinsma, Wessel, Turner, Richard E.
Multi-output regression models must exploit dependencies between outputs to maximise predictive performance. The application of Gaussian processes (GPs) to this setting typically yields models that are computationally demanding and have limited representational power. We present the Gaussian Process Autoregressive Regression (GPAR) model, a scalable multi-output GP model that is able to capture nonlinear, possibly input-varying, dependencies between outputs in a simple and tractable way: the product rule is used to decompose the joint distribution over the outputs into a set of conditionals, each of which is modelled by a standard GP. GPAR's efficacy is demonstrated on a variety of synthetic and real-world problems, outperforming existing GP models and achieving state-of-the-art performance on the tasks with existing benchmarks.
L4: Practical loss-based stepsize adaptation for deep learning
Rolinek, Michal, Martius, Georg
We propose a stepsize adaptation scheme for stochastic gradient descent. It operates directly with the loss function and rescales the gradient in order to make fixed predicted progress on the loss. We demonstrate its capabilities by strongly improving the performance of Adam and Momentum optimizers. The enhanced optimizers with default hyperparameters consistently outperform their constant stepsize counterparts, even the best ones, without a measurable increase in computational cost. The performance is validated on multiple architectures including ResNets and the Differential Neural Computer. A prototype implementation as a TensorFlow optimizer is released.
Combinatorial Preconditioners for Proximal Algorithms on Graphs
Möllenhoff, Thomas, Ye, Zhenzhang, Wu, Tao, Cremers, Daniel
We present a novel preconditioning technique for proximal optimization methods that relies on graph algorithms to construct effective preconditioners. Such combinatorial preconditioners arise from partitioning the graph into forests. We prove that certain decompositions lead to a theoretically optimal condition number. We also show how ideal decompositions can be realized using matroid partitioning and propose efficient greedy variants thereof for large-scale problems. Coupled with specialized solvers for the resulting scaled proximal subproblems, the preconditioned algorithm achieves competitive performance in machine learning and vision applications.
High-dimensional Bayesian inference via the Unadjusted Langevin Algorithm
We consider in this paper the problem of sampling a high-dimensional probability distribution $\pi$ having a density \wrt\ the Lebesgue measure on $\mathbb{R}^d$, known up to a normalization factor $x \mapsto \pi(x)= \mathrm{e}^{-U(x)}/\int_{\mathbb{R}^d} \mathrm{e}^{-U(y)} \mathrm{d}y$. Such problem naturally occurs for example in Bayesian inference and machine learning. Under the assumption that $U$ is continuously differentiable, $\nabla U$ is globally Lipschitz and $U$ is strongly convex, we obtain non-asymptotic bounds for the convergence to stationarity in Wasserstein distance of order $2$ and total variation distance of the sampling method based on the Euler discretization of the Langevin stochastic differential equation, for both constant and decreasing step sizes. The dependence on the dimension of the state space of the obtained bounds is studied to demonstrate the applicability of this method. The convergence of an appropriately weighted empirical measure is also investigated and bounds for the mean square error and exponential deviation inequality are reported for functions which are measurable and bounded. An illustration to Bayesian inference for binary regression is presented.
VBALD - Variational Bayesian Approximation of Log Determinants
Granziol, Diego, Wagstaff, Edward, Ru, Bin Xin, Osborne, Michael, Roberts, Stephen
Evaluating the log determinant of a positive definite matrix is ubiquitous in machine learning. Applications thereof range from Gaussian processes, minimum-volume ellipsoids, metric learning, kernel learning, Bayesian neural networks, Determinental Point Processes, Markov random fields to partition functions of discrete graphical models. In order to avoid the canonical, yet prohibitive, Cholesky $\mathcal{O}(n^{3})$ computational cost, we propose a novel approach, with complexity $\mathcal{O}(n^{2})$, based on a constrained variational Bayes algorithm. We compare our method to Taylor, Chebyshev and Lanczos approaches and show state of the art performance on both synthetic and real-world datasets.
Nearly 70 Percent of Taxpayers Support Use of AI to Improve Accuracy of Filings, Accenture Global Survey Finds
Nearly 70 Percent of Taxpayers Support Use of AI to Improve Accuracy of Filings, Accenture Global Survey Finds'Digital tax assistant' could prove especially beneficial to 40 percent of taxpayers who reported making filing errors NEW YORK; Feb. 20, 2018 – Nearly 70 percent of taxpayers in 12 countries said they would use AI to improve the accuracy of tax filings, according to a new study by Accenture (NYSE: ACN), which also found that more than 40 percent of taxpayers reported making a filing error in the last 24 months. The Accenture Digital Taxpayers Research asked more than 6,500 taxpayers across Europe, Asia-Pacific and North America who interacted with their tax authority in the prior 12 months about their experiences with, attitudes about and expectations of revenue authorities. The findings indicate that in an era in which people around the world expect easy and simple consumer experiences, tax rules and regulations still confuse citizens. For instance, 38 percent of respondents said they are not confident they pay the right amount of tax, and 44 percent said they feel their tax knowledge could be improved. While most respondents said they have limited contact with their revenue authorities after filing a tax form, half (51 percent) reported contacting their revenue authority once or twice in the past year, with 20 percent reporting three or more contacts.
Connecting AI and IoT with blockchain-based platforms
To fully realize the promise of IoT and the wealth of data it can provide, it's essential to connect it to artificial intelligence, aka. There are, however, several challenges associated with both running AI at the edge of networks and sharing databases securely. Combining AI and IoT enables "making a series of valuable predictions about devices that are embedded within the real world," said Christian Catalini, assistant professor of technological innovation, entrepreneurship and strategic management, and founder of the Cryptoeconomics Lab at Massachusetts Institute of Technology. "This allows you to change how you respond to changes in the environment and to new information coming in." But you can only make valuable predictions if you have access to good data. "Blockchain can help create marketplaces for data and their exchange, and also improve data privacy," Catalini added.