Education
Scalable Lifelong Learning with Active Task Selection
Ruvolo, Paul (Bryn Mawr College) | Eaton, Eric (Bryn Mawr College)
The recently developed Efficient Lifelong Learning Algorithm (ELLA) acquires knowledge incrementally over a sequence of tasks, learning a repository of latent model components that are sparsely shared between models. ELLA shows strong performance in comparison to other multi-task learning algorithms, achieving nearly identical performance to batch multi-task learning methods while learning tasks sequentially in three orders of magnitude (over 1,000x) less time. In this paper, we evaluate several curriculum selection methods that allow ELLA to actively select the next task for learning in order to maximize performance on future learning tasks. Through experiments with three real and one synthetic data set, we demonstrate that active curriculum selection allows an agent to learn up to 50% more efficiently than when the agent has no control over the task order.
Towards Pareto Descent Directions in Sampling Experts for Multiple Tasks in an On-Line Learning Paradigm
Ghosh, Shaona (University of Southampton,UK) | Lovell, Chris (University of Southampton) | Gunn, Steve R. (University of Southampton)
In many real-life design problems, there is a requirement to simultaneously balance multiple tasks or objectives in the system that are conflicting in nature, where minimizing one objective causes another to increase in value, thereby resulting in trade-offs between the objectives. For example, in embedded multi-core mobile devices and very large scale data centers, there is a continuous problem of simultaneously balancing interfering goals of maximal power savings and minimal performance delay with varying trade-off values for different application workloads executing on them. Typically, the optimal trade-offs for the executing workloads, lie on a difficult to determine optimal Pareto front. The nature of the problem requires learning over the lifetime of the mobile device or server with continuous evaluation and prediction of the trade-off settings on the system that balances the interfering objectives optimally. Towards this, we propose an on-line learning method, where the weights of experts for addressing the objectives are updated based on a convex combination of their relative performance in addressing all objectives simultaneously. An additional importance vector that assigns relative importance to each objective at every round is used, and is sampled from a convex cone pointed at the origin Our preliminary results show that the convex combination of the importance vector and the gradient of the potential functions of the learner's regret with respect to each objective ensure that in the next round, the drift (instantaneous regret vector), is the Pareto descent direction that enables better convergence to the optimal Pareto front.
Multi-Engine Machine Translation as a Lifelong Machine Learning Problem
Federmann, Christian (German Research Center for Artificial Intelligence)
We describe an approach for multi-engine machine translation that uses machine learning methods to train one or several classifiers for a given set of candidate translations. Contrary to existing approaches in quality estimation which only consider a single translation at a time, we explicitly model pairwise comparison with our feature vectors. We discuss several challenges our method is facing and discuss how lifelong machine learning could be applied to resolve these. We also show how the proposed architecture can be extended to allow human feedback to be included into the training process, improving the system's selection process over time.
Scalable Lifelong Learning with Active Task Selection
Ruvolo, Paul (Bryn Mawr College) | Eaton, Eric (Bryn Mawr College)
The recently developed Efficient Lifelong Learning Algorithm (ELLA) acquires knowledge incrementally over a sequence of tasks, learning a repository of latent model components that are sparsely shared between models. ELLA shows strong performance in comparison to other multi-task learning algorithms, achieving nearly identical performance to batch multi-task learning methods while learning tasks sequentially in three orders of magnitude (over 1,000x) less time. In this paper, we evaluate several curriculum selection methods that allow ELLA to actively select the next task for learning in order to maximize performance on future learning tasks. Through experiments with three real and one synthetic data set, we demonstrate that active curriculum selection allows an agent to learn up to 50% more efficiently than when the agent has no control over the task order.
Symbolic Play and Analogy: a Way to Foster Childrenโs Creativity
Sefer, Jasmina (Institute for Educational Research, Belgrade)
The author discusses the relationship between symbolic play, abstract thinking, and divergent and associative thinking based on analogies, and finally connects symbolic play with the creative process. Play and the creative act are seen as similar by definition, since they are characterized as divergent, regulative, expressive and autotelic processes. Symbolic play is not only a product of the animistic and concrete logical way of thinking in childhood but also represents a mode of abstract thinking at the fictional symbolic level, which provides different options important for creativity development. Symbolic play is based on analogies with reality, and in this way reality is transformed in the imagination to be comprehended by the child. This transformation, which takes place in the nest of analogy at the symbolic level, is a key for creative production. Analogies in symbolic play are created through the divergent associative thinking process, also basic for any creative activity. The author has already used play as a tool to enhance creative behavior among young students in primary schools, and currently one project is being implemented in Serbia by the Institute for Educational Research with the intention of promoting initiative, cooperation and creativity by using play among other learning methods.
Towards Pareto Descent Directions in Sampling Experts for Multiple Tasks in an On-Line Learning Paradigm
Ghosh, Shaona (University of Southampton,UK) | Lovell, Chris (University of Southampton) | Gunn, Steve R. (University of Southampton)
In many real-life design problems, there is a requirement to simultaneously balance multiple tasks or objectives in the system that are conflicting in nature, where minimizing one objective causes another to increase in value, thereby resulting in trade-offs between the objectives. For example, in embedded multi-core mobile devices and very large scale data centers, there is a continuous problem of simultaneously balancing interfering goals of maximal power savings and minimal performance delay with varying trade-off values for different application workloads executing on them. Typically, the optimal trade-offs for the executing workloads, lie on a difficult to determine optimal Pareto front. The nature of the problem requires learning over the lifetime of the mobile device or server with continuous evaluation and prediction of the trade-off settings on the system that balances the interfering objectives optimally. Towards this, we propose an on-line learning method, where the weights of experts for addressing the objectives are updated based on a convex combination of their relative performance in addressing all objectives simultaneously. An additional importance vector that assigns relative importance to each objective at every round is used, and is sampled from a convex cone pointed at the origin Our preliminary results show that the convex combination of the importance vector and the gradient of the potential functions of the learner's regret with respect to each objective ensure that in the next round, the drift (instantaneous regret vector), is the Pareto descent direction that enables better convergence to the optimal Pareto front.
Information-Theoretic Objective Functions for Lifelong Learning
Zhang, Byoung-Tak (Seoul National University)
Conventional paradigms of machine learning assume all the training data are available when learning starts. However, in lifelong learning, the examples are observed sequentially as learning unfolds, and the learner should continually explore the world and reorganize and refine the internal model or knowledge of the world. This leads to a fundamental challenge: How to balance long-term and short-term goals and how to trade-off between information gain and model complexity? These questions boil down to โwhat objective functions can best guide a lifelong learning agent?โ Here we develop a sequential Bayesian framework for lifelong learning, build a taxonomy of lifelong-learning paradigms, and examine information-theoretic objective functions for each paradigm, with an emphasis on predictive and active learning. The objective functions can provide theoretical criteria for designing algorithms and determining effective strategies for selective sampling, representation discovery, knowledge transfer, and continual update over a lifetime of experience.
Lifelong Learning of Structure in the Space of Policies
Hawasly, Majd (University of Edinburgh) | Ramamoorthy, Subramanian (University of Edinburgh)
We address the problem faced by an autonomous agent that must achieve quick responses to a family of qualitatively-related tasks, such as a robot interacting with different types of human participants. We work in the setting where the tasks share a state-action space and have the same qualitative objective but differ in the dynamics and reward process. We adopt a transfer approach where the agent attempts to exploit common structure in learnt policies to accelerate learning in a new one. Our technique consists of a few key steps. First, we use a probabilistic model to describe the regions in state space which successful trajectories seem to prefer. Then, we extract policy fragments from previously-learnt policies for these regions as candidates for reuse. These fragments may be treated as options with corresponding domains and termination conditions extracted by unsupervised learning. Then, the set of reusable policies is used when learning novel tasks, and the process repeats. The utility of this method is demonstrated through experiments in the simulated soccer domain, where the variability comes from the different possible behaviours of opponent teams, and the agent needs to perform well against novel opponents.
A proximal Newton framework for composite minimization: Graph learning without Cholesky decompositions and matrix inversions
Dinh, Quoc Tran, Kyrillidis, Anastasios, Cevher, Volkan
We propose an algorithmic framework for convex minimization problems of a composite function with two terms: a self-concordant function and a possibly nonsmooth regularization term. Our method is a new proximal Newton algorithm that features a local quadratic convergence rate. As a specific instance of our framework, we consider the sparse inverse covariance matrix estimation in graph learning problems. Via a careful dual formulation and a novel analytic step-size selection procedure, our approach for graph learning avoids Cholesky decompositions and matrix inversions in its iteration making it attractive for parallel and distributed implementations.
Variational Inference in Nonconjugate Models
Mean-field variational methods are widely used for approximate posterior inference in many probabilistic models. In a typical application, mean-field methods approximately compute the posterior with a coordinate-ascent optimization algorithm. When the model is conditionally conjugate, the coordinate updates are easily derived and in closed form. However, many models of interest---like the correlated topic model and Bayesian logistic regression---are nonconjuate. In these models, mean-field methods cannot be directly applied and practitioners have had to develop variational algorithms on a case-by-case basis. In this paper, we develop two generic methods for nonconjugate models, Laplace variational inference and delta method variational inference. Our methods have several advantages: they allow for easily derived variational algorithms with a wide class of nonconjugate models; they extend and unify some of the existing algorithms that have been derived for specific models; and they work well on real-world datasets. We studied our methods on the correlated topic model, Bayesian logistic regression, and hierarchical Bayesian logistic regression.