Country
Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
Eysenbach, Benjamin, Salakhutdinov, Ruslan, Levine, Sergey
The history of learning for control has been an exciting back and forth between two broad classes of algorithms: planning and reinforcement learning. Planning algorithms effectively reason over long horizons, but assume access to a local policy and distance metric over collision-free paths. Reinforcement learning excels at learning policies and the relative values of states, but fails to plan over long horizons. Despite the successes of each method in various domains, tasks that require reasoning over long horizons with limited feedback and high-dimensional observations remain exceedingly challenging for both planning and reinforcement learning algorithms. Frustratingly, these sorts of tasks are potentially the most useful, as they are simple to design (a human only need to provide an example goal state) and avoid reward shaping, which can bias the agent towards finding a sub-optimal solution. We introduce a general control algorithm that combines the strengths of planning and reinforcement learning to effectively solve these tasks. Our aim is to decompose the task of reaching a distant goal state into a sequence of easier tasks, each of which corresponds to reaching a subgoal. Planning algorithms can automatically find these waypoints, but only if provided with suitable abstractions of the environment -- namely, a graph consisting of nodes and edges. Our main insight is that this graph can be constructed via reinforcement learning, where a goal-conditioned value function provides edge weights, and nodes are taken to be previously seen observations in a replay buffer. Using graph search over our replay buffer, we can automatically generate this sequence of subgoals, even in image-based environments. Our algorithm, search on the replay buffer (SoRB), enables agents to solve sparse reward tasks over one hundred steps, and generalizes substantially better than standard RL algorithms.
When to use parametric models in reinforcement learning?
van Hasselt, Hado, Hessel, Matteo, Aslanides, John
We examine the question of when and how parametric models are most useful in reinforcement learning. In particular, we look at commonalities and differences between parametric models and experience replay. Replay-based learning algorithms share important traits with model-based approaches, including the ability to plan: to use more computation without additional data to improve predictions and behaviour. We discuss when to expect benefits from either approach, and interpret prior work in this context. We hypothesise that, under suitable conditions, replay-based algorithms should be competitive to or better than model-based algorithms if the model is used only to generate fictional transitions from observed states for an update rule that is otherwise model-free. We validated this hypothesis on Atari 2600 video games. The replay-based algorithm attained state-of-the-art data efficiency, improving over prior results with parametric models.
General Video Game Rule Generation
Khalifa, Ahmed, Green, Michael Cerny, Perez-Liebana, Diego, Togelius, Julian
We introduce the General Video Game Rule Generation problem, and the eponymous software framework which will be used in a new track of the General Video Game AI (GVGAI) competition. The problem is, given a game level as input, to generate the rules of a game that fits that level. This can be seen as the inverse of the General Video Game Level Generation problem. Conceptualizing these two problems as separate helps breaking the very hard problem of generating complete games into smaller, more manageable subproblems. The proposed framework builds on the GVGAI software and thus asks the rule generator for rules defined in the Video Game Description Language. We describe the API, and three different rule generators: a random, a constructive and a search-based generator. Early results indicate that the constructive generator generates playable and somewhat interesting game rules but has a limited expressive range, whereas the search-based generator generates remarkably diverse rulesets, but with an uneven quality.
Polynomial-time Updates of Epistemic States in a Fragment of Probabilistic Epistemic Argumentation (Technical Report)
Potyka, Nico, Polberg, Sylwia, Hunter, Anthony
Probabilistic epistemic argumentation allows for reasoning about argumentation problems in a way that is well founded by probability theory. Epistemic states are represented by probability functions over possible worlds and can be adjusted to new beliefs using update operators. While the use of probability functions puts this approach on a solid foundational basis, it also causes computational challenges as the amount of data to process depends exponentially on the number of arguments. This leads to bottlenecks in applications such as modelling opponent's beliefs for persuasion dialogues. We show how update operators over probability functions can be related to update operators over much more compact representations that allow polynomial-time updates. We discuss the cognitive and probabilistic-logical plausibility of this approach and demonstrate its applicability in computational persuasion.
Neural Variational Inference For Estimating Uncertainty in Knowledge Graph Embeddings
Cowen-Rivers, Alexander I., Minervini, Pasquale, Rocktaschel, Tim, Bovsnjak, Matko, Riedel, Sebastian, Wang, Jun
Recent advances in Neural Variational Inference allowed for a renaissance in latent variable models in a variety of domains involving high-dimensional data. While traditional variational methods derive an analytical approximation for the intractable distribution over the latent variables, here we construct an inference network conditioned on the symbolic representation of entities and relation types in the Knowledge Graph, to provide the variational distributions. The new framework results in a highly-scalable method. Under a Bernoulli sampling framework, we provide an alternative justification for commonly used techniques in large-scale stochastic variational inference, which drastically reduce training time at a cost of an additional approximation to the variational lower bound. We introduce two models from this highly scalable probabilistic framework, namely the Latent Information and Latent Fact models, for reasoning over knowledge graph-based representations. Our Latent Information and Latent Fact models improve upon baseline performance under certain conditions. We use the learnt embedding variance to estimate predictive uncertainty during link prediction, and discuss the quality of these learnt uncertainty estimates. Our source code and datasets are publicly available online at https://github.com/alexanderimanicowenrivers/Neural-Variational-Knowledge-Graphs.
Unsupervised Question Answering by Cloze Translation
Lewis, Patrick, Denoyer, Ludovic, Riedel, Sebastian
Obtaining training data for Question Answering (QA) is time-consuming and resource-intensive, and existing QA datasets are only available for limited domains and languages. In this work, we explore to what extent high quality training data is actually required for Extractive QA, and investigate the possibility of unsupervised Extractive QA. We approach this problem by first learning to generate context, question and answer triples in an unsupervised manner, which we then use to synthesize Extractive QA training data automatically. To generate such triples, we first sample random context paragraphs from a large corpus of documents and then random noun phrases or named entity mentions from these paragraphs as answers. Next we convert answers in context to "fill-in-the-blank" cloze questions and finally translate them into natural questions. We propose and compare various unsupervised ways to perform cloze-to-natural question translation, including training an unsupervised NMT model using non-aligned corpora of natural questions and cloze questions as well as a rule-based approach. We find that modern QA models can learn to answer human questions surprisingly well using only synthetic training data. We demonstrate that, without using the SQuAD training data at all, our approach achieves 56.4 F1 on SQuAD v1 (64.5 F1 when the answer is a Named entity mention), outperforming early supervised models.
A Unified Linear-Time Framework for Sentence-Level Discourse Parsing
Lin, Xiang, Joty, Shafiq, Jwalapuram, Prathyusha, Bari, M Saiful
We propose an efficient neural framework for sentence-level discourse analysis in accordance with Rhetorical Structure Theory (RST). Our framework comprises a discourse segmenter to identify the elementary discourse units (EDU) in a text, and a discourse parser that constructs a discourse tree in a top-down fashion. Both the segmenter and the parser are based on Pointer Networks and operate in linear time. Our segmenter yields an $F_1$ score of 95.4, and our parser achieves an $F_1$ score of 81.7 on the aggregated labeled (relation) metric, surpassing previous approaches by a good margin and approaching human agreement on both tasks (98.3 and 83.0 $F_1$).
Scientists who scanned European al Qaeda supporters' brains found reduced activity
Brain scans on fanatical Islamists show they have a reduced capacity for rational thought, new research suggests. Members of a radical Islamist group were asked how willing they were to'fight and die' for their ideas. Their brains were then scanned during the process. The results showed that when questioned, the part of the brain that engages in evaluating costs and consequences showed reduced activity. The scientists say this shows that when it comes to values held'sacred' to the radicals, they are immune to arguments involving costs and benefits.
#288: On Artificial Intelligence for Wildlife Conservation, with Milind Tambe
Dr. Tambe describes his team's use of security games to combat poaching, and his experience deploying his algorithms to inform park ranger schedules internationally. Dr. Milind Tambe is the Helen N. and Emmett H. Jones Professor in Engineering at the University of Southern California, and Professor in the Computer Science and Industrial and Systems Engineering Departments. He is a founding co-director of the CAIS Center for AI in Society, where he advises students and conducts research on multiagent teamwork, distributed constraint optimization, and security games. The security games framework developed by Dr. Tambe has been deployed and tested nationally and internationally, and led to his co-founding of company Avata Intelligence.
E3 2019: Luigi is a star in Nintendo's 2019 video game lineup with 'Luigi's Mansion 3'
Mario's brother Luigi is the ghostbusting star in'Luigi's Mansion 3,' a new video game coming to Nintendo Switch later this year. LOS ANGELES – Mario is usually front and center for Nintendo, but the video game maker has propelled his brother Luigi to the forefront at this year's Electronic Entertainment Expo. The upcoming game "Luigi's Mansion 3," out later this year for Nintendo Switch, got headlining treatment during the Nintendo Direct online video message Tuesday, just before the doors opened at the E3 expo, which runs through Thursday. The setup: A scary adventure ensues after what Luigi and Mario thought was going to be a fun vacation getaway at an upscale resort, instead turns out to be a trip to haunted hotel. When Mario and Princess Peach go missing, Luigi must become a ghostbuster and find them.