Education
Algorithmic Improvements for Deep Reinforcement Learning applied to Interactive Fiction
Jain, Vishal, Fedus, William, Larochelle, Hugo, Precup, Doina, Bellemare, Marc G.
Text-based games are a natural challenge domain for deep reinforcement learning algorithms. Their state and action spaces are combinatorially large, their reward function is sparse, and they are partially observable: the agent is informed of the consequences of its actions through textual feedback. In this paper we emphasize this latter point and consider the design of a deep reinforcement learning agent that can play from feedback alone. Our design recognizes and takes advantage of the structural characteristics of text-based games. We first propose a contextualisation mechanism, based on accumulated reward, which simplifies the learning problem and mitigates partial observability. We then study different methods that rely on the notion that most actions are ineffectual in any given situation, following Zahavy et al.'s idea of an admissible action. We evaluate these techniques in a series of text-based games of increasing difficulty based on the TextWorld framework, as well as the iconic game Zork. Empirically, we find that these techniques improve the performance of a baseline deep reinforcement learning agent applied to text-based games.
Stigmergic Independent Reinforcement Learning for Multi-Agent Collaboration
Xing, Xu, Rongpeng, Li, Zhifeng, Zhao, Honggang, Zhang
--With the rapid evolution of wireless mobile devices, it emerges stronger incentive to design proper collaboration mechanisms among the intelligent agents. Following their individual observations, multiple intelligent agents could cooperate and gradually approach the final collective objective through continuously learning from the environment. In that regard, independent reinforcement learning (IRL) is often deployed within the multi-agent collaboration to alleviate the dilemma of non-stationary learning environment. However, behavioral strategies of the intelligent agents in IRL could only be formulated upon their local individual observations of the global environment, and appropriate communication mechanisms must be introduced to reduce their behavioral localities. In this paper, we tackle the communication problem among the intelligent agents in IRL by jointly adopting two mechanisms with different scales. For the large scale, we introduce the stigmergy mechanism as an indirect communication bridge among the independent learning agents and carefully design a mathematical representation to indicate the impact of digital pheromone. For the small scale, we propose a conflict-avoidance mechanism between adjacent agents by implementing an additionally embedded neural network to provide more opportunities for participants with higher action priorities. Besides, we also present a federal training method to effectively optimize the neural networks within each agent in a decentralized manner . Finally, we establish a simulation scenario where a number of mobile agents in a certain area move automatically to form a specified target shape, and demonstrate the superiorities of our proposed methods through extensive simulations. I NTRODUCTION With the rapid development of mobile wireless communication and IoTs (Internet of Things) technologies, many scenarios gradually arise where the collaboration among the involved intelligent agents is highly required, such as the deployment of unmanned aerial vehicles (UA Vs) [1]-[3], the distributed control in the field of industry automation [4]-[6], and mobile crowd sensing and computing (MCSC) [7], [8]. In these scenarios, traditional centralized control methods are usually impracticable because of the restriction from limited computing resources as well as the demand for ultra-low latency and ultra-high reliability. As an alternative, multi-agent collaboration can be introduced into these scenarios to reduce the pressure at the central controller side. As one of the primary goals in the field of artificial intelligence (AI), assisting autonomous agents to act optimally through the "trial-and-error" interaction process with the expected environment is regarded as an important target of reinforcement learning (RL) [9]-[11].
Information-Geometric Set Embeddings (IGSE): From Sets to Probability Distributions
This letter introduces an abstract learning problem called the ``set embedding'': The objective is to map sets into probability distributions so as to lose less information. We relate set union and intersection operations with corresponding interpolations of probability distributions. We also demonstrate a preliminary solution with experimental results on toy set embedding examples.
Improving Model Robustness Using Causal Knowledge
Kyono, Trent, van der Schaar, Mihaela
For decades, researchers in fields, such as the natural and social sciences, have been verifying causal relationships and investigating hypotheses that are now well-established or understood as truth. These causal mechanisms are properties of the natural world, and thus are invariant conditions regardless of the collection domain or environment. We show in this paper how prior knowledge in the form of a causal graph can be utilized to guide model selection, i.e., to identify from a set of trained networks the models that are the most robust and invariant to unseen domains. Our method incorporates prior knowledge (which can be incomplete) as a Structural Causal Model (SCM) and calculates a score based on the likelihood of the SCM given the target predictions of a candidate model and the provided input variables. We show on both publicly available and synthetic datasets that our method is able to identify more robust models in terms of generalizability to unseen out-of-distribution test examples and domains where covariates have shifted.
GRIm-RePR: Prioritising Generating Important Features for Pseudo-Rehearsal
Atkinson, Craig, McCane, Brendan, Szymanski, Lech, Robins, Anthony
Pseudo-rehearsal allows neural networks to learn a sequence of tasks without forgetting how to perform in earlier tasks. Preventing forgetting is achieved by introducing a generative network which can produce data from previously seen tasks so that it can be rehearsed along side learning the new task. This has been found to be effective in both supervised and reinforcement learning. Our current work aims to further prevent forgetting by encouraging the generator to accurately generate features important for task retention. More specifically, the generator is improved by introducing a second discriminator into the Generative Adversarial Network which learns to classify between real and fake items from the intermediate activation patterns that they produce when fed through a continual learning agent. Using Atari 2600 games, we experimentally find that improving the generator can considerably reduce catastrophic forgetting compared to the standard pseudo-rehearsal methods used in deep reinforcement learning. Furthermore, we propose normalising the Q-values taught to the long-term system as we observe this substantially reduces catastrophic forgetting by minimising the interference between tasks' reward functions.
Contrastive Learning of Structured World Models
Kipf, Thomas, van der Pol, Elise, Welling, Max
A structured understanding of our world in terms of objects, relations, and hierarchies is an important component of human cognition. Learning such a structured world model from raw sensory data remains a challenge. As a step towards this goal, we introduce Contrastively-trained Structured World Models (C-SWMs). C-SWMs utilize a contrastive approach for representation learning in environments with compositional structure. We structure each state embedding as a set of object representations and their relations, modeled by a graph neural network. This allows objects to be discovered from raw pixel observations without direct supervision as part of the learning process. We evaluate C-SWMs on compositional environments involving multiple interacting objects that can be manipulated independently by an agent, simple Atari games, and a multi-object physics simulation. Our experiments demonstrate that C-SWMs can overcome limitations of models based on pixel reconstruction and outperform typical representatives of this model class in highly structured environments, while learning interpretable object-based representations.
AI Curricula for K-12 Classrooms
Schools like those in the Pennsylvania Montour School District have mandated AI in the grades 5-8 curriculum, and they are expanded the initiative in other grades as well. Educators have embedded artificial intelligence in STEM courses, and other subjects like Music, Computer Science and Media Arts also include AI in their curricula. Additionally, the district requires their students to take a stand-alone AI Ethics course that teaches students design and values.
Build 111 Projects, Earn 10 Certifications - Now With Python
We've been working hard on Version 7.0 of the freeCodeCamp curriculum. Some of these improvements - including 4 new Python certifications - will go live in early 2019. Note: if you're already going through the current version of the curriculum, keep going. As you'll see, there's no reason to stop. Will take a person with very basic computer knowledge...
Utility Companies Prepare for AI-Powered Cyber Threats
The automated nature of such attacks means that they can be launched at speeds far in excess of what humans are capable of, he said, suggesting that attacks could happen on a microsecond-by-microsecond level. "We're going to have to understand the implications of, not people-to-machine attacks, but machine-to-machine attacks," said Mr. Fanning. Some security teams are using AI defensively, but cybersecurity leaders across sectors worry that the same technology could propel sophisticated attacks that will be difficult to fend off. A congressional report published last year raised the possibility of AI-based attacks overwhelming grid defenses. Utilities need to invest in defenses and do so quickly, said Mark James, an adjunct professor of law at Vermont Law School and a co-author of a report on state power utilities' cybersecurity practices, published this month.
Bots in the Library? Colleges Try AI to Help Researchers (But With Caution) - EdSurge News
The newest librarian at the University of Oklahoma is a robot. It's a chatbot, which library officials plan to add to the library's website this summer to answer some of the most common questions students come in with, as well as to help them get started with their research. The system can tackle things like "where can I print?" or "what databases do you have about biology?" Anything it can't answer gets sent to a human librarian. The bot is just one example of how college libraries and technologists are experimenting with artificial intelligence to support students and professors in their research. Algorithms may soon help them prepare their literature reviews by quickly finding the most important papers in an area, and help match researchers with peers in other disciplines doing similar work to form new collaborations.