Education
Programmable Reinforcement Learning Agents
Andre, David, Russell, Stuart J.
We present an expressive agent design language for reinforcement learning that allows the user to constrain the policies considered by the learning process.The language includes standard features such as parameterized subroutines, temporary interrupts, aborts, and memory variables, but also allows for unspecified choices in the agent program. For learning that which isn't specified, we present provably convergent learning algorithms. We demonstrate by example that agent programs written in the language are concise as well as modular. This facilitates state abstraction and the transferability of learned skills. 1 Introduction The field of reinforcement learning has recently adopted the idea that the application of prior knowledge may allow much faster learning and may indeed be essential if realworld environments are to be addressed. For learning behaviors, the most obvious form of prior knowledge provides a partial description of desired behaviors. Several languages for partial descriptions have been proposed, including Hierarchical Abstract Machines (HAMs) [8], semi-Markov options [12], and the MAXQ framework [4]. This paper describes extensions to the HAM language that substantially increase its expressive power, using constructs borrowed from programming languages. Obviously, increasing expressiveness makes it easier for the user to supply whatever prior knowledge is available, and to do so more concisely.
Programmable Reinforcement Learning Agents
Andre, David, Russell, Stuart J.
We present an expressive agent design language for reinforcement learning that allows the user to constrain the policies considered by the learning process.The language includes standard features such as parameterized subroutines, temporary interrupts, aborts, and memory variables, but also allows for unspecified choices in the agent program. For learning that which isn't specified, we present provably convergent learning algorithms. We demonstrate by example that agent programs written in the language are concise as well as modular. This facilitates state abstraction and the transferability of learned skills. 1 Introduction The field of reinforcement learning has recently adopted the idea that the application of prior knowledge may allow much faster learning and may indeed be essential if realworld environments are to be addressed. For learning behaviors, the most obvious form of prior knowledge provides a partial description of desired behaviors. Several languages for partial descriptions have been proposed, including Hierarchical Abstract Machines (HAMs) [8], semi-Markov options [12], and the MAXQ framework [4]. This paper describes extensions to the HAM language that substantially increase its expressive power, using constructs borrowed from programming languages. Obviously, increasing expressiveness makes it easier for the user to supply whatever prior knowledge is available, and to do so more concisely.
Model Complexity, Goodness of Fit and Diminishing Returns
Cadez, Igor V., Smyth, Padhraic
Igor V. Cadez Information and Computer Science University of California Irvine, CA 92697-3425, U.S.A. PadhraicSmyth Information and Computer Science University of California Irvine, CA 92697-3425, U.S.A. Abstract We investigate a general characteristic of the tradeoff in learning problems between goodness-of-fit and model complexity. Specifically wecharacterize a general class of learning problems where the goodness-of-fit function can be shown to be convex within firstorder asa function of model complexity. This general property of "diminishing returns" is illustrated on a number of real data sets and learning problems, including finite mixture modeling and multivariate linear regression. 1 Introduction, Motivation, and Related Work Assume we have a data set D Such learning tasks can typically be characterized by the existence of a model and a loss function. A fitted model of complexity k is a function of the data points D and depends on a specific set of fitted parameters B. The loss function (goodnessof-fit) isa functional of the model and maps each specific model to a scalar used to evaluate the model, e.g., likelihood for density estimation or sum-of-squares for regression. Figure 1 illustrates a typical empirical curve for loss function versus complexity, for mixtures of Markov models fitted to a large data set of 900,000 sequences.
Programmable Reinforcement Learning Agents
Andre, David, Russell, Stuart J.
We present an expressive agent design language for reinforcement learning thatallows the user to constrain the policies considered by the learning process.Thelanguage includes standard features such as parameterized subroutines,temporary interrupts, aborts, and memory variables, but also allows for unspecified choices in the agent program. For learning that which isn't specified, we present provably convergent learning algorithms. Wedemonstrate by example that agent programs written in the language are concise as well as modular. This facilitates state abstraction and the transferability of learned skills. 1 Introduction The field of reinforcement learning has recently adopted the idea that the application of prior knowledge may allow much faster learning and may indeed be essential if realworld environmentsare to be addressed. For learning behaviors, the most obvious form of prior knowledge provides a partial description of desired behaviors. Several languages for partial descriptions have been proposed, including Hierarchical Abstract Machines (HAMs) [8], semi-Markov options [12], and the MAXQ framework [4]. This paper describes extensions to the HAM language that substantially increase its expressive power,using constructs borrowed from programming languages. Obviously, increasing expressivenessmakes it easier for the user to supply whatever prior knowledge is available, and to do so more concisely.
Robust Reinforcement Learning
KenjiDoya ATR International; CREST, JST 2-2 Hikaridai Seika-cho Soraku-gun Kyoto 619-0288 JAPAN doya@isd.atr.co.jp Abstract This paper proposes a new reinforcement learning (RL) paradigm that explicitly takes into account input disturbance as well as modeling errors.The use of environmental models in RL is quite popular for both off-line learning by simulations and for online action planning. However, the difference between the model and the real environment can lead to unpredictable, often unwanted results. Based on the theory of H oocontrol, we consider a differential game in which a'disturbing' agent (disturber) tries to make the worst possible disturbance while a'control' agent (actor) tries to make the best control input. The problem is formulated as finding a minmax solutionof a value function that takes into account the norm of the output deviation and the norm of the disturbance. We derive online learning algorithms for estimating the value function and for calculating the worst disturbance and the best control in reference tothe value function.
Intelligent Tutoring Systems with Conversational Dialogue
Graesser, Arthur C., VanLehn, Kurt, Rose, Carolyn P., Jordan, Pamela W., Harter, Derek
Many of the intelligent tutoring systems that have been developed during the last 20 years have proven to be quite successful, particularly in the domains of mathematics, science, and technology. We have been working on a new generation of intelligent tutoring systems that hold mixed-initiative conversational dialogues with the learner. The tutoring systems present challenging problems and questions to the learner, the learner types in answers in English, and there is a lengthy multiturn dialogue as complete solutions or answers evolve. This article presents the tutoring systems that we have been developing.
Introduction to the Special Issue on Intelligent User Interfaces
Recent years have witnessed significant progress in intelligent user interfaces. Emerging from the intersection of AI and human-computer interaction, research on intelligent user interfaces is experiencing a renaissance, both in the overall level of activity and in raw research achievements. Research on intelligent user interfaces exploits developments in a broad range of foundational AI work, ranging from knowledge representation and computational linguistics to planning and vision. Because intelligent user interfaces are designed to facilitate problem-solving activities where reasoning is shared between users and the machine, they are currently transitioning from the laboratory to applications in the workplace, home, and classroom.
Pedagogical Agent Research at CARTE
They express both thoughts and California (USC)/Information Sciences Institute emotions; emotional expression is important to (ISI) is to develop new technologies that portray characteristics of enthusiasm and empathy promote effective learning and increase learner that are important for human teachers. These technologies are intended They are knowledgeable about the subject matter to result in interactive learning materials that being learned, of pedagogical strategies, and support the learning process and that complement also have knowledge about how to find and and enhance existing technologies relevant obtain relevant knowledge from available to learning such as the World Wide Web. Our work draws significant inspiration from Figure 1 shows one of the guidebots that we human learning and teaching. We piece of equipment called a high-pressure air seek a better understanding of the characteristics compressor aboard United States Navy ships. As learners view instructional materials, guidebots can provide useful commentary on these materials.
Intelligent Tutoring Systems with Conversational Dialogue
Graesser, Arthur C., VanLehn, Kurt, Rose, Carolyn P., Jordan, Pamela W., Harter, Derek
Many of the intelligent tutoring systems that have been developed during the last 20 years have proven to be quite successful, particularly in the domains of mathematics, science, and technology. They produce significant learning gains beyond classroom environments. They are capable of engaging most students' attention and interest for hours. We have been working on a new generation of intelligent tutoring systems that hold mixed-initiative conversational dialogues with the learner. The tutoring systems present challenging problems and questions to the learner, the learner types in answers in English, and there is a lengthy multiturn dialogue as complete solutions or answers evolve. This article presents the tutoring systems that we have been developing. AutoTutor is a conversational agent, with a talking head, that helps college students learn about computer literacy. andes, atlas, and why2 help adults learn about physics. Instead of being mere information-delivery systems, our systems help students actively construct knowledge through conversations.