Goto

Collaborating Authors

 Education


Reinforcement Learning in Partially Observable Markov Decision Processes using Hybrid Probabilistic Logic Programs

arXiv.org Artificial Intelligence

We present a probabilistic logic programming framework to reinforcement learning, by integrating reinforce-ment learning, in POMDP environments, with normal hybrid probabilistic logic programs with probabilistic answer set seman-tics, that is capable of representing domain-specific knowledge. We formally prove the correctness of our approach. We show that the complexity of finding a policy for a reinforcement learning problem in our approach is NP-complete. In addition, we show that any reinforcement learning problem can be encoded as a classical logic program with answer set semantics. We also show that a reinforcement learning problem can be encoded as a SAT problem. We present a new high level action description language that allows the factored representation of POMDP. Moreover, we modify the original model of POMDP so that it be able to distinguish between knowledge producing actions and actions that change the environment.


A Large-Deviation Analysis of the Maximum-Likelihood Learning of Markov Tree Structures

arXiv.org Machine Learning

The problem of maximum-likelihood (ML) estimation of discrete tree-structured distributions is considered. Chow and Liu established that ML-estimation reduces to the construction of a maximum-weight spanning tree using the empirical mutual information quantities as the edge weights. Using the theory of large-deviations, we analyze the exponent associated with the error probability of the event that the ML-estimate of the Markov tree structure differs from the true tree structure, given a set of independently drawn samples. By exploiting the fact that the output of ML-estimation is a tree, we establish that the error exponent is equal to the exponential rate of decay of a single dominant crossover event. We prove that in this dominant crossover event, a non-neighbor node pair replaces a true edge of the distribution that is along the path of edges in the true tree graph connecting the nodes in the non-neighbor pair. Using ideas from Euclidean information theory, we then analyze the scenario of ML-estimation in the very noisy learning regime and show that the error exponent can be approximated as a ratio, which is interpreted as the signal-to-noise ratio (SNR) for learning tree distributions. We show via numerical experiments that in this regime, our SNR approximation is accurate.


An Introduction to Conditional Random Fields

arXiv.org Machine Learning

Often we wish to predict a large number of variables that depend on each other as well as on other observed variables. Structured prediction methods are essentially a combination of classification and graphical modeling, combining the ability of graphical models to compactly model multivariate data with the ability of classification methods to perform prediction using large sets of input features. This tutorial describes conditional random fields, a popular probabilistic method for structured prediction. CRFs have seen wide application in natural language processing, computer vision, and bioinformatics. We describe methods for inference and parameter estimation for CRFs, including practical issues for implementing large scale CRFs. We do not assume previous knowledge of graphical modeling, so this tutorial is intended to be useful to practitioners in a wide variety of fields.


Towards a Computational Model of Why Some Students Learn Faster than Others

AAAI Conferences

Learners that have better metacognition acquire knowledge faster than others who do not. If we had better models of such learning, we would be able to build a better metacognitive educational system. In this paper, we propose a computational model that uses a probabilistic context free grammar induction algorithm yielding metacognitive learning by acquiring deep features to assist future learning. We discuss the challenges of integrating this model into a synthetic student, and possible future studies in using this model to better understand human learning. Preliminary results suggest that both stronger prior knowledge and a better learning strategy can speed up the learning process. Some model variations generate human-like error pattern.


Preface: Meta-Cognitive Educational Systems: One Step Forward

AAAI Conferences

The AAAI Fall Symposium on Meta-Cognitive Educational - What are the theoretical foundations and how are they articulated Systems: One Step Forward is the second edition of the successful in CBLEs? MCES implemented as CBLEs are designed to interact with - What are the main aspects of metacognition, selfregulation users, and support their learning and decision-making processes. Can MCES actually foster they need to plan their learning activities, to adapt their learners to be self-regulating agents? How can a MCES learning strategies to meet learning goals, become aware of be autonomous and increase its knowledge to match the changing task conditions, and the dynamic aspects of the learners evolving skills and knowledge? MCES may not be embodied, prior to, during, and after they have been involved in but does it help if they act as intentional agents? the learning environment.


Acquiring Vocabulary through Human Robot Interaction: A Learning Architecture for Grounding Words with Multiple Meanings

AAAI Conferences

This paper presents a robust methodology for grounding vocabulary in robots. A social language grounding experiment is designed, where, a human instructor teaches a robotic agent the names of the objects present in a visually shared environment. Any system for grounding vocabulary has to incorporate the properties of gradual evolution and lifelong learning. The learning model of the robot is adopted from an ongoing work on developing systems that conform to these properties. Significant modifications have been introduced to the adopted model, especially to handle words with multiple meanings. A novel classification strategy has been developed for improving the performance of each classifier for each learned category. A set of six new nearest-neighbor based classifiers have also been integrated into the agent architecture. A series of experiments were conducted to test the performance of the new model on vocabulary acquisition. The robot was shown to be robust at acquiring vocabulary and has the potential to learn a far greater number of words (with either single or multiple meanings).


Preparing to Talk: Interaction between a Linguistically Enabled Agent and a Human Teacher

AAAI Conferences

As a precursor to learning to use language an infant has to acquire preliminary linguistic skills, including the ability to recognize and produce word forms without meaning. This develops out of babbling, through vocal interaction with carers. We report on evidence from developmental psychology and from neuroscientific research that supports a dual process approach to language learning. We describe a simulation of the transition from babbling to the recognition of first word forms in a simulated robot interacting with a human teacher. This precedes interactions with the real iCub robot.


Building a Job Lanscape from Directional Transition Data

AAAI Conferences

The analysis of career paths suffers from a lack of exploratory tools and dynamic models, due in part to the inherent high dimensionality of the problem. Paths may be understood as directed traversals through a graph whose nodes consist of "job types," which we define as industry and occupation pairs. We want to develop tools to understand and detect high-level features of  both the labor market and the workers moving through it — career dynamics. To do this, we map the discrete space of jobs into a d-dimensional continuous space; proximity between jobs will mean that they are "close" to each other in a non-negligible subset of career paths. This embedding allows one to visualize the job landscape.  Moreover, we can map individual or groups of career paths to this space, extract features of their collective structure, and construct statistical tests comparing groups by means of this mapping.


Toward Fast Mapping for Robot Adjective Learning

AAAI Conferences

Fast mapping is a phenomenon by which children learn the meanings of novel adjectives after a very small number of exposures when the new word is contrasted with a known word. The present study was a preliminary test of whether machine learners could use such contrasts in unconstrained speech to learn adjective meanings and categories. Six decision tree-based learning methods were evaluated that use contrasting examples in order to work toward an adjective fast-mapping system for machine learners. Subjects tended to compare objects using adjectives of the same category, implying that such contrasts may be a useful source of data about adjective meaning, though none of the learning algorithms showed strong advantages over any other.


The Role of Prompting and Feedback in Facilitating Students’ Learning about Science with MetaTutor

AAAI Conferences

An experiment was conducted to test the efficacy of a new intelligent hypermedia system, MetaTutor, which is intended to prompt and scaffold the use of self-regulated learning (SRL) processes during learning about a human body system. Sixty-eight (N=68) undergraduate students learned about the human circulatory system under one of three conditions: prompt and feedback (PF), prompt-only (PO), and control (C) condition. The PF condition received timely prompts from animated pedagogical agents to engage in planning processes, monitoring processes, and learning strategies and also received immediate directive feedback from the agents concerning the deployment of the processes. The PO condition received the same timely prompts, but did not receive any feedback following the deployment of the processes. Finally, the control condition learned without any assistance from the agents during the learning session. All participants had two hours to learn using a 41-page hypermedia environment which included texts describing and static diagrams depicting various topics concerning the human circulatory system. Results indicate that the PF condition had significantly higher learning efficiency scores, when compared to the control condition. There were no significant differences between the PF and PO conditions. These results are discussed in the context of development of a fully-adaptive hypermedia learning system intended to scaffold self-regulated learning.