Goto

Collaborating Authors

 Education


Streamed Learning: One-Pass SVMs

AAAI Conferences

We present a streaming model for large-scale classification (in the context of โ„“2 -SVM) by leveraging connections between learning and computational geometry. The streaming model imposes the constraint that only a single pass over the data is allowed. The โ„“2 -SVM is known to have an equivalent formulation in terms of the minimum enclosing ball (MEB) problem, and an efficient algorithm based on the idea of core sets exists (CVM) [Tsang et al., 2005]. CVM learns a (1 + ฮต)-approximate MEB for a set of points and yields an approximate solution to corresponding SVM instance. However CVM works in batch mode requiring multiple passes over the data. This paper presents a single-pass SVM which is based on the minimum enclosing ball of streaming data. We show that the MEB updates for the streaming case can be easily adapted to learn the SVM weight vector in a way similar to using online stochastic gradient updates. Our algorithm performs polylogarithmic computation at each example, and requires very small and constant storage. Experimental results show that, even in such restrictive settings, we can learn efficiently in just one pass and get accuracies comparable to other state-of-the-art SVM solvers (batch and online). We also give an analysis of the algorithm, and discuss some open issues and possible extensions.


Local Learning Regularized Nonnegative Matrix Factorization

AAAI Conferences

Nonnegative Matrix Factorization (NMF) has been widely used in machine learning and data mining. It aims to find two nonnegative matrices whose product can well approximate the nonnegative data matrix, which naturally lead to parts-based representation. In this paper, we present a local learning regularized nonnegative matrix factorization (LLNMF) for clustering. It imposes an additional constraint on NMF that the cluster label of each point can be predicted by the points in its neighborhood. This constraint encodes both the discriminative information and the geometric structure, and is good at clustering data on manifold. An iterative multiplicative updating algorithm is proposed to optimize the objective, and its convergence is guaranteed theoretically. Experiments on many benchmark data sets demonstrate that the proposed method outperforms NMF as well as many state of the art clustering methods.


Efficient Online Learning and Prediction of Users' Desktop Actions

AAAI Conferences

We investigate prediction of users' desktop activities in the Unix domain. The learning techniques we explore do not require explicit user teaching. We show that simple efficient many-class learning can perform well for action prediction, significantly improving over previously published results and baselines. This finding is promising for various human-computer interaction scenarios where a rich set of potentially predictive features is available, where there can be many different actions to predict, and where there can be considerable nonstationarity.


Efficient Skill Learning Using Abstraction Selection

AAAI Conferences

We present an algorithm for selecting an appropriate abstraction when learning a new skill. We show empirically that it can consistently select an appropriate abstraction using very little sample data, and that it significantly improves skill learning performance in a reasonably large real-valued reinforcement learning domain.


Machine Learning in Ecosystem Informatics and Sustainability

AAAI Conferences

Ecosystem Informatics brings together mathematical and computational tools to address scientific and policy challenges in the ecosystem sciences. These challenges include novel sensors for collecting data, algorithms for automated data cleaning, learning methods for building statistical models from data and for fitting mechanistic models to data, and algorithms for designing optimal policies for biosphere management. This presentation discusses these challenges and then describes recent work on the first two of these--new methods for automated arthropod population counting and linear Gaussian DBNs for automated cleaning of sensor network data.


Toward a Category Theory Design of Ontological Knowledge Bases

arXiv.org Artificial Intelligence

I discuss (ontologies_and_ontological_knowledge_bases / formal_methods_and_theories) duality and its category theory extensions as a step toward a solution to Knowledge-Based Systems Theory. In particular I focus on the example of the design of elements of ontologies and ontological knowledge bases of next three electronic courses: Foundations of Research Activities, Virtual Modeling of Complex Systems and Introduction to String Theory.


Granularity-Adaptive Proof Presentation

arXiv.org Artificial Intelligence

When mathematicians present proofs they usually adapt their explanations to their didactic goals and to the (assumed) knowledge of their addressees. Modern automated theorem provers, in contrast, present proofs usually at a fixed level of detail (also called granularity). Often these presentations are neither intended nor suitable for human use. A challenge therefore is to develop user- and goal-adaptive proof presentation techniques that obey common mathematical practice. We present a flexible and adaptive approach to proof presentation that exploits machine learning techniques to extract a model of the specific granularity of proof examples and employs this model for the automated generation of further proofs at an adapted level of granularity.


Knowledge Engineering with Didactic Knowledge โ€” First Steps towards an Ultimate Goal

AAAI Conferences

Generally, learning systems suffer from a lack of an explicit and adaptable didactic design. A previously introduced modeling approach called storyboarding is setting the stage to apply Knowledge Engineering Technologies to verify and validate the didactics behind a learning process. Moreover, didactics can be refined according to revealed weaknesses and proven excellence. Successful didactic patterns can be explored by applying mining techniques to the various ways students went through the storyboard and their associated level of success.


Invited Talks

AAAI Conferences

Vincent Aleven Intelligent tutoring systems (ITS) are highly effective in supporting student learning, but are difficult to build. The Cognitive Tutor Authoring Tools (CTAT) project started over 6 years ago with the goals of making it easier for experienced programmers, and possible for non-programmers to create an ITS. CTAT supports tutor building through programming by demonstration, an approach that has been successful in a range of application areas, but that has been applied to only a very limited degree to ITS authoring. Using CTAT, an author creates a tutor by demonstrating correct and incorrect problem solving behaviors, rather than by writing code. The resulting tutors, called exampletracing tutors, evaluate student behavior by flexibly comparing it against the demonstrated problem-solving examples.


Promoting Reflection and its Effect on Learning in a Programming Tutor

AAAI Conferences

We studied the effect of post-practice reflection on learning, using programming tutors, and multiple-choice format for reflection. We conducted in-vivo controlled studies with introductory programming students from multiple schools over 3 semesters, and used mixed-factor ANOVA to analyze the collected data. We found that reflecting on the concept underlying each problem neither promotes greater learning, measured as pre-post increase in the average score per problem, nor promotes faster learning, measured as the problems solved per concept learned. We conjecture that the benefits of reflecting on the concept underlying each problem may be limited if a tutor already promotes deep understanding of the domain.