Goto

Collaborating Authors

 Technology


Minimax-Optimal Bounds for Detectors Based on Estimated Prior Probabilities

arXiv.org Machine Learning

In many signal detection and classification problems, we have knowledge of the distribution under each hypothesis, but not the prior probabilities. This paper is aimed at providing theory to quantify the performance of detection via estimating prior probabilities from either labeled or unlabeled training data. The error or {\em risk} is considered as a function of the prior probabilities. We show that the risk function is locally Lipschitz in the vicinity of the true prior probabilities, and the error of detectors based on estimated prior probabilities depends on the behavior of the risk function in this locality. In general, we show that the error of detectors based on the Maximum Likelihood Estimate (MLE) of the prior probabilities converges to the Bayes error at a rate of $n^{-1/2}$, where $n$ is the number of training data. If the behavior of the risk function is more favorable, then detectors based on the MLE have errors converging to the corresponding Bayes errors at optimal rates of the form $n^{-(1+\alpha)/2}$, where $\alpha>0$ is a parameter governing the behavior of the risk function with a typical value $\alpha = 1$. The limit $\alpha \rightarrow \infty$ corresponds to a situation where the risk function is flat near the true probabilities, and thus insensitive to small errors in the MLE; in this case the error of the detector based on the MLE converges to the Bayes error exponentially fast with $n$. We show the bounds are achievable no matter given labeled or unlabeled training data and are minimax-optimal in labeled case.


Counting-Based Search: Branching Heuristics for Constraint Satisfaction Problems

Journal of Artificial Intelligence Research

Designing a search heuristic for constraint programming that is reliable across problem domains has been an important research topic in recent years. This paper concentrates on one family of candidates: counting-based search. Such heuristics seek to make branching decisions that preserve most of the solutions by determining what proportion of solutions to each individual constraint agree with that decision. Whereas most generic search heuristics in constraint programming rely on local information at the level of the individual variable, our search heuristics are based on more global information at the constraint level. We design several algorithms that are used to count the number of solutions to specific families of constraints and propose some search heuristics exploiting such information. The experimental part of the paper considers eight problem domains ranging from well-established benchmark puzzles to rostering and sport scheduling. An initial empirical analysis identifies heuristic maxSD as a robust candidate among our proposals. We then evaluate the latter against the state of the art, including the latest generic search heuristics, restarts, and discrepancy-based tree traversals. Experimental results show that counting-based search generally outperforms other generic heuristics.


Empirical estimation of entropy functionals with confidence

arXiv.org Machine Learning

This paper introduces a class of k-nearest neighbor ($k$-NN) estimators called bipartite plug-in (BPI) estimators for estimating integrals of non-linear functions of a probability density, such as Shannon entropy and R\'enyi entropy. The density is assumed to be smooth, have bounded support, and be uniformly bounded from below on this set. Unlike previous $k$-NN estimators of non-linear density functionals, the proposed estimator uses data-splitting and boundary correction to achieve lower mean square error. Specifically, we assume that $T$ i.i.d. samples ${X}_i \in \mathbb{R}^d$ from the density are split into two pieces of cardinality $M$ and $N$ respectively, with $M$ samples used for computing a k-nearest-neighbor density estimate and the remaining $N$ samples used for empirical estimation of the integral of the density functional. By studying the statistical properties of k-NN balls, explicit rates for the bias and variance of the BPI estimator are derived in terms of the sample size, the dimension of the samples and the underlying probability distribution. Based on these results, it is possible to specify optimal choice of tuning parameters $M/T$, $k$ for maximizing the rate of decrease of the mean square error (MSE). The resultant optimized BPI estimator converges faster and achieves lower mean squared error than previous $k$-NN entropy estimators. In addition, a central limit theorem is established for the BPI estimator that allows us to specify tight asymptotic confidence intervals.


Active Learning of Halfspaces under a Margin Assumption

arXiv.org Machine Learning

We derive and analyze a new, efficient, pool-based active learning algorithm for halfspaces, called ALuMA. Most previous algorithms show exponential improvement in the label complexity assuming that the distribution over the instance space is close to uniform. This assumption rarely holds in practical applications. Instead, we study the label complexity under a large-margin assumption -- a much more realistic condition, as evident by the success of margin-based algorithms such as SVM. Our algorithm is computationally efficient and comes with formal guarantees on its label complexity. It also naturally extends to the non-separable case and to non-linear kernels. Experiments illustrate the clear advantage of ALuMA over other active learning algorithms.


Interaction Histories and Short Term Memory: Enactive Development of Turn-taking Behaviors in a Childlike Humanoid Robot

arXiv.org Artificial Intelligence

In this article, an enactive architecture is described that allows a humanoid robot to learn to compose simple actions into turn-taking behaviors while playing interaction games with a human partner. The robot's action choices are reinforced by social feedback from the human in the form of visual attention and measures of behavioral synchronization. We demonstrate that the system can acquire and switch between behaviors learned through interaction based on social feedback from the human partner. The role of reinforcement based on a short term memory of the interaction is experimentally investigated. Results indicate that feedback based only on the immediate state is insufficient to learn certain turn-taking behaviors. Therefore some history of the interaction must be considered in the acquisition of turn-taking, which can be efficiently handled through the use of short term memory.


Type-elimination-based reasoning for the description logic SHIQbs using decision diagrams and disjunctive datalog

arXiv.org Artificial Intelligence

We propose a novel, type-elimination-based method for reasoning in the description logic SHIQbs including DL-safe rules. To this end, we first establish a knowledge compilation method converting the terminological part of an ALCIb knowledge base into an ordered binary decision diagram (OBDD) which represents a canonical model. This OBDD can in turn be transformed into disjunctive Datalog and merged with the assertional part of the knowledge base in order to perform combined reasoning. In order to leverage our technique for full SHIQbs, we provide a stepwise reduction from SHIQbs to ALCIb that preserves satisfiability and entailment of positive and negative ground facts. The proposed technique is shown to be worst case optimal w.r.t. combined and data complexity and easily admits extensions with ground conjunctive queries.


Cultural Analytics of Large Datasets from Flickr

AAAI Conferences

Deluge became a metaphor to describe the amount of information to which we are subjected, and very often we feel we are drowning while our access to information is rising. Devising mechanisms for exploring massive image sets according to perceptual attributes is still a challenge, even more when dealing with user-generated social media content. Such images tend to be heterogenous, and using metadata-only can be misleading. This paper describes a set of tools designed to analyze large sets of user-created art related images using image features describing color, texture, composition and orientation. The proposed pipeline permits to discriminate Flickr groups in terms of feature vectors and clustering parameters. The algorithms are general enough to be applied to other domains in which the main question is about the variability of the images.


Composing Traveling Paths from Location-Based Services

AAAI Conferences

With the emergence of location-based services, such as Foursquare and Gowalla, users are allowed to easily perform check-in actions anywhere and anytime. The location-based check-in not only enables personal geospatial journeys but also serves as a kind of fine-grained source for trip planning. In this work, we aim to collectively compose traveling paths by leveraging the check-in data through mining the moving behaviors of users. A novel system, TP-Comp, is developed. To compose travel paths, TP-Comp not only allows users to specify starting/end and/or must-go locations, but also provides the flexibility of the time constraint requirement (i.e., the expected duration of the trip). By considering a sequence of check-in points as a traveling path, we mine the frequent sequences with some ranking mechanism to achieve the goal. Our TP-Comp targets at travelers who are unfamiliar to the objective area/city and have time limitation in the trip.


Feasibility Study on Detection of Transportation Information Exploiting Twitter as a Sensor

AAAI Conferences

The concept of a smart community has recently been attracting great attention as a means of utilizing energy effectively. One of the modules constituting the smart community is an intelligent transportation system, in which various sensors track movements of people and vehicles in real time to optimize migration pathways or means. Social media have the potential to serve as sensors, since people often post transportation information on such media. This paper presents a feasibility study on detecting information, focusing on train status information, by exploiting Twitter as a sensor. We dealt with two issues: (1) for the ambiguity of textual information expressed in tweets, we utilized heuristic rules in text manipulation, and (2) for the differences in the numbers of tweets among train lines, we optimized parameter values in statistical analysis for each train line. The experimental results show that the F-measure of detecting the information was more than 0.85 and the time taken to detect the information was less than 4 minutes. As a result we confirmed the high potential of detecting transportation information through Twitter.


Tag Recommendation by Link Prediction Based on Supervised Machine Learning

AAAI Conferences

In this work, we explore applying a link prediction approach to tag recommendation in broad folksonomies. The original idea of the approach is to mine the dynamic of the tagging activity in order to compute the most suitable tag for a given user and a given resource. The tagging history of each user is modeled by a temporal sequence of bipartite graphs linking tags to resources. Given a target user and a target resource, we first compute a set of similar users. The tagging history of the identified set of users is merged in one temporal sequence on bipartite graphs. The obtained sequence is used to learn a model of link prediction in bipartite graphs. The learned model is then applied to predict tags to be linked to the target resource and a list of top similar resources. We get hence several ranked lists tags, one list for each considered resource. These ranked lists are then merged, applying classical preference merging methods in order to obtain the final output: a list of ranked tags that will be recommended to the user. We show through experiments conducted on real datasets extracted for the CiteULike folksonomy the soundness of the proposed approach.