Goto

Collaborating Authors

 Country


Variance function estimation in high-dimensions

arXiv.org Machine Learning

We consider the high-dimensional heteroscedastic regression model, where the mean and the log variance are modeled as a linear combination of input variables. Existing literature on high-dimensional linear regres- sion models has largely ignored non-constant error variances, even though they commonly occur in a variety of applications ranging from biostatis- tics to finance. In this paper we study a class of non-convex penalized pseudolikelihood estimators for both the mean and variance parameters. We show that the Heteroscedastic Iterative Penalized Pseudolikelihood Optimizer (HIPPO) achieves the oracle property, that is, we prove that the rates of convergence are the same as if the true model was known. We demonstrate numerical properties of the procedure on a simulation study and real world data.


Approximate Dynamic Programming By Minimizing Distributionally Robust Bounds

arXiv.org Machine Learning

Approximate dynamic programming is a popular method for solving large Markov decision processes. This paper describes a new class of approximate dynamic programming (ADP) methods- distributionally robust ADP-that address the curse of dimensionality by minimizing a pessimistic bound on the policy loss. This approach turns ADP into an optimization problem, for which we derive new mathematical program formulations and analyze its properties. DRADP improves on the theoretical guarantees of existing ADP methods-it guarantees convergence and L1 norm based error bounds. The empirical evaluation of DRADP shows that the theoretical guarantees translate well into good performance on benchmark problems.


Latent Multi-group Membership Graph Model

arXiv.org Machine Learning

We develop the Latent Multi-group Membership Graph (LMMG) model, a model of networks with rich node feature structure. In the LMMG model, each node belongs to multiple groups and each latent group models the occurrence of links as well as the node feature structure. The LMMG can be used to summarize the network structure, to predict links between the nodes, and to predict missing features of a node. We derive efficient inference and learning algorithms and evaluate the predictive performance of the LMMG on several social and document network datasets.


Hypothesis testing using pairwise distances and associated kernels (with Appendix)

arXiv.org Machine Learning

We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as established in machine learning. The equivalence holds when energy distances are computed with semimetrics of negative type, in which case a kernel may be defined such that the RKHS distance between distributions corresponds exactly to the energy distance. We determine the class of probability distributions for which kernels induced by semimetrics are characteristic (that is, for which embeddings of the distributions to an RKHS are injective). Finally, we investigate the performance of this family of kernels in two-sample and independence tests: we show in particular that the energy distance most commonly employed in statistics is just one member of a parametric family of kernels, and that other choices from this family can yield more powerful tests.


Interactivity and Multimedia in Case-Based Recommendation

AAAI Conferences

The increasingly prevalent view that recommendation is a conversation between user and system is driving a renewed interest in approaches to system design that involve the user in meaningful ways. In addition to this the proliferation of mobile devices and the near-ubiquity of sensing technologies means that there are now many opportunities to capture real-life experiences, in real-time, providing a new source of raw material for case-based reasoning. In this paper we consider the availability of real-world exercise information, in this cases corresponding to jogging routes, and meth- ods by which we can involve a user in recommending such routes. We describe the Exercise Builder, a proof-of-concept application that attempts to help visitors to a new city to plan their jogging routes by combining case retrieval, interactive adaptation, and multimedia explanation in a single online service.


Robot Localization Using Overhead Camera and LEDs

AAAI Conferences

Determining the position of a robot in an environment, termed localization, is one of the challenges facing roboticist. Localization is essential to solving more complex problems such as locomotion, path planning and environmental learning. Our lab is developing a multi-agent system to use multiple small robots to accomplish tasks normally completed by larger robots. However, because of the reduced size of these robots, methods previously used to determine the position of the robot, such as GPS, cannot be employed. The problem we are facing is that we need to be able to determine the position of each of the robots in this multi-agent system simultaneously. We have developed a system to help track and identify robots using an overhead camera and LEDs, mounted on the robots, to efficiently solve the localization problem.


Classifying Scientific Performance on a Metric-by-Metric Basis

AAAI Conferences

In this paper, we outline a system for evaluating the performance of scientific research across a number of outcome metrics (e.g. publications, sales, new hires). Our system is designed to classify research performance into a number of metrics, evaluate each metricโ€™s performance using only data on other metrics, and to cast predictions of future performance by metric. This study shows how data mining techniques can be used to provide a predictive analytic approach to the management of resources for scientific research.


Ant Hunt: Towards a Validated Model of Live Ant Hunting Behavior

AAAI Conferences

Biologists seek concise, testable models of behavior for the animals they study. We suggest a robot programming paradigm in which animal behaviors are described as robot controllers to support a cycle of hypothesis generation and testing of animal models. In this work we illustrate that approach by modeling the hunting behavior of a captive colony of Aphaenogaster cockerelli , a desert harvester ant. In laboratory animal experiments we introduce live prey (fruit flies) into the foraging arena of the colony. We observe the behavior of the ants, and we measure aspects of their performance in capturing the prey. Based on these observations we create a model of their behavior using Clay, a Java library developed for coding hybrid controllers in a behavior-based manner. We then validate that model in quantitative comparisons with the live animal behavior.


Efficiency Improvements for Parallel Subgraph Miners

AAAI Conferences

Algorithms for finding frequent and/or interesting subgraphs in a single large graph scenario are computationally intensive because of the graph isomorphism and the subgraph isomorphism problem. These problems are compounded by the size of most real-world datasets which have sizes in the order of 105 or 106. The SUBDUE algorithm developed by Cook and Holder finds the most compressing subgraph in a large graph. In order to perform the same task on real-world data sets efficiently, Cook et al. developed a parallel approach to SUBDUE called the SP-SUBDUE based on the MPI framework. This paper extends the work done by Cook et al. to improve the efficiency of MPI SUBDUE by modifying the evaluation phase. Our experiments show an improvement in speed-up while retaining the quality of the results of serial SUBDUE. The techniques that we have used in this study can also be used in similar algorithms which use static partitioning of the data and re-evaluation of locally interesting patterns over all the nodes of the cluster.


Speech Acts, Dialogues and the Common Ground

AAAI Conferences

The formal semantics of speech acts, even in the classical framework of illocutionary logic, requires considerations that go beyond individual speech activity and beyond the interpretation of individual sentences. We show how the formal semantics of speech acts can be extended to take into account the social effects and interactive aspects of illocutionary activity. To illustrate our approach, we focus on the semantics of assertions and descriptive discourse, contrasting the individual aspect of speaker's meaning and the epistemic effects of assertion making. The approach presented in this paper generalizes to all other types of illocutionary acts, adding specific content to the conversational record that registers the common ground of speakers and hearers as a dialogue unfolds.