Statistical Learning
Generalized Random Utility Models with Multiple Types
We propose a model for demand estimation in multi-agent, differentiated product settings and present an estimation algorithm that uses reversible jump MCMC techniques to classify agents' types. Our model extends the popular setup in Berry, Levinsohn and Pakes (1995) to allow for the data-driven classification of agents' types using agent-level data. We focus on applications involving data on agents' ranking over alternatives, and present theoretical conditions that establish the identifiability of the model and uni-modality of the likelihood/posterior. Results on both real and simulated data provide support for the scalability of our approach.
c4851e8e264415c4094e4e85b0baa7cc-Reviews.html
This paper considers automatic classification of unstructured social group activity videos. To bridge the semantic gap between low-level features and the class-labels, the authors adopt a latent topic model based on replicated softmax to extract topics as mid-level representations for video classification. The main idea of this paper is the integration of sparse Bayesian learning and replicated softmax, which leads to the proposed model referred to "relevance topic model (RTM)". In RTM, the discriminative topics and sparse classifier weights are learned jointly, and the authors proposes variational EM algorithm for model parameter estimation and inference. The authors test their algorithm on a benchmark dataset and demonstrate better performance compared to other supervised topic models and some baseline algorithms.
Relevance Topic Model for Unstructured Social Group Activity Recognition
Unstructured social group activity recognition in web videos is a challenging task due to 1) the semantic gap between class labels and low-level visual features and 2) the lack of labeled training data. To tackle this problem, we propose a "relevance topic model" for jointly learning meaningful mid-level representations upon bagof-words (BoW) video representations and a classifier with sparse weights. In our approach, sparse Bayesian learning is incorporated into an undirected topic model (i.e., Replicated Softmax) to discover topics which are relevant to video classes and suitable for prediction. Rectified linear units are utilized to increase the expressive power of topics so as to explain better video data containing complex contents and make variational inference tractable for the proposed model. An efficient variational EM algorithm is presented for model parameter estimation and inference. Experimental results on the Unstructured Social Activity Attribute dataset show that our model achieves state of the art performance and outperforms other supervised topic model in terms of classification accuracy, particularly in the case of a very small number of labeled training videos.
Modeling Clutter Perception using Parametric Proto-object Partitioning
Visual clutter, the perception of an image as being crowded and disordered, affects aspects of our lives ranging from object detection to aesthetics, yet relatively little effort has been made to model this important and ubiquitous percept. Our approach models clutter as the number of proto-objects segmented from an image, with proto-objects defined as groupings of superpixels that are similar in intensity, color, and gradient orientation features. We introduce a novel parametric method of clustering superpixels by modeling mixture of Weibulls on Earth Mover's Distance statistics, then taking the normalized number of proto-objects following partitioning as our estimate of clutter perception. We validated this model using a new 90-image dataset of real world scenes rank ordered by human raters for clutter, and showed that our method not only predicted clutter extremely well (Spearman's ρ = 0.8038, p < 0.001), but also outperformed all existing clutter perception models and even a behavioral object segmentation ground truth. We conclude that the number of proto-objects in an image affects clutter perception more than the number of objects or features.
c3e878e27f52e2a57ace4d9a76fd9acf-Reviews.html
Simply showing that there is information about actions, actors (and their conjunction) in any part of the brain does not mean they have tackled a "(three-fold) challenge". One would find (presumably) very similar response properties in any part of the mirror neuron system. Furthermore, if one analysed retinal cells, one would also find this information. Perhaps you could highlight the fact that you have found information or invariance properties at the level of the single neuron - that could not be found at low levels in the visual hierarchy - to make your point more clearly?
Speeding up Permutation Testing in Neuroimaging Chris Hinrichs Qinyuan Sun Sterling C. Johnson
Multiple hypothesis testing is a significant problem in nearly all neuroimaging studies. In order to correct for this phenomena, we require a reliable estimate of the Family-Wise Error Rate (FWER). The well known Bonferroni correction method, while simple to implement, is quite conservative, and can substantially under-power a study because it ignores dependencies between test statistics. Permutation testing, on the other hand, is an exact, non-parametric method of estimating the FWER for a given α-threshold, but for acceptably low thresholds the computational burden can be prohibitive. In this paper, we show that permutation testing in fact amounts to populating the columns of a very large matrix P. By analyzing the spectrum of this matrix, under certain conditions, we see that P has a low-rank plus a low-variance residual decomposition which makes it suitable for highly sub-sampled -- on the order of 0.5% -- matrix completion methods. Based on this observation, we propose a novel permutation testing methodology which offers a large speedup, without sacrificing the fidelity of the estimated FWER. Our evaluations on four different neuroimaging datasets show that a computational speedup factor of roughly 50 can be achieved while recovering the FWER distribution up to very high accuracy. Further, we show that the estimated α-threshold is also recovered faithfully, and is stable.
Learning Trajectory Preferences for Manipulators via Iterative Improvement
We consider the problem of learning good trajectories for manipulation tasks. This is challenging because the criterion defining a good trajectory varies with users, tasks and environments. In this paper, we propose a co-active online learning framework for teaching robots the preferences of its users for object manipulation tasks. The key novelty of our approach lies in the type of feedback expected from the user: the human user does not need to demonstrate optimal trajectories as training data, but merely needs to iteratively provide trajectories that slightly improve over the trajectory currently proposed by the system. We argue that this co-active preference feedback can be more easily elicited from the user than demonstrations of optimal trajectories, which are often challenging and non-intuitive to provide on high degrees of freedom manipulators. Nevertheless, theoretical regret bounds of our algorithm match the asymptotic rates of optimal trajectory algorithms. We demonstrate the generalizability of our algorithm on a variety of grocery checkout tasks, for whom, the preferences were not only influenced by the object being manipulated but also by the surrounding environment.
Non-Linear Domain Adaptation with Boosting Carlos Becker
A common assumption in machine vision is that the training and test samples are drawn from the same distribution. However, there are many problems when this assumption is grossly violated, as in bio-medical applications where different acquisitions can generate drastic variations in the appearance of the data due to changing experimental conditions. This problem is accentuated with 3D data, for which annotation is very time-consuming, limiting the amount of data that can be labeled in new acquisitions for training. In this paper we present a multi-task learning algorithm for domain adaptation based on boosting. Unlike previous approaches that learn task-specific decision boundaries, our method learns a single decision boundary in a shared feature space, common to all tasks. We use the boosting-trick to learn a non-linear mapping of the observations in each task, with no need for specific a-priori knowledge of its global analytical form. This yields a more parameter-free domain adaptation approach that successfully leverages learning on new tasks where labeled data is scarce. We evaluate our approach on two challenging bio-medical datasets and achieve a significant improvement over the state of the art.
On the Sample Complexity of Subspace Learning
A large number of algorithms in machine learning, from principal component analysis (PCA), and its non-linear (kernel) extensions, to more recent spectral embedding and support estimation methods, rely on estimating a linear subspace from samples. In this paper we introduce a general formulation of this problem and derive novel learning error estimates. Our results rely on natural assumptions on the spectral properties of the covariance operator associated to the data distribution, and hold for a wide class of metrics between subspaces. As special cases, we discuss sharp error estimates for the reconstruction properties of PCA and spectral support estimation. Key to our analysis is an operator theoretic approach that has broad applicability to spectral learning methods.
bca82e41ee7b0833588399b1fcd177c7-Reviews.html
The authors propose a parallel algorithm for the DPMM that parallelizes a RJMCMC sampler that jumps between finite models. While the parallelization and the RJMCMC sampler are proposed together, I will separate them for the purpose of this review, in order to ask questions about each part separately. First, the RJMCMC algorithm (by which I mean, the algorithm we would have on a single cluster). Here, we use a reversible-jump MCMC algorithm to jump between finite-dimensional Dirichlet distributions. As an aside, since \bar{\pi}_{K 1} is not used in the mixture model (the mixture model is defined on the renormalized occupied K components), it would seem to make more sense to define a K-dimensional, rather than a K-1 - dimensional, Dirichlet distribution; this is valid under marginalization properties of the Dirichlet distribution, since equation 10 samples from a distribution proportional to \pi_1 ... \pi_K To jump between model dimensionalities, the authors propose a split/merge RJMCMC step that is reminiscent of that of Green and Richardson.