Goto

Collaborating Authors

 Statistical Learning


Reproducing Kernel Banach Spaces with the l1 Norm

arXiv.org Machine Learning

Targeting at sparse learning, we construct Banach spaces B of functions on an input space X with the properties that (1) B possesses an l1 norm in the sense that it is isometrically isomorphic to the Banach space of integrable functions on X with respect to the counting measure; (2) point evaluations are continuous linear functionals on B and are representable through a bilinear form with a kernel function; (3) regularized learning schemes on B satisfy the linear representer theorem. Examples of kernel functions admissible for the construction of such spaces are given.


Analysis of a Random Forests Model

arXiv.org Machine Learning

Random forests are a scheme proposed by Leo Breiman in the 2000's for building a predictor ensemble with a set of decision trees that grow in randomly selected subspaces of data. Despite growing interest and practical use, there has been little exploration of the statistical properties of random forests, and little is known about the mathematical forces driving the algorithm. In this paper, we offer an in-depth analysis of a random forests model suggested by Breiman in \cite{Bre04}, which is very close to the original algorithm. We show in particular that the procedure is consistent and adapts to sparsity, in the sense that its rate of convergence depends only on the number of strong features and not on how many noise variables are present.


Random Feature Maps for Dot Product Kernels

arXiv.org Machine Learning

Approximating non-linear kernels using feature maps has gained a lot of interest in recent years due to applications in reducing training and testing times of SVM classifiers and other kernel based learning algorithms. We extend this line of work and present low distortion embeddings for dot product kernels into linear Euclidean spaces. We base our results on a classical result in harmonic analysis characterizing all dot product kernels and use it to define randomized feature maps into explicit low dimensional Euclidean spaces in which the native dot product provides an approximation to the dot product kernel with high confidence.


Meditation Training and Neurofeedback Using a Personal EEG Device

AAAI Conferences

Baseline and meditation data was obtained from 31 longterm meditation practitioners using the single-sensor right Over the past several years, a host of simple consumer prefrontal EEG system produced by Neurosky, Inc. Each electroencephalography (EEG) devices have been released subject was asked to complete a 5 minute resting period in at relatively inexpensive price points. These devices allow which they were asked to close their eyes and let their single or multi-channel recording of EEG, generally mind wander (without meditating). This was followed by employing user-friendly design, e.g.


Tracking Epidemics with Natural Language Processing and Crowdsourcing

AAAI Conferences

The first indication of a new outbreak is often in unstructured data (natural language) and reported openly in traditional or social media as a new `flu-like' or `malaria-like' illness weeks or months before the new pathogen is eventually isolated. We present a system for tracking these early signals globally, using natural language processing and crowdsourcing. By comparison, search-log-based approaches, while innovative and inexpensive, are often a trailing signal that follow open reports in plain language. Concentrating on discovering outbreak-related reports in big open data, we show how crowdsourced workers can create near-real-time training data for adaptive active-learning models, addressing the lack of broad coverage training data for tracking epidemics. This is well-suited to an outbreak information-flow context, where sudden bursts of information about new diseases/locations need to be manually processed quickly at short notice.


A Semantic Metadirectory of Services Based on Web Mining Techniques

AAAI Conferences

In the current web, developers are able to create new applications by composing already existing services from third-party vendors. However, the vast amount of choices, technologies and repositories can make it a tedious task. This paper describes a semantic metadirectory of services that helps in the process of discovering services. We propose a semantic service discovery process and description of existing service repositories, such as Programmable Web and Yahoo Pipes, which are two service repositories which provide plenty of services that can be reused by developers to build new web applications. The challenges behind integrating these repositories involved the problems of defining a common model, identifying relevant data and integrating and ranking the extracted data.


Adaptive Learning Agents for Sustainable Building Energy Management.

AAAI Conferences

Nearly 20% of total energy consumption in the United States is accounted for in heating, ventilation, and air conditioning (HVAC) systems. Smart sensing and adaptive energy management agents can greatly decrease the energy usage of HVAC systems in many building applications, for example by enabling the operator to shut off HVAC to unoccupied rooms. We implement a multimodal sensor agent that is nonintrusive and low-cost, combining information such as motion detection, CO2 reading, sound level, ambient light,and door state sensing. We show that in our live test bed at the USC campus, these sensor agents can be used to accurately estimate the number of occupants in each room using machine learning techniques, and that these techniques can also be applied to predict future occupancy by creating agent models of the occupants. These predictions will be used by control agents to enable the HVAC system increase its efficiency by continuously adapting to occupancy forecasts of each room.


A Regularization Approach for Prediction of Edges and Node Features in Dynamic Graphs

arXiv.org Machine Learning

We consider the two problems of predicting links in a dynamic graph sequence and predicting functions defined at each node of the graph. In many applications, the solution of one problem is useful for solving the other. Indeed, if these functions reflect node features, then they are related through the graph structure. In this paper, we formulate a hybrid approach that simultaneously learns the structure of the graph and predicts the values of the node-related functions. Our approach is based on the optimization of a joint regularization objective. We empirically test the benefits of the proposed method with both synthetic and real data. The results indicate that joint regularization improves prediction performance over the graph evolution and the node features.


Semi-blind Sparse Image Reconstruction with Application to MRFM

arXiv.org Machine Learning

We propose a solution to the image deconvolution problem where the convolution kernel or point spread function (PSF) is assumed to be only partially known. Small perturbations generated from the model are exploited to produce a few principal components explaining the PSF uncertainty in a high dimensional space. Unlike recent developments on blind deconvolution of natural images, we assume the image is sparse in the pixel basis, a natural sparsity arising in magnetic resonance force microscopy (MRFM). Our approach adopts a Bayesian Metropolis-within-Gibbs sampling framework. The performance of our Bayesian semi-blind algorithm for sparse images is superior to previously proposed semi-blind algorithms such as the alternating minimization (AM) algorithm and blind algorithms developed for natural images. We illustrate our myopic algorithm on real MRFM tobacco virus data.


Asymptotic Confidence Sets for General Nonparametric Regression and Classification by Regularized Kernel Methods

arXiv.org Machine Learning

Regularized kernel methods such as, e.g., support vector machines and least-squares support vector regression constitute an important class of standard learning algorithms in machine learning. Theoretical investigations concerning asymptotic properties have manly focused on rates of convergence during the last years but there are only very few and limited (asymptotic) results on statistical inference so far. As this is a serious limitation for their use in mathematical statistics, the goal of the article is to fill this gap. Based on asymptotic normality of many of these methods, the article derives a strongly consistent estimator for the unknown covariance matrix of the limiting normal distribution. In this way, we obtain asymptotically correct confidence sets for $\psi(f_{P,\lambda_0})$ where $f_{P,\lambda_0}$ denotes the minimizer of the regularized risk in the reproducing kernel Hilbert space $H$ and $\psi:H\rightarrow\mathds{R}^m$ is any Hadamard-differentiable functional. Applications include (multivariate) pointwise confidence sets for values of $f_{P,\lambda_0}$ and confidence sets for gradients, integrals, and norms.