Goto

Collaborating Authors

 Instructional Material


Checking in with Andrew Ng at Baidu's Blooming Silicon Valley Research Lab

IEEE Spectrum Robotics

Scatterings of completed buildings, sporting new plantings of drought-tolerant grasses, are already occupied; other buildings are going up quickly, including a new fire station. There's Nissan's new Silicon Valley research center, a well-financed medical device startup called Spiracur, a digital cash startup called Quisk, and a biotech startup incubator. And there is Baidu's Silicon Valley AI Lab--my destination along this dusty road crowded with construction vehicles. It's good to spend time in a new research lab; there's not only fresh paint and hip decor--like living walls of plants--there are fresh, excited faces, and empty desks waiting to be filled. In mid-2014, I spent a morning on just the other side of nearby Moffett Field watching a far more somber group of researchers moving out of a suddenly closed division of Microsoft Research.


How to learn Machine Learning?

#artificialintelligence

Some time ago I started a journey into one of the most exciting fields in Computer Science -- Machine Learning. This is my subjective guide for anyone who would like to explore this topic, but don't know how to start. Your first steps should lead to Stanford Machine Learning class at Coursera by Andrew Ng. This course is simply brilliant! Along a way, you will be given everything you need to know, including algebra review.


Intelligent Conversational Agents as Facilitators and Coordinators for Group Work in Distributed Learning Environments (MOOCs)

AAAI Conferences

Artificially intelligent conversational agents have been demonstrated to positively impact team based learning in classrooms and hold even greater potential for impact in the now widespread Massive Open Online Courses (MOOCs) if certain challenges can be overcome. These challenges include team formation, coordination and management of group processes in teams working together while distributed both in time and space. Our work begins with an architecture for orchestrating conversational agent based support for group learning called Bazaar, which has facilitated numerous successful studies of learning in the past including some early investigations in MOOC contexts. In this paper, we briefly describe our experience in designing, developing and deploying agent supported collaborative learning activities in 3 different MOOCs in three iterations. Findings from this iterative design process provide an empirical foundation for a reusable framework for facilitating similar activities in future MOOCs.


Sequential Monte Carlo Methods for System Identification

arXiv.org Machine Learning

One of the key challenges in identifying nonlinear and possibly non-Gaussian state space models (SSMs) is the intractability of estimating the system state. Sequential Monte Carlo (SMC) methods, such as the particle filter (introduced more than two decades ago), provide numerical solutions to the nonlinear state estimation problems arising in SSMs. When combined with additional identification techniques, these algorithms provide solid solutions to the nonlinear system identification problem. We describe two general strategies for creating such combinations and discuss why SMC is a natural tool for implementing these strategies.


Blind Source Separation: Fundamentals and Recent Advances (A Tutorial Overview Presented at SBrT-2001)

arXiv.org Machine Learning

A number of people are found in a room and involved in loud conversations in groups, just as it would happen in a cocktail party. There might also be some background noise, which could be music, car noise from outside, etc. Each person in this room is therefore forced to listen to a mixture of speech sounds coming from various directions, along with some noise. These sounds may come directly to one's ear or have first suffered a sequence of reverberations because of their reflections on the room's walls. The problem of focusing one's listening attention on a particular speaker among this cacophony of conversations and noise has been known as the cocktail party problem [6]. It consists of separating a mixture of speech signals of different characteristics with noise added to it. The signals are a-priori unknown (one listens only to a combination of them) as is also the way they have been mixed. The above scenario is a good analog for many other examples of situations that demand for a separation of mixed signals with no presupposed knowledge on the signals and the system mixing them.


Dual Smoothing and Level Set Techniques for Variational Matrix Decomposition

arXiv.org Machine Learning

We focus on the robust principal component analysis (RPCA) problem, and review a range of old and new convex formulations for the problem and its variants. We then review dual smoothing and level set techniques in convex optimization, present several novel theoretical results, and apply the techniques on the RPCA problem. In the final sections, we show a range of numerical experiments for simulated and real-world problems.


Online Low-Rank Subspace Learning from Incomplete Data: A Bayesian View

arXiv.org Machine Learning

Extracting the underlying low-dimensional space where high-dimensional signals often reside has long been at the center of numerous algorithms in the signal processing and machine learning literature during the past few decades. At the same time, working with incomplete (partly observed) large scale datasets has recently been commonplace for diverse reasons. This so called {\it big data era} we are currently living calls for devising online subspace learning algorithms that can suitably handle incomplete data. Their envisaged objective is to {\it recursively} estimate the unknown subspace by processing streaming data sequentially, thus reducing computational complexity, while obviating the need for storing the whole dataset in memory. In this paper, an online variational Bayes subspace learning algorithm from partial observations is presented. To account for the unawareness of the true rank of the subspace, commonly met in practice, low-rankness is explicitly imposed on the sought subspace data matrix by exploiting sparse Bayesian learning principles. Moreover, sparsity, {\it simultaneously} to low-rankness, is favored on the subspace matrix by the sophisticated hierarchical Bayesian scheme that is adopted. In doing so, the proposed algorithm becomes adept in dealing with applications whereby the underlying subspace may be also sparse, as, e.g., in sparse dictionary learning problems. As shown, the new subspace tracking scheme outperforms its state-of-the-art counterparts in terms of estimation accuracy, in a variety of experiments conducted on simulated and real data.


Peer Grading in a Course on Algorithms and Data Structures: Machine Learning Algorithms do not Improve over Simple Baselines

arXiv.org Machine Learning

Peer grading is the process of students reviewing each others' work, such as homework submissions, and has lately become a popular mechanism used in massive open online courses (MOOCs). Intrigued by this idea, we used it in a course on algorithms and data structures at the University of Hamburg. Throughout the whole semester, students repeatedly handed in submissions to exercises, which were then evaluated both by teaching assistants and by a peer grading mechanism, yielding a large dataset of teacher and peer grades. We applied different statistical and machine learning methods to aggregate the peer grades in order to come up with accurate final grades for the submissions (supervised and unsupervised, methods based on numeric scores and ordinal rankings). Surprisingly, none of them improves over the baseline of using the mean peer grade as the final grade. We discuss a number of possible explanations for these results and present a thorough analysis of the generated dataset.


Compressed Online Dictionary Learning for Fast fMRI Decomposition

arXiv.org Machine Learning

ABSTRACT We present a method for fast resting-state fMRI spatial decompositions of very large datasets, based on the reduction of the temporal dimension before applying dictionary learning on concatenated individual records from groups of subjects. Introducing a measure of correspondence between spatial decompositions of rest fMRI, we demonstrates that time-reduced dictionary learning produces result as reliable as non-reduced decompositions. We also show that this reduction significantly improves computational scalability. Index Terms-- resting-state fMRI, sparse decomposition, dictionary learning, online learning, rangefinder 1. INTRODUCTION Resting-state fMRI data analysis traditionally implies, as an initial step, to decompose a set of raw 4D records (time-series sampled in a volumic voxel grid) into a sum of spatially located functional networks that isolate a part of the brain signals. Functional networks, that can be seen as a set of brain activation maps, form a relevant basis for the experiment signals that captures its essence in a low-dimensional space.


Image Denoising with Kernels based on Natural Image Relations

arXiv.org Machine Learning

A successful class of image denoising methods is based on Bayesian approaches working in wavelet representations. However, analytical estimates can be obtained only for particular combinations of analytical models of signal and noise, thus precluding its straightforward extension to deal with other arbitrary noise sources. In this paper, we propose an alternative non-explicit way to take into account the relations among natural image wavelet coefficients for denoising: we use support vector regression (SVR) in the wavelet domain to enforce these relations in the estimated signal. Since relations among the coefficients are specific to the signal, the regularization property of SVR is exploited to remove the noise, which does not share this feature. The specific signal relations are encoded in an anisotropic kernel obtained from mutual information measures computed on a representative image database. Training considers minimizing the Kullback-Leibler divergence (KLD) between the estimated and actual probability functions of signal and noise in order to enforce similarity. Due to its non-parametric nature, the method can eventually cope with different noise sources without the need of an explicit re-formulation, as it is strictly necessary under parametric Bayesian formalisms. Results under several noise levels and noise sources show that: (1) the proposed method outperforms conventional wavelet methods that assume coefficient independence, (2) it is similar to state-of-the-art methods that do explicitly include these relations when the noise source is Gaussian, and (3) it gives better numerical and visual performance when more complex, realistic noise sources are considered. Therefore, the proposed machine learning approach can be seen as a more flexible (model-free) alternative to the explicit description of wavelet coefficient relations for image denoising.