Goto

Collaborating Authors

 Genre


Noisy subspace clustering via matching pursuits

arXiv.org Machine Learning

Sparsity-based subspace clustering algorithms have attracted significant attention thanks to their excellent performance in practical applications. A prominent example is the sparse subspace clustering (SSC) algorithm by Elhamifar and Vidal, which performs spectral clustering based on an adjacency matrix obtained by sparsely representing each data point in terms of all the other data points via the Lasso. When the number of data points is large or the dimension of the ambient space is high, the computational complexity of SSC quickly becomes prohibitive. Dyer et al. observed that SSC-OMP obtained by replacing the Lasso by the greedy orthogonal matching pursuit (OMP) algorithm results in significantly lower computational complexity, while often yielding comparable performance. The central goal of this paper is an analytical performance characterization of SSC-OMP for noisy data. Moreover, we introduce and analyze the SSC-MP algorithm, which employs matching pursuit (MP) in lieu of OMP. Both SSC-OMP and SSC-MP are proven to succeed even when the subspaces intersect and when the data points are contaminated by severe noise. The clustering conditions we obtain for SSC-OMP and SSC-MP are similar to those for SSC and for the thresholding-based subspace clustering (TSC) algorithm due to Heckel and B\"olcskei. Analytical results in combination with numerical results indicate that both SSC-OMP and SSC-MP with a data-dependent stopping criterion automatically detect the dimensions of the subspaces underlying the data. Moreover, experiments on synthetic and real data show that SSC-MP compares very favorably to SSC, SSC-OMP, TSC, and the nearest subspace neighbor (NSN) algorithm, both in terms of clustering performance and running time. In addition, we find that, in contrast to SSC-OMP, the performance of SSC-MP is very robust with respect to the choice of parameters in the stopping criteria.


Identity Matters in Deep Learning

arXiv.org Machine Learning

An emerging design principle in deep learning is that each layer of a deep artificial neural network should be able to easily express the identity transformation. This idea not only motivated various normalization techniques, such as \emph{batch normalization}, but was also key to the immense success of \emph{residual networks}. In this work, we put the principle of \emph{identity parameterization} on a more solid theoretical footing alongside further empirical progress. We first give a strikingly simple proof that arbitrarily deep linear residual networks have no spurious local optima. The same result for linear feed-forward networks in their standard parameterization is substantially more delicate. Second, we show that residual networks with ReLu activations have universal finite-sample expressivity in the sense that the network can represent any function of its sample provided that the model has more parameters than the sample size. Directly inspired by our theory, we experiment with a radically simple residual architecture consisting of only residual convolutional layers and ReLu activations, but no batch normalization, dropout, or max pool. Our model improves significantly on previous all-convolutional networks on the CIFAR10, CIFAR100, and ImageNet classification benchmarks.


SCOPE: Scalable Composite Optimization for Learning on Spark

arXiv.org Machine Learning

Many machine learning models, such as logistic regression~(LR) and support vector machine~(SVM), can be formulated as composite optimization problems. Recently, many distributed stochastic optimization~(DSO) methods have been proposed to solve the large-scale composite optimization problems, which have shown better performance than traditional batch methods. However, most of these DSO methods are not scalable enough. In this paper, we propose a novel DSO method, called \underline{s}calable \underline{c}omposite \underline{op}timization for l\underline{e}arning~({SCOPE}), and implement it on the fault-tolerant distributed platform \mbox{Spark}. SCOPE is both computation-efficient and communication-efficient. Theoretical analysis shows that SCOPE is convergent with linear convergence rate when the objective function is convex. Furthermore, empirical results on real datasets show that SCOPE can outperform other state-of-the-art distributed learning methods on Spark, including both batch learning methods and DSO methods.


Self-calibrating Neural Networks for Dimensionality Reduction

arXiv.org Machine Learning

Recently, a novel family of biologically plausible online algorithms for reducing the dimensionality of streaming data has been derived from the similarity matching principle. In these algorithms, the number of output dimensions can be determined adaptively by thresholding the singular values of the input data matrix. However, setting such threshold requires knowing the magnitude of the desired singular values in advance. Here we propose online algorithms where the threshold is self-calibrating based on the singular values computed from the existing observations. To derive these algorithms from the similarity matching cost function we propose novel regularizers. As before, these online algorithms can be implemented by Hebbian/anti-Hebbian neural networks in which the learning rule depends on the chosen regularizer. We demonstrate both mathematically and via simulation the effectiveness of these online algorithms in various settings.


Bayes Theorem: A Visual Introduction For Beginners

#artificialintelligence

From Google search results to Netflix recommendations and investment strategies, Bayes Theorem (also often called Bayes Rule or Bayes Formula) is used across countless industries to help calculate and assess probability. Bayesian statistics is taught in most first-year statistics classes across the nation, but there is one major problem that many students (and others who are interested in the theorem) face. The theorem is not intuitive for most people, and understanding how it works can be a challenge, especially because it is often taught without visual aids. In this guide, we unpack the various components of the theorem and provide a basic overview of how it works – and with illustrations to help. Three scenarios – the flu, breathalyzer tests, and peacekeeping – are used throughout the booklet to teach how problems involving Bayes Theorem can be approached and solved.


Are you smart enough to work at Google?

@machinelearnbot

This was the title of a very popular book published in 2012, featuring several job interview questions (brain teasers) asked by Google's hiring managers to candidates. They apparently dropped all these questions, as they found out that they were not good indicators of career success. I had one phone interview with Google long ago, and was rejected right away. The interviewer was just focused on very technical details, and spent all her time arguing about Lasso regression, and was clearly looking for a specialist, dismissing people with a broad range of skills and non-standard approach to solving tech problems. Big companies do not value things like intuition, innovation, vision or a disruptive mindset (despite claiming the contrary), and for good reasons.


Easy fixes to tech problems

FOX News

From the looks of it, 2017 will be a pretty amazing year for technology. Every day consumers are talking about virtual reality, self-driving cars and a limitless menu of on-demand services. The future really has arrived. But even as digital tech gets more streamlined and powerful, glitches keep popping up. Just when we think we have superhuman control of our lives, a device fails to work and we have no clue how to fix it.


AI Teaching Assistant Helped Students Online--and No One Knew the Difference

#artificialintelligence

Meet Jill Watson, a first-time teaching assistant at Georgia Tech assigned to moderate an online forum for a computer science class. Jill was 1 of 9 TAs assigned to help answer questions about coursework and projects from the 300 students enrolled in the advanced course. During the first few weeks in January, Jill really struggled. This was Knowledge-Based Artificial Intelligence, after all, a course with the goal to "build AI agents capable of human-level intelligence and gain insights into human cognition." It was also a requirement for graduate students to earn their master's degree.


Yann LeCun Lecture 1/8 Why Deep Learning ?

#artificialintelligence

Want to watch this again later? Report Need to report the video? Report Need to report the video? Need to report the video? This feature is not available right now.


Technology Vs. Human - Who Is Going To Win? An Interview With Gerd Leonhard

#artificialintelligence

I remember meeting Gerd Leonhard [Futurist, Author and a raft of other titles] for the first time in a particularly crowded Benugo in Covent Garden. The meeting came after several near misses and during one of his gigs in London and we decided just to wing it. The fries were unmemorable but the conversation probably set me on the path I find myself travelling today. I have quizzed him on his latest book "Technology Vs. Humanity" [Amazon] in which he poses some interesting questions about the future of the human race and technology but essentially asks; "Are you on team human, or not...? We are at a pivot point in human history [and you need to choose]."