Genre
Data Science and Machine Learning Workshop
This specialised field demands multiple skills not easy to obtain through conventional curricula. So, be part of the data revolution by attending this workshop to learn the fundamentals of data science and machine learning and leave armed with practical skills to extract value from data. With the knowledge and skills gained from this workshop, you will be able to tackle complicated big data and machine learning challenges. Attend the workshop and develop the foundation level competence in data science and finding, manipulating, managing, interpreting and visualizing data.
Amazon Alexa Prize offers 2.5M to get voice assistant chatting
Amazon is hoping some universities start getting chatty with Alexa. The online retailer announced on Thursday the Alexa Prize, a 2.5 million award for college students who develop technology to make it more natural to talk with Amazon's Alexa virtual assistant. The company said the goal of the competition is to build a "socialbot" on Alexa that will converse with people about popular topics and news events. The university team with the best performing socialbot will win a 500,000 prize. Amazon also will give 1 million to the winning team's university if its socialbot is able to talk "coherently and engagingly" with humans for 20 minutes.
A year of Alphabet: Great for Google, less so for moonshots
Reorganizing itself under the umbrella company Alphabet has done wonders for Google -- but less so for a grab bag of eclectic projects ranging from robotic cars to internet-beaming balloons, which are suffering costly growing pains. A year after Alphabet took shape, Google's revenue growth has accelerated -- an unusual development for a company of its size. That success, however, also underscores Alphabet's dependence on the fickle business of placing digital ads in core Google products like search, Gmail and YouTube video. As a result, it remains vulnerable to swings in marketing budgets and stiffening competition from another equally ambitious rival, Facebook. Alphabet was supposed to speed the process of turning offshoot businesses into new technological jackpots.
How to steal the mind of an AI: Machine-learning models vulnerable to reverse engineering
Amazon, Baidu, Facebook, Google and Microsoft, among other technology companies, have been investing heavily in artificial intelligence and related disciplines like machine learning because they see the technology enabling services that become a source of revenue. Consultancy Accenture earlier this week quantified this enthusiasm, predicting that AI "could double annual economic growth rates by 2035 by changing the nature of work and spawning a new relationship between man and machine" and by boosting labor productivity by 40 per cent. Certainly things could work out well for Accenture, which a day later announced a partnership with Google to help companies deploy Google technology like machine learning. It's as if the global services firm has a stake in the future it foresees. But the machine learning algorithms underpinning this harmonious union of people and circuits aren't secure. In a paper [PDF] presented in August at the 25th Annual Usenix Security Symposium, researchers at รcole Polytechnique Fรฉdรฉrale de Lausanne, Cornell University, and The University of North Carolina at Chapel Hill showed that machine learning models can be stolen and that basic security measures don't really mitigate attacks.
Two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS)
Firouzi, Hamed, Hero, Alfred, Rajaratnam, Bala
This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). SPARCS can be applied to high dimensional prediction problems in experimental science, medicine, finance, and engineering, as illustrated by the following. Suppose one wishes to run a sequence of experiments to learn a sparse multivariate predictor of a dependent variable $Y$ (disease prognosis for instance) based on a $p$ dimensional set of independent variables $\mathbf X=[X_1,\ldots, X_p]^T$ (assayed biomarkers). Assume that the cost of acquiring the full set of variables $\mathbf X$ increases linearly in its dimension. SPARCS breaks the data collection into two stages in order to achieve an optimal tradeoff between sampling cost and predictor performance. In the first stage we collect a few ($n$) expensive samples $\{y_i,\mathbf x_i\}_{i=1}^n$, at the full dimension $p\gg n$ of $\mathbf X$, winnowing the number of variables down to a smaller dimension $l < p$ using a type of cross-correlation or regression coefficient screening. In the second stage we collect a larger number $(t-n)$ of cheaper samples of the $l$ variables that passed the screening of the first stage. At the second stage, a low dimensional predictor is constructed by solving the standard regression problem using all $t$ samples of the selected variables. SPARCS is an adaptive online algorithm that implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds for the Familywise Error Rate (FWER), specify high dimensional convergence rates for support recovery, and establish optimal sample allocation rules to the first and second stages.
Tuning Parameter Calibration in High-dimensional Logistic Regression With Theoretical Guarantees
Feature selection is a standard approach to understanding and modeling high-dimensional classification data, but the corresponding statistical methods hinge on tuning parameters that are difficult to calibrate. In particular, existing calibration schemes in the logistic regression framework lack any finite sample guarantees. In this paper, we introduce a novel calibration scheme for penalized logistic regression. It is based on simple tests along the tuning parameter path and satisfies optimal finite sample bounds. It is also amenable to easy and efficient implementations, and it rivals or outmatches existing methods in simulations and real data applications.
Convergence of a Grassmannian Gradient Descent Algorithm for Subspace Estimation From Undersampled Data
Subspace learning and matrix factorization problems have a great many applications in science and engineering, and efficient algorithms are critical as dataset sizes continue to grow. Many relevant problem formulations are non-convex, and in a variety of contexts it has been observed that solving the non-convex problem directly is not only efficient but reliably accurate. We discuss convergence theory for a particular method: first order incremental gradient descent constrained to the Grassmannian. The output of the algorithm is an orthonormal basis for a $d$-dimensional subspace spanned by an input streaming data matrix. We study two sampling cases: where each data vector of the streaming matrix is fully sampled, or where it is undersampled by a sampling matrix $A_t\in \R^{m\times n}$ with $m\ll n$. We propose an adaptive stepsize scheme that depends only on the sampled data and algorithm outputs. We prove that with fully sampled data, the stepsize scheme maximizes the improvement of our convergence metric at each iteration, and this method converges from any random initialization to the true subspace, despite the non-convex formulation and orthogonality constraints. For the case of undersampled data, we establish monotonic improvement on the defined convergence metric for each iteration with high probability.
A Birth and Death Process for Bayesian Network Structure Inference
Bayesian networks (Pearl [13]) are convenient graphical expressions for high dimensional probability distributions representing complex relationships between a large number of random variables. A Bayesian network is a directed acyclic graph consisting of nodes which represent random variables and arrows which correspond to probabilistic dependencies between them. There has been a great deal of interest in recent years on the NPhard problem of learning the structure (placement of directed edges) of Bayesian networks from data ([1],[2],[4],[5], [6],[8],[9],[11],[12]). Much of this has been driven by the study of genetic regulatory networks in molecular biology due to advances in technology and, specifically, microarray techniques that allow scientists to rapidly measure expression levels of genes in cells. As an integral part of machine learning, Bayesian networks have also been used for pattern recognition, language processing including speech recognition, and credit risk analysis. Structure learning typically involves defining a network score function and is then, in theory, a straightforward optimization problem.
Provable Burer-Monteiro factorization for a class of norm-constrained matrix problems
Park, Dohyung, Kyrillidis, Anastasios, Bhojanapalli, Srinadh, Caramanis, Constantine, Sanghavi, Sujay
We study the projected gradient descent method on low-rank matrix problems with a strongly convex objective. We use the Burer-Monteiro factorization approach to implicitly enforce low-rankness; such factorization introduces non-convexity in the objective. We focus on constraint sets that include both positive semi-definite (PSD) constraints and specific matrix norm-constraints. Such criteria appear in quantum state tomography and phase retrieval applications. We show that non-convex projected gradient descent favors local linear convergence in the factored space. We build our theory on a novel descent lemma, that non-trivially extends recent results on the unconstrained problem. The resulting algorithm is Projected Factored Gradient Descent, abbreviated as ProjFGD, and shows superior performance compared to state of the art on quantum state tomography and sparse phase retrieval applications.
Data-Driven Learning of a Union of Sparsifying Transforms Model for Blind Compressed Sensing
Ravishankar, Saiprasad, Bresler, Yoram
Compressed sensing is a powerful tool in applications such as magnetic resonance imaging (MRI). It enables accurate recovery of images from highly undersampled measurements by exploiting the sparsity of the images or image patches in a transform domain or dictionary. In this work, we focus on blind compressed sensing (BCS), where the underlying sparse signal model is a priori unknown, and propose a framework to simultaneously reconstruct the underlying image as well as the unknown model from highly undersampled measurements. Specifically, our model is that the patches of the underlying image(s) are approximately sparse in a transform domain. We also extend this model to a union of transforms model that better captures the diversity of features in natural images. The proposed block coordinate descent type algorithms for blind compressed sensing are highly efficient, and are guaranteed to converge to at least the partial global and partial local minimizers of the highly non-convex BCS problems. Our numerical experiments show that the proposed framework usually leads to better quality of image reconstructions in MRI compared to several recent image reconstruction methods. Importantly, the learning of a union of sparsifying transforms leads to better image reconstructions than a single adaptive transform.