Goto

Collaborating Authors

 Statistical Learning


Scikit-learn Tutorial: Machine Learning in Python

#artificialintelligence

Scikit-learn is a free machine learning library for Python. It features various algorithms like support vector machine, random forests, and k-neighbours, and it also supports Python numerical and scientific libraries like NumPy and SciPy. In this tutorial we will learn how to easily apply Machine Learning with the help of the scikit-learn library, which was created to make doing machine learning in Python easier and more robust. To do this, we'll be using the Sales_Win_Loss data set from IBM's Watson repository. We will import the data set using pandas, explore the data using pandas methods like head(), tail(), dtypes(), and then try our hand at using plotting techniques from Seaborn to visualize our data. Then we'll dive into scikit-learn and use preprocessing.LabelEncoder() in scikit-learn to process the data, and train_test_split() to split the data set into test and train samples. We will also use a cheat sheet to help us decide which algorithms to use for the data set. Finally we will use three different algorithms (Naive-Bayes, LinearSVC, K-Neighbors Classifier) to make predictions and compare their performance using methods like accuracy_score() provided by the scikit-learn library. We will also visualize the performance score of different models using scikit-learn and Yellowbrick visualization. If you need to brush up on these topics, check out these pandas and data visualization blog posts. For this tutorial, we will use the Sales-Win-Loss data set available on the IBM Watson website.


AI predicts precursors to heart attacks

#artificialintelligence

According to the Center for Disease Control and Prevention, over 610,000 people die of heart disease every year, which is the leading cause of death for both men and women in the U.S. Fortunately, scientists at IBM and pharmaceutical giant AstraZeneca are investigating a machine learning framework that can suss out early ASC warning signs. It's described in a newly published paper ("Outcome-Driven Clustering of Acute Coronary Syndrome Patients using Multi-Task Neural Network with Attention") on the preprint server Arxiv.org. The team sourced a dataset containing the age, gender, personal disease history, habits, laboratory test results, procedures, ACS type, and nearly 40 other characteristics of 26,986 adult hospitalized patients across 38 urban and rural hospitals in China, which they fed to a neural network -- i.e., layers of mathematical functions loosely modeled after biological neurons. Said neural network was architected to predict four factors simultaneously: whether they'd experienced a major adverse cardiac event, or MACE, prior to ACS; whether they'd received antiplatelet medicine to prevent blood clots from forming in the coronary arteries; whether they'd been given beta-blockers, which reduce blood pressure; and whether they were prescribed statins, a class of drugs that help lower cholesterol levels (and in turn prevent heart attacks and stroke). The paper's authors next employed k-means clustering -- a statistical technique in which data points are allocated to collections by similarities -- to organize the patients into seven groups based on the classification data obtained from the neural network.


Essential Machine Learning with Linear Models in RAPIDS: part 1 of a series.

#artificialintelligence

This blog is the first in a series about regression analysis in RAPIDS, an open GPU data science platform. There are many varieties of regression techniques, and we're working to include them all in RAPIDS. In this blog edition, I use Ordinary Least Squares (OLS) and Ridge regression to choose a model to predict Washington, D.C. bikeshare rentalsยน. I want to take a moment to tell the origin story of regression analysis, which will explain why it has that name. I believe that of all the common machine learning techniques (K-means, kNN, PCA), "regression analysis" has the most opaque name.


Nonlinear input design as optimal control of a Hamiltonian system

arXiv.org Machine Learning

We propose an input design method for a general class of parametric probabilistic models, including nonlinear dynamical systems with process noise. The goal of the procedure is to select inputs such that the parameter posterior distribution concentrates about the true value of the parameters; however, exact computation of the posterior is intractable. By representing (samples from) the posterior as trajectories from a certain Hamiltonian system, we transform the input design task into an optimal control problem. The method is illustrated via numerical examples, including MRI pulse sequence design.


A Rank-1 Sketch for Matrix Multiplicative Weights

arXiv.org Machine Learning

We show that a simple randomized sketch of the matrix multiplicative weight (MMW) update enjoys the same regret bounds as MMW, up to a small constant factor. Unlike MMW, where every step requires full matrix exponentiation, our steps require only a single product of the form $e^A b$, which the Lanczos method approximates efficiently. Our key technique is to view the sketch as a randomized mirror projection, and perform mirror descent analysis on the expected projection. Our sketch solves the online eigenvector problem, improving the best known complexity bounds. We also apply this sketch to a simple no-regret scheme for semidefinite programming in saddle-point form, where it matches the best known guarantees.


Fast Parallel Algorithms for Feature Selection

arXiv.org Machine Learning

In this paper, we analyze a fast parallel algorithm to efficiently select and build a set of $k$ random variables from a large set of $n$ candidate elements. This combinatorial optimization problem can be viewed in the context of feature selection for the prediction of a response variable. Using the adaptive sampling technique, which has recently been shown to exponentially speed up submodular maximization algorithms, we propose a new parallelizable algorithm that dramatically speeds up previous selection algorithms by reducing the number of rounds from $\mathcal O(k)$ to $\mathcal O(\log n)$ for objectives that do not conform to the submodularity property. We introduce a new metric to quantify the closeness of the objective function to submodularity and analyze the performance of adaptive sampling under this regime. We also conduct experiments on synthetic and real datasets and show that the empirical performance of adaptive sampling on not-submodular objectives greatly outperforms its theoretical lower bound. Additionally, the empirical running time drastically improved in all experiments without comprising the terminal value, showing the practicality of adaptive sampling.


Neural Empirical Bayes

arXiv.org Machine Learning

We formulate a novel framework that unifies kernel density estimation and empirical Bayes, where we address a broad set of problems in unsupervised learning with a geometric interpretation rooted in the concentration of measure phenomenon. We start by energy estimation based on a denoising objective which recovers the original/clean data X from its measured/noisy version Y with empirical Bayes least squares estimator. The setup is rooted in kernel density estimation, but the log-pdf in Y is parametrized with a neural network, and crucially, the learning objective is derived for any level of noise/kernel bandwidth. Learning is efficient with double backpropagation and stochastic gradient descent. An elegant physical picture emerges of an interacting system of high-dimensional spheres around each data point, together with a globally-defined probability flow field. The picture is powerful: it leads to a novel sampling algorithm, a new notion of associative memory, and it is instrumental in designing experiments. We start with extreme denoising experiments. Walk-jump sampling is defined by Langevin MCMC walks in Y, along with asynchronous empirical Bayes jumps to X. Robbins associative memory is defined by a deterministic flow to attractors of the learned probability flow field. Finally, we observed the emergence of remarkably rich creative modes in the regime of highly overlapping spheres.


A heuristic approach for lactate threshold estimation for training decision-making: An accessible and easy to use solution for recreational runners

arXiv.org Machine Learning

In this work, a heuristic as operational tool to estimate the lactate threshold and to facilitate its integration into the training process of recreational runners is proposed. To do so, we formalize the principles for the lactate threshold estimation from empirical data and an iterative methodology that enables experience based learning. This strategy arises as a robust and adaptive approach to solve data analysis problems. We compare the results of the heuristic with the most commonly used protocol by making a first quantitative error analysis to show its reliability. Additionally, we provide a computational algorithm so that this quantitative analysis can be easily performed in other lactate threshold protocols. With this work, we have shown that a heuristic %60 of 'endurance running speed reserve', serves for the same purpose of the most commonly used protocol in recreational runners, but improving its operational limitations of accessibility and consistent use.


Novel quantitative indicators of digital ophthalmoscopy image quality

arXiv.org Machine Learning

With the advent of smartphone indirect ophthalmoscopy, teleophthalmology - the use of specialist ophthalmology assets at a distance from the patient - has experienced a breakthrough, promising enormous benefits especially for healthcare in distant, inaccessible or opthalmologically underserved areas, where specialists are either unavailable or too few in number. However, accurate teleophthalmology requires high-quality ophthalmoscopic imagery. This paper considers three feature families - statistical metrics, gradient-based metrics and wavelet transform coefficient derived indicators - as possible metrics to identify unsharp or blurry images. By using standard machine learning techniques, the suitability of these features for image quality assessment is confirmed, albeit on a rather small data set. With the increased availability and decreasing cost of digital ophthalmoscopy on one hand and the increased prevalence of diabetic retinopathy worldwide on the other, creating tools that can determine whether an image is likely to be diagnostically suitable can play a significant role in accelerating and streamlining the teleophthalmology process. This paper highlights the need for more research in this area, including the compilation of a diverse database of ophthalmoscopic imagery, annotated with quality markers, to train the Point of Acquisition error detection algorithms of the future.


Generative Graph Convolutional Network for Growing Graphs

arXiv.org Machine Learning

Modeling generative process of growing graphs has wide applications in social networks and recommendation systems, where cold start problem leads to new nodes isolated from existing graph. Despite the emerging literature in learning graph representation and graph generation, most of them can not handle isolated new nodes without nontrivial modifications. The challenge arises due to the fact that learning to generate representations for nodes in observed graph relies heavily on topological features, whereas for new nodes only node attributes are available. Here we propose a unified generative graph convolutional network that learns node representations for all nodes adaptively in a generative model framework, by sampling graph generation sequences constructed from observed graph data. We optimize over a variational lower bound that consists of a graph reconstruction term and an adaptive Kullback-Leibler divergence regularization term. We demonstrate the superior performance of our approach on several benchmark citation network datasets.