Statistical Learning
Regulate Your Regression Model With Ridge, LASSO and ElasticNet
Linear models have a wide appeal. Even with a basic understanding of Excel, it is possible to create a model that explains patterns in data. After attaching weights (coefficients) to explanatory variables (features), it is easy to assess the importance of individual variables when explaining the data. It is not surprising that linear models have been around for many decades, and are widely used throughout many domains, ranging from psychology to business administration and from machine learning to statistics. Despite the superficial simplicity of linear models, many things can go wrong with them.
R Programming: Non-deterministic testing
Paired t-test: On many occasions, the two samples are not independent because they involve the same sampling unit. Measuring each subject or entity twice results in pairs of observations. In such situations, tests for paired data are used to perform hypothesis tests. This paired t-test relies on the assumption that the differences are normally distributed. The sign test is a non-parametric alternative to the paired t-test.
Develop and Operationalize ML models using plain SQL on Google BigQuery
Not too long ago, data deficiency was a major impediment towards making informed decisions, understanding customer behavior, predictions and forecasting. In the modern digital age, where data continuously streams in all shapes, sizes and from all directions, enterprises are constantly challenged with sifting through petabytes of data to infer key indicators. Making sense of "the right data at the right time" yields a huge competitive edge. Blending real-time streams, batch processing, external data sources and machine learning -- Google BigQuery transcends traditional data warehouse solutions with the ability to offer business insights into data across 3 dimensions -- historical, real-time and predictive. BigQuery democratizes machine learning by letting users develop and operationalize ML models with just SQL skills.
Machine Learning 103: Loss Functions
In two previous articles I covered two of the most basic models used in machine learning -- linear regression and logistic regression. In both cases, we were interested in searching for the set of model parameters m that result in the best model predictions d' of the observed targets d, and in both cases this was done by minimizing some loss function L(m), which measures the error between d' and d. A good proportion of machine learning -- from simple linear regression to deep learning models, essentially involves the minimization of some sort of loss function -- and yet, many data science or machine learning books/tutorials/materials tend to place more emphasis on the model itself than on the loss function! In this article, we will continue on where we left off from the previous two articles and focus on loss functions before exploring more advanced models in future articles! Now, just as "best" is a very subjective word, so are loss functions!
Computing Graph Edit Distance with Algorithms on Quantum Devices
Incudini, Massimiliano, Tarocco, Fabio, Mengoni, Riccardo, Di Pierro, Alessandra, Mandarino, Antonio
Distance measures provide the foundation for many popular algorithms in Machine Learning and Pattern Recognition. Different notions of distance can be used depending on the types of the data the algorithm is working on. For graph-shaped data, an important notion is the Graph Edit Distance (GED) that measures the degree of (dis)similarity between two graphs in terms of the operations needed to make them identical. As the complexity of computing GED is the same as NP-hard problems, it is reasonable to consider approximate solutions. In this paper we present a QUBO formulation of the GED problem. This allows us to implement two different approaches, namely quantum annealing and variational quantum algorithms that run on the two types of quantum hardware currently available: quantum annealer and gate-based quantum computer, respectively. Considering the current state of noisy intermediate-scale quantum computers, we base our study on proof-of-principle tests of their performance.
Robust SVM Optimization in Banach spaces
Sbihi, Mohammed, Couellan, Nicolas
We address the issue of binary classification in Banach spaces in presence of uncertainty. We show that a number of results from classical support vector machines theory can be appropriately generalised to their robust counterpart in Banach spaces. These include the Representer Theorem, strong duality for the associated Optimization problem as well as their geometric interpretation. Furthermore, we propose a game theoretic interpretation by expressing a Nash equilibrium problem formulation for the more general problem of finding the closest points in two closed convex sets when the underlying space is reflexive and smooth.
Sampling Approximately Low-Rank Ising Models: MCMC meets Variational Methods
Koehler, Frederic, Lee, Holden, Risteski, Andrej
We consider Ising models on the hypercube with a general interaction matrix $J$, and give a polynomial time sampling algorithm when all but $O(1)$ eigenvalues of $J$ lie in an interval of length one, a situation which occurs in many models of interest. This was previously known for the Glauber dynamics when *all* eigenvalues fit in an interval of length one; however, a single outlier can force the Glauber dynamics to mix torpidly. Our general result implies the first polynomial time sampling algorithms for low-rank Ising models such as Hopfield networks with a fixed number of patterns and Bayesian clustering models with low-dimensional contexts, and greatly improves the polynomial time sampling regime for the antiferromagnetic/ferromagnetic Ising model with inconsistent field on expander graphs. It also improves on previous approximation algorithm results based on the naive mean-field approximation in variational methods and statistical physics. Our approach is based on a new fusion of ideas from the MCMC and variational inference worlds. As part of our algorithm, we define a new nonconvex variational problem which allows us to sample from an exponential reweighting of a distribution by a negative definite quadratic form, and show how to make this procedure provably efficient using stochastic gradient descent. On top of this, we construct a new simulated tempering chain (on an extended state space arising from the Hubbard-Stratonovich transform) which overcomes the obstacle posed by large positive eigenvalues, and combine it with the SGD-based sampler to solve the full problem.
Low-rank features based double transformation matrices learning for image classification
Cai, Yu-Hong, Wu, Xiao-Jun, Chen, Zhe
Linear regression is a supervised method that has been widely used in classification tasks. In order to apply linear regression to classification tasks, a technique for relaxing regression targets was proposed. However, methods based on this technique ignore the pressure on a single transformation matrix due to the complex information contained in the data. A single transformation matrix in this case is too strict to provide a flexible projection, thus it is necessary to adopt relaxation on transformation matrix. This paper proposes a double transformation matrices learning method based on latent low-rank feature extraction. The core idea is to use double transformation matrices for relaxation, and jointly projecting the learned principal and salient features from two directions into the label space, which can share the pressure of a single transformation matrix. Firstly, the low-rank features are learned by the latent low rank representation (LatLRR) method which processes the original data from two directions. In this process, sparse noise is also separated, which alleviates its interference on projection learning to some extent. Then, two transformation matrices are introduced to process the two features separately, and the information useful for the classification is extracted. Finally, the two transformation matrices can be easily obtained by alternate optimization methods. Through such processing, even when a large amount of redundant information is contained in samples, our method can also obtain projection results that are easy to classify. Experiments on multiple data sets demonstrate the effectiveness of our approach for classification, especially for complex scenarios.
A hypothesis-driven method based on machine learning for neuroimaging data analysis
Gorriz, JM, Martin-Clemente, R., Puntonet, C. G., Ortiz, A., Ramirez, J., Suckling, J.
There remains an open question about the usefulness and the interpretation of Machine learning (MLE) approaches for discrimination of spatial patterns of brain images between samples or activation states. In the last few decades, these approaches have limited their operation to feature extraction and linear classification tasks for between-group inference. In this context, statistical inference is assessed by randomly permuting image labels or by the use of random effect models that consider between-subject variability. These multivariate MLE-based statistical pipelines, whilst potentially more effective for detecting activations than hypotheses-driven methods, have lost their mathematical elegance, ease of interpretation, and spatial localization of the ubiquitous General linear Model (GLM). Recently, the estimation of the conventional GLM has been demonstrated to be connected to an univariate classification task when the design matrix is expressed as a binary indicator matrix. In this paper we explore the complete connection between the univariate GLM and MLE \emph{regressions}. To this purpose we derive a refined statistical test with the GLM based on the parameters obtained by a linear Support Vector Regression (SVR) in the \emph{inverse} problem (SVR-iGLM). Subsequently, random field theory (RFT) is employed for assessing statistical significance following a conventional GLM benchmark. Experimental results demonstrate how parameter estimations derived from each model (mainly GLM and SVR) result in different experimental design estimates that are significantly related to the predefined functional task. Moreover, using real data from a multisite initiative the proposed MLE-based inference demonstrates statistical power and the control of false positives, outperforming the regular GLM.
Introduction to Probabilistic Classification: A Machine Learning Perspective
You are capable of training and evaluating classification models, both linear and non-linear model structures. Now, you want class probabilities instead of class labels. This is the article you are looking for. This article walks you through the different evaluation metrics, its pros and cons and optimal model training for multiple ML models. Imagine creating a model with the sole purpose of classifying cats and dogs.