Statistical Learning
On the Use of Minimum Penalties in Statistical Learning
Sherwood, Ben, Price, Bradley S.
Modern multivariate machine learning and statistical methodologies estimate parameters of interest while leveraging prior knowledge of the association between outcome variables. The methods that do allow for estimation of relationships do so typically through an error covariance matrix in multivariate regression which does not scale to other types of models. In this article we proposed the MinPEN framework to simultaneously estimate regression coefficients associated with the multivariate regression model and the relationships between outcome variables using mild assumptions. The MinPen framework utilizes a novel penalty based on the minimum function to exploit detected relationships between responses. An iterative algorithm that generalizes current state of the art methods is proposed as a solution to the non-convex optimization that is required to obtain estimates. Theoretical results such as high dimensional convergence rates, model selection consistency, and a framework for post selection inference are provided. We extend the proposed MinPen framework to other exponential family loss functions, with a specific focus on multiple binomial responses. Tuning parameter selection is also addressed. Finally, simulations and two data examples are presented to show the finite sample properties of this framework.
Gaussian Mixture Estimation from Weighted Samples
Frisch, Daniel, Hanebeck, Uwe D.
Given a set of samples, the parameters of a GM are determined in such a way as to best fit the samples in a maximum likelihood way. Solutions for equally weighted samples are readily available, expectation-maximization (EM) based methods being the most prevalent because of low computational requirements and ease of implementation. So it comes as a surprise that GM estimation for weighted samples is hard to find in literature. It might be even more surprising that the standard reference [1] gives incorrect results, see Figure 1. 2. Context Applications for sample-to-density function approximation include clustering of unlabled data [2, 3], multi-target tracking [4, 5], group tracking [6], multilateration [7, 8], and arbitrary density representation in nonlinear filters [9, 10]. A popular basic solution to this is the k-means algorithm. It does not find a complete density representation, only the means of the individual clusters. The k-means algorithm uses hard sample-tomean associations, therefore yields merely approximate solutions but can be computationally optimized using k-d trees [11, 12]. Moreover, the global optimum can be found deterministically [13], therefore it can be used to provide an initial guess for more elaborate algorithms. A sample-to-density approximation that is optimal in a maximum likelihood sense can be searched with numerical optimization techniques such as the Newton algorithm that has quadratic convergence but high computational demand per iteration, quasi-Newton methods, the method of scoring, or the conjugate gradient method with slower convergence but less computational effort per iteration [14].
Polynomial magic! Hermite polynomials for private data generation
Park, Mijung, Vinaroz, Margarita, Charusaie, Mohammad-Amin, Harder, Frederik
Kernel mean embedding is a useful tool to compare probability measures. Despite its usefulness, kernel mean embedding considers infinite-dimensional features, which are challenging to handle in the context of differentially private data generation. A recent work [13] proposes to approximate the kernel mean embedding of data distribution using finite-dimensional random features, where the sensitivity of the features becomes analytically tractable. More importantly, this approach significantly reduces the privacy cost, compared to other known privatization methods (e.g., DP-SGD), as the approximate kernel mean embedding of the data distribution is privatized only once and can then be repeatedly used during training of a generator without incurring any further privacy cost. However, the required number of random features is excessively high, often ten thousand to a hundred thousand, which worsens the sensitivity of the approximate kernel mean embedding. To improve the sensitivity, we propose to replace random features with Hermite polynomial features. Unlike the random features, the Hermite polynomial features are ordered, where the features at the low orders contain more information on the distribution than those at the high orders. Hence, a relatively low order of Hermite polynomial features can more accurately approximate the mean embedding of the data distribution compared to a significantly higher number of random features. As a result, using the Hermite polynomial features, we significantly improve the privacy-accuracy trade-off, reflected in the high quality and diversity of the generated data, when tested on several heterogeneous tabular datasets, as well as several image benchmark datasets.
Fully differentiable model discovery
Model discovery aims at autonomously discovering differential equations underlying a dataset. Approaches based on Physics Informed Neural Networks (PINNs) have shown great promise, but a fully-differentiable model which explicitly learns the equation has remained elusive. In this paper we propose such an approach by combining neural network based surrogates with Sparse Bayesian Learning (SBL). We start by reinterpreting PINNs as multitask models, applying multitask learning using uncertainty, and show that this leads to a natural framework for including Bayesian regression techniques. We then construct a robust model discovery algorithm by using SBL, which we showcase on various datasets. Concurrently, the multitask approach allows the use of probabilistic approximators, and we show a proof of concept using normalizing flows to directly learn a density model from single particle data. Our work expands PINNs to various types of neural network architectures, and connects neural network-based surrogates to the rich field of Bayesian parameter inference.
Bayesian Boosting for Linear Mixed Models
Zhang, Boyao, Griesbach, Colin, Kim, Cora, Müller-Voggel, Nadia, Bergherr, Elisabeth
Linear mixed models (LMM) (Laird and Ware, 1982) are widely used in longitudinal data analysis as they incorporate random effects to deal with group-specific heterogeneity. Data involving repeated observations of the same variables are common in epidemiology, medical statistics and many other fields. Likelihood-based methods are often used to make inference for (generalized) linear mixed models (Bates et al., 2000; Gumedze and Dunne, 2011). Schelldorfer et al. (2011) and Groll and Tutz (2014) introduced separately the L1-penalized estimation for high-dimensional linear mixed models. Fong et al. (2010) argued that for small sample sizes likelihood-based inference can be unreliable with variance components being difficult to estimate and suggested to use the Bayesian method. When the random effects distribution is misspecified, the resulting maximum likelihood estimators are inconsistent and biased (Neuhaus et al., 1992; Heagerty and Kurland, 2001; Litière et al., 2008). Fahrmeir and Lang (2001) presented a fully Bayesian inference via Markov Chain Monte Carlo (MCMC) simulation in generalized additive and semiparametric mixed models. Rosa et al. (2003) described a normal/independent residual distributions for robust inference and suggested also the Bayesian framework. Bayesian inference for mixed models can be conducted with for example BayesX, a program with MCMC simulation techniques (Lang and Brezger, 2000).
DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
Hazimeh, Hussein, Zhao, Zhe, Chowdhery, Aakanksha, Sathiamoorthy, Maheswaran, Chen, Yihua, Mazumder, Rahul, Hong, Lichan, Chi, Ed H.
The Mixture-of-experts (MoE) architecture is showing promising results in multi-task learning (MTL) and in scaling high-capacity neural networks. State-of-the-art MoE models use a trainable sparse gate to select a subset of the experts for each input example. While conceptually appealing, existing sparse gates, such as Top-k, are not smooth. The lack of smoothness can lead to convergence and statistical performance issues when training with gradient-based methods. In this paper, we develop DSelect-k: the first, continuously differentiable and sparse gate for MoE, based on a novel binary encoding formulation. Our gate can be trained using first-order methods, such as stochastic gradient descent, and offers explicit control over the number of experts to select. We demonstrate the effectiveness of DSelect-k in the context of MTL, on both synthetic and real datasets with up to 128 tasks. Our experiments indicate that MoE models based on DSelect-k can achieve statistically significant improvements in predictive and expert selection performance. Notably, on a real-world large-scale recommender system, DSelect-k achieves over 22% average improvement in predictive performance compared to the Top-k gate. We provide an open-source TensorFlow implementation of our gate.
Bias-Robust Bayesian Optimization via Dueling Bandits
Kirschner, Johannes, Krause, Andreas
We consider Bayesian optimization in settings where observations can be adversarially biased, for example by an uncontrolled hidden confounder. Our first contribution is a reduction of the confounded setting to the dueling bandit model. Then we propose a novel approach for dueling bandits based on information-directed sampling (IDS). Thereby, we obtain the first efficient kernelized algorithm for dueling bandits that comes with cumulative regret guarantees. Our analysis further generalizes a previously proposed semi-parametric linear bandit model to non-linear reward functions, and uncovers interesting links to doubly-robust estimation.
Yes, XGBoost is cool, but have you heard of CatBoost?
If you've worked as a data scientist, competed in Kaggle competitions, or even browsed data science articles on the internet, there's a high chance that you've heard of XGBoost. Even today, it is often the go-to algorithm for many Kagglers and data scientists working on general machine learning tasks. While XGBoost is popular for good reasons, it does have some limitations, which I mentioned in my article below. Odds are, you've probably heard of XGBoost, have you ever heard of CatBoost? CatBoost is another open-source gradient boosting library that was created by researchers at Yandex.
Lasso (l1) and Ridge (l2) Regularization Techniques
What is the need for Ridge and Lasso Regression? When we create our linear model with the best-fitted line and come on testing phase then because of increased variation, our model is over-fitted, So It will not work well in the future also not provide appropriate accuracy. Therefore, to reduce overfitting, ridge and lasso regression came into the picture. Both are powerful techniques with a slight difference used for creating such models that are efficient and computationally fit to reduce over-fitting. It is a process to classify the classes and provide additional information to prevent over-fitting.
A 2020 taxonomy of algorithms inspired on living beings behavior
Since the emerge of ideas about simulation of life in last decades, several algorithms have been proposed to solve complex problems inspired on nature phenomena; i.e. evolutionary computation or artificial life. A role of a naturalist or biologist is taken with the purpose for studying all living forms in a new ecosystem and trying to make a classification of all discoveries to form a taxonomy of living beings. This role is taken as a computer naturalist to make a compilation of algorithms inspired on behavior of living beings. There are several bio-inspired algorithms; however, this work focus on actions of living beings like the growth of plants, reproduction of mushrooms, living of bacteria, the individuals behavior of animals, etc.; however, highlights the interactions between individuals of a group of different animals like school of fishes, flock of birds, herd of mammals, or swarm of insects. Focusing on algorithms inspired in actions of living beings that belongs to any kingdom of the nature; nevertheless, it is important to locate all algorithms as possible. Only basic algorithms are considered, but derivations, variants and hybrids are omitted; at least, algorithms which involves an inspiration of any living being. Location of bio-inspired algorithms related with a specific species is made by a review of several papers of surveys which involve nature bio-inspired, swarm intelligence, and metaheuristics algorithms; however, several of these surveys consider different points of view. It was consider only survey papers from ten years old ago because it is expected a more complete reviews since then. Surveys span in many cases all kind of algorithms; however many of them have been proposed recently; it maybe because the year 2020 is iconic.