Statistical Learning
Supervised Contrastive Learning
Khosla, Prannay, Teterwak, Piotr, Wang, Chen, Sarna, Aaron, Tian, Yonglong, Isola, Phillip, Maschinot, Aaron, Liu, Ce, Krishnan, Dilip
Cross entropy is the most widely used loss function for supervised training of image classification models. In this paper, we propose a novel training methodology that consistently outperforms cross entropy on supervised learning tasks across different architectures and data augmentations. We modify the batch contrastive loss, which has recently been shown to be very effective at learning powerful representations in the self-supervised setting. We are thus able to leverage label information more effectively than cross entropy. Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes. In addition to this, we leverage key ingredients such as large batch sizes and normalized embeddings, which have been shown to benefit self-supervised learning. On both ResNet-50 and ResNet-200, we outperform cross entropy by over 1%, setting a new state of the art number of 78.8% among methods that use AutoAugment data augmentation. The loss also shows clear benefits for robustness to natural corruptions on standard benchmarks on both calibration and accuracy. Compared to cross entropy, our supervised contrastive loss is more stable to hyperparameter settings such as optimizers or data augmentations.
A Complete Characterization of Projectivity for Statistical Relational Models
Jaeger, Manfred, Schulte, Oliver
A generative probabilistic model for relational data consists of a family of probability distributions for relational structures over domains of different sizes. In most existing statistical relational learning (SRL) frameworks, these models are not projective in the sense that the marginal of the distribution for size-$n$ structures on induced sub-structures of size $k
Sparse Generalized Canonical Correlation Analysis: Distributed Alternating Iteration based Approach
Cai, Jia, Lv, Kexin, Huo, Junyi, Huang, Xiaolin, Yang, Jie
Sparse canonical correlation analysis (CCA) is a useful statistical tool to detect latent information with sparse structures. However, sparse CCA works only for two datasets, i.e., there are only two views or two distinct objects. To overcome this limitation, in this paper, we propose a sparse generalized canonical correlation analysis (GCCA), which could detect the latent relations of multiview data with sparse structures. Moreover, the introduced sparsity could be considered as Laplace prior on the canonical variates. Specifically, we convert the GCCA into a linear system of equations and impose $\ell_1$ minimization penalty for sparsity pursuit. This results in a nonconvex problem on Stiefel manifold, which is difficult to solve. Motivated by Boyd's consensus problem, an algorithm based on distributed alternating iteration approach is developed and theoretical consistency analysis is investigated elaborately under mild conditions. Experiments on several synthetic and real world datasets demonstrate the effectiveness of the proposed algorithm.
Machine Learning using C for Linear and Logistic Regression
The applications of machine learning transcend boundaries and industries so why should we let tools and languages hold us back? Yes, Python is the language of choice in the industry right now but a lot of us come from a background where Python isn't taught! The computer science faculty in universities are still teaching programming in C – so that's what most of us end up learning first. I understand why you should learn Python – it's the primary language in the industry and it has all the libraries you need to get started with machine learning. But what if your university doesn't teach it?
Model-based targeted dimensionality reduction for neuronal population data
Aoi, Mikio, Pillow, Jonathan W.
Summarizing high-dimensional data using a small number of parameters is a ubiquitous first step in the analysis of neuronal population activity. Recently developed methods use "targeted" approaches that work by identifying multiple, distinct low-dimensional subspaces of activity that capture the population response to individual experimental task variables, such as the value of a presented stimulus or the behavior of the animal. These methods have gained attention because they decompose total neural activity into what are ostensibly different parts of a neuronal computation. However, existing targeted methods have been developed outside of the confines of probabilistic modeling, making some aspects of the procedures ad hoc, or limited in flexibility or interpretability. Here we propose a new model-based method for targeted dimensionality reduction based on a probabilistic generative model of the population response data.
Learning Sampling and Model-Based Signal Recovery for Compressed Sensing MRI
Huijben, Iris A. M., Veeling, Bastiaan S., van Sloun, Ruud J. G.
Compressed sensing (CS) MRI relies on adequate undersampling of the k-space to accelerate the acquisition without compromising image quality. Consequently, the design of optimal sampling patterns for these k-space coefficients has received significant attention, with many CS MRI methods exploiting variable-density probability distributions. Realizing that an optimal sampling pattern may depend on the downstream task (e.g. image reconstruction, segmentation, or classification), we here propose joint learning of both task-adaptive k-space sampling and a subsequent model-based proximal-gradient recovery network. The former is enabled through a probabilistic generative model that leverages the Gumbel-softmax relaxation to sample across trainable beliefs while maintaining differentiability. The proposed combination of a highly flexible sampling model and a model-based (sampling-adaptive) image reconstruction network facilitates exploration and efficient training, yielding improved MR image quality compared to other sampling baselines.
Active Learning for Gaussian Process Considering Uncertainties with Application to Shape Control of Composite Fuselage
Yue, Xiaowei, Wen, Yuchen, Hunt, Jeffrey H., Shi, Jianjun
This paper has been accepted by IEEE Transactions on Automation Science and Engineering. 1 This preprint is an accepted version, not the IEEE published version. Abstract--In the machine learning domain, active learning is an iterative data selection algorithm for maximizing information acquisition and improving model performance with limited training samples. It is very useful, especially for the industrial applications where training samples are expensive, time-consuming, or difficult to obtain. Existing methods mainly focus on active learning for classification, and a few methods are designed for regression such as linear regression or Gaussian process. Uncertainties from measurement errors and intrinsic input noise inevitably exist in the experimental data, which further affects the modeling performance. The existing active learning methods do not incorporate these uncertainties for Gaussian process. In this paper, we propose two new active learning algorithms for the Gaussian process with uncertainties, which are variance-based weighted active learning algorithm and D-optimal weighted active learning algorithm. Through numerical study, we show that the proposed approach can incorporate the impact from uncertainties, and realize better prediction performance. This approach has been applied to improving the predictive modeling for automatic shape control of composite fuselage. I. INTRODUCTION Active learning is a type of iterative supervised learning which focuses on maximizing information acquisition with limited samples. In statistics literature, this process is also called optimal experimental design, or sequential design. The main idea of active learning is to iteratively pose "query" or "design" to explore the most informative new experimental samples according to the information obtained from the current samples. In many machine learning applications, especially in some industrial systems, the explanatory data are rich and easy to get, but the response data are very expensive, time-consuming, or difficult to obtain. For example, when training autonomous driving algorithms, a lot of media (e.g., images, videos) require that oracle users mark them with particular labels, such as "vehicle", "street sign" or "road lines". It can be tedious, redundant and time-consuming to annotate lots of these instances.
Doubly-stochastic mining for heterogeneous retrieval
Rawat, Ankit Singh, Menon, Aditya Krishna, Veit, Andreas, Yu, Felix, Reddi, Sashank J., Kumar, Sanjiv
Information retrieval concerns finding documents that are most relevant for a given query, and is a canonical real-world use case for machine learning [Manning et al., 2008]. The simplest incarnation of retrieval models involves learning a real-valued scoring function that ranks, for each example, the set of possible labels it may be matched to. A core challenge is scalability: there may be billions of examples (e.g., user queries) and labels (e.g., videos in a recommendation system), each of whose scores naïvely needs to be updated at every training iteration. Effective means of addressing both problems have been widely studied [Mikolov et al., 2013, Jean et al., 2015, Reddi et al., 2019]. A distinct challenge is heterogeneity: the distribution over examples is often a mixture of diverse subpopulations (e.g., queries may arise from geographically disparate user bases). Naïve training on such data may lead to models that perform disproportionately well on one subpopulation at the expense of others; e.g., if queries originate from multiple countries, the retrieval model may only perform well on queries from the dominant country. Such behaviour is clearly undesirable.
Alternating Minimization Converges Super-Linearly for Mixed Linear Regression
Ghosh, Avishek, Ramchandran, Kannan
We address the problem of solving mixed random linear equations. We have unlabeled observations coming from multiple linear regressions, and each observation corresponds to exactly one of the regression models. The goal is to learn the linear regressors from the observations. Classically, Alternating Minimization (AM) (which is a variant of Expectation Maximization (EM)) is used to solve this problem. AM iteratively alternates between the estimation of labels and solving the regression problems with the estimated labels. Empirically, it is observed that, for a large variety of non-convex problems including mixed linear regression, AM converges at a much faster rate compared to gradient based algorithms. However, the existing theory suggests similar rate of convergence for AM and gradient based methods, failing to capture this empirical behavior. In this paper, we close this gap between theory and practice for the special case of a mixture of $2$ linear regressions. We show that, provided initialized properly, AM enjoys a \emph{super-linear} rate of convergence in certain parameter regimes. To the best of our knowledge, this is the first work that theoretically establishes such rate for AM. Hence, if we want to recover the unknown regressors upto an error (in $\ell_2$ norm) of $\epsilon$, AM only takes $\mathcal{O}(\log \log (1/\epsilon))$ iterations. Furthermore, we compare AM with a gradient based heuristic algorithm empirically and show that AM dominates in iteration complexity as well as wall-clock time.
Adversarial examples and where to find them
Risse, Niklas, Göpfert, Christina, Göpfert, Jan Philip
Adversarial robustness of trained models has attracted considerable attention over recent years, within and beyond the scientific community. This is not only because of a straight-forward desire to deploy reliable systems, but also because of how adversarial attacks challenge our beliefs about deep neural networks. Demanding more robust models seems to be the obvious solution -- however, this requires a rigorous understanding of how one should judge adversarial robustness as a property of a given model. In this work, we analyze where adversarial examples occur, in which ways they are peculiar, and how they are processed by robust models. We use robustness curves to show that $\ell_\infty$ threat models are surprisingly effective in improving robustness for other $\ell_p$ norms; we introduce perturbation cost trajectories to provide a broad perspective on how robust and non-robust networks perceive adversarial perturbations as opposed to random perturbations; and we explicitly examine the scale of certain common data sets, showing that robustness thresholds must be adapted to the data set they pertain to. This allows us to provide concrete recommendations for anyone looking to train a robust model or to estimate how much robustness they should require for their operation. The code for all our experiments is available at www.github.com/niklasrisse/adversarial-examples-and-where-to-find-them .