Statistical Learning
Generalizing to Unseen Domains: A Survey on Domain Generalization
Wang, Jindong, Lan, Cuiling, Liu, Chang, Ouyang, Yidong, Qin, Tao
Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increased interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. For years, great progress has been achieved. This paper presents the first review for recent advances in domain generalization. First, we provide a formal definition of domain generalization and discuss several related fields. Next, we thoroughly review the theories related to domain generalization and carefully analyze the theory behind generalization. Then, we categorize recent algorithms into three classes and present them in detail: data manipulation, representation learning, and learning strategy, each of which contains several popular algorithms. Third, we introduce the commonly used datasets and applications. Finally, we summarize existing literature and present some potential research topics for the future.
Machine Learning Basics Course for Beginners in 3 Hours
Please watch: "Mask R CNN Implementation How to Install and Run using TensorFlow 2.0 2021" https://www.youtube.com/watch?v tcu4pr948n0 -- --Welcome to th... Welcome to the Fun and Easy machine learning concepts FULL COURSE in 3 hours, where you will be learning popular theoretical topics in: Machine Learning, Neural Networks, and Computer Vision. This course is designed to be simple and fun without all of complex math and boring explanations. Each theoretical lecture is crafted using whiteboard animations like this one which maximizes concentration, and knowledge retention making you feel like a Machine Learning expert once you have completed this free training.
Fresh Out Of College & No Experience? Here's How To Get An AI Job
AI is currently booming and with the kind of advancements that are happening, there is no stopping. Every industry, from manufacturing, retail, pharmaceuticals, to healthcare and finance, uses AI and machine learning (a subset of AI) tools to automate mundane tasks, sift through several GBs of data to make an accurate business decision, and improve customer service amongst many other tasks. So while this is all true and evident, everyone wants to hop on the train that is artificial intelligence. Artificial Intelligence is a unique field. As a young graduate, you might not have any solid information about the field and no relevant experience. A few modules in colleges will not be much of help, and employers have a hard time looking for candidates with relevant experience.
Machine Learning a Systems Engineering Perspective
Systems engineering seeks to understand the big picture by breaking complex projects into manageable well-defined sub-systems. This article will leverage fundamental systems engineering principles to introduce Machine Learning as a system composed of interacting elements. The usage of terminology throughout this article is an elaboration of the fundamental idea that a system is a purposeful whole consisting of interacting parts. Each element that is part of these system is atomic (i.e., not further decomposable) in nature and modeled using descriptive features. This reduces the complexity and supports task independence by allowing management authorities to get a high-level perspective of a machine learning pipeline.
Continual Density Ratio Estimation in an Online Setting
Chen, Yu, Liu, Song, Diethe, Tom, Flach, Peter
In online applications with streaming data, awareness of how far the training or test set has shifted away from the original dataset can be crucial to the performance of the model. However, we may not have access to historical samples in the data stream. To cope with such situations, we propose a novel method, Continual Density Ratio Estimation (CDRE), for estimating density ratios between the initial and current distributions ($p/q_t$) of a data stream in an iterative fashion without the need of storing past samples, where $q_t$ is shifting away from $p$ over time $t$. We demonstrate that CDRE can be more accurate than standard DRE in terms of estimating divergences between distributions, despite not requiring samples from the original distribution. CDRE can be applied in scenarios of online learning, such as importance weighted covariate shift, tracing dataset changes for better decision making. In addition, (CDRE) enables the evaluation of generative models under the setting of continual learning. To the best of our knowledge, there is no existing method that can evaluate generative models in continual learning without storing samples from the original distribution.
BIKED: A Dataset and Machine Learning Benchmarks for Data-Driven Bicycle Design
Regenwetter, Lyle, Curry, Brent, Ahmed, Faez
In this paper, we present "BIKED," a dataset comprised of 4500 individually designed bicycle models sourced from hundreds of designers. We expect BIKED to enable a variety of data-driven design applications for bicycles and generally support the development of data-driven design methods. The dataset is comprised of a variety of design information including assembly images, component images, numerical design parameters, and class labels. In this paper, we first discuss the processing of the dataset and present the various features provided. We then illustrate the scale, variety, and structure of the data using several unsupervised clustering studies. Next, we explore a variety of data-driven applications. We provide baseline classification performance for 10 algorithms trained on differing amounts of training data. We then contrast classification performance of three deep neural networks using parametric data, image data, and a combination of the two. Using one of the trained classification models, we conduct a Shapley Additive Explanations Analysis to better understand the extent to which certain design parameters impact classification predictions. Next, we test bike reconstruction and design synthesis using two Variational Autoencoders (VAEs) trained on images and parametric data. We furthermore contrast the performance of interpolation and extrapolation tasks in the original parameter space and the latent space of a VAE. Finally, we discuss some exciting possibilities for other applications beyond the few actively explored in this paper and summarize overall strengths and weaknesses of the dataset.
Interpretable Machines: Constructing Valid Prediction Intervals with Random Forests
An important issue when using Machine Learning algorithms in recent research is the lack of interpretability. Although these algorithms provide accurate point predictions for various learning problems, uncertainty estimates connected with point predictions are rather sparse. A contribution to this gap for the Random Forest Regression Learner is presented here. Based on its Out-of-Bag procedure, several parametric and non-parametric prediction intervals are provided for Random Forest point predictions and theoretical guarantees for its correct coverage probability is delivered. In a second part, a thorough investigation through Monte-Carlo simulation is conducted evaluating the performance of the proposed methods from three aspects: (i) Analyzing the correct coverage rate of the proposed prediction intervals, (ii) Inspecting interval width and (iii) Verifying the competitiveness of the proposed intervals with existing methods. The simulation yields that the proposed prediction intervals are robust towards non-normal residual distributions and are competitive by providing correct coverage rates and comparably narrow interval lengths, even for comparably small samples.
Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds
Feng, Yihao, Tang, Ziyang, Zhang, Na, Liu, Qiang
Off-policy evaluation (OPE) is the task of estimating the expected reward of a given policy based on offline data previously collected under different policies. Therefore, OPE is a key step in applying reinforcement learning to real-world domains such as medical treatment, where interactive data collection is expensive or even unsafe. As the observed data tends to be noisy and limited, it is essential to provide rigorous uncertainty quantification, not just a point estimation, when applying OPE to make high stakes decisions. This work considers the problem of constructing non-asymptotic confidence intervals in infinite-horizon off-policy evaluation, which remains a challenging open question. We develop a practical algorithm through a primal-dual optimization-based approach, which leverages the kernel Bellman loss (KBL) of Feng et al.(2019) and a new martingale concentration inequality of KBL applicable to time-dependent data with unknown mixing conditions. Our algorithm makes minimum assumptions on the data and the function class of the Q-function, and works for the behavior-agnostic settings where the data is collected under a mix of arbitrary unknown behavior policies. We present empirical results that clearly demonstrate the advantages of our approach over existing methods.
More data or more parameters? Investigating the effect of data structure on generalization
d'Ascoli, Stรฉphane, Gabriรฉ, Marylou, Sagun, Levent, Biroli, Giulio
One of the central features of deep learning is the generalization abilities of neural networks, which seem to improve relentlessly with over-parametrization. In this work, we investigate how properties of data impact the test error as a function of the number of training examples and number of training parameters; in other words, how the structure of data shapes the "generalization phase space". We first focus on the random features model trained in the teacher-student scenario. The synthetic input data is composed of independent blocks, which allow us to tune the saliency of low-dimensional structures and their relevance with respect to the target function. Using methods from statistical physics, we obtain an analytical expression for the train and test errors for both regression and classification tasks in the high-dimensional limit. The derivation allows us to show that noise in the labels and strong anisotropy of the input data play similar roles on the test error. Both promote an asymmetry of the phase space where increasing the number of training examples improves generalization further than increasing the number of training parameters. Our analytical insights are confirmed by numerical experiments involving fully-connected networks trained on MNIST and CIFAR10.
Active Testing: Sample-Efficient Model Evaluation
Kossen, Jannik, Farquhar, Sebastian, Gal, Yarin, Rainforth, Tom
We introduce active testing: a new framework for sample-efficient model evaluation. While approaches like active learning reduce the number of labels needed for model training, existing literature largely ignores the cost of labeling test data, typically unrealistically assuming large test sets for model evaluation. This creates a disconnect to real applications where test labels are important and just as expensive, e.g. for optimizing hyperparameters. Active testing addresses this by carefully selecting the test points to label, ensuring model evaluation is sample-efficient. To this end, we derive theoretically-grounded and intuitive acquisition strategies that are specifically tailored to the goals of active testing, noting these are distinct to those of active learning. Actively selecting labels introduces a bias; we show how to remove that bias while reducing the variance of the estimator at the same time. Active testing is easy to implement, effective, and can be applied to any supervised machine learning method. We demonstrate this on models including WideResNet and Gaussian processes on datasets including CIFAR-100.