Goto

Collaborating Authors

 Statistical Learning


The 4 Machine Learning Models Imperative for Business Transformation

#artificialintelligence

Machine learning is hot right now, and for good reason. We're going to break down what you need to know about what goes into a model and give you four machine learning models your business should have in production right now. The Lead/Opportunity Conversions Model The lifeblood of every business is new leads and opportunities. Having a machine learning model in place to predict where you're more likely to convert those leads can be an effective guide to growth. The Attrition/Customer Retention Model Once you have a customer in your ecosystem, it's in your best interest to keep that customer for the long haul. The attrition/customer retention model can tell you who has a high propensity to churn, so you can market to your existing base effectively. The Lifetime Value Model Increasing the lifetime value of your customers or clients is critical. Having a model in place that offers behavior-driven insight will help you keep your customers in your pipeline longer.


Why Is The K-Nearest Neighbors (KNN) Called A "Lazy Algorithm"?

#artificialintelligence

Understanding why K-Nearest Neighbors is called a "lazy learner" K-Nearest Neighbors or KNN is one of the simplest machine learning algorithms. This algorithm is very easy to implement and equally easy to understand. It is a supervised machine learning algorithm. This means we need a reference dataset to predict the category/group of the future data point. Once the KNN model is trained on the reference dataset, it classifies the new data points based on the points (neighbors) that are most similar to it.


Empirical Evaluation of Biased Methods for Alpha Divergence Minimization

arXiv.org Machine Learning

There has been great recent interest in methods to minimize other alpha-divergences, such as the "inclusive" KL divergence, KL(pโ€–q). Some methods employ unbiased gradient estimators (Dieng et al., 2017; Kuleshov and Ermon, 2017). These estimators often suffer from a high variance, difficulting optimization (Geffner and Domke, 2020). Another class of methods estimate a gradient using self-normalized importance sampling (Bornschein and Bengio, 2014; Finke and Thiery, 2019; Li and Turner, 2016). While these estimators may control variance, they do so at the cost of some bias.


Graph Consistency based Mean-Teaching for Unsupervised Domain Adaptive Person Re-Identification

arXiv.org Artificial Intelligence

Recent works show that mean-teaching is an effective framework for unsupervised domain adaptive person re-identification. However, existing methods perform contrastive learning on selected samples between teacher and student networks, which is sensitive to noises in pseudo labels and neglects the relationship among most samples. Moreover, these methods are not effective in cooperation of different teacher networks. To handle these issues, this paper proposes a Graph Consistency based Mean-Teaching (GCMT) method with constructing the Graph Consistency Constraint (GCC) between teacher and student networks. Specifically, given unlabeled training images, we apply teacher networks to extract corresponding features and further construct a teacher graph for each teacher network to describe the similarity relationships among training images. To boost the representation learning, different teacher graphs are fused to provide the supervise signal for optimizing student networks. GCMT fuses similarity relationships predicted by different teacher networks as supervision and effectively optimizes student networks with more sample relationships involved. Experiments on three datasets, i.e., Market-1501, DukeMTMCreID, and MSMT17, show that proposed GCMT outperforms state-of-the-art methods by clear margin. Specially, GCMT even outperforms the previous method that uses a deeper backbone. Experimental results also show that GCMT can effectively boost the performance with multiple teacher and student networks. Our code is available at https://github.com/liu-xb/GCMT .


Learning Gaussian Graphical Models with Latent Confounders

arXiv.org Machine Learning

Gaussian Graphical models (GGM) are widely used to estimate the network structures in many applications ranging from biology to finance. In practice, data is often corrupted by latent confounders which biases inference of the underlying true graphical structure. In this paper, we compare and contrast two strategies for inference in graphical models with latent confounders: Gaussian graphical models with latent variables (LVGGM) and PCA-based removal of confounding (PCA+GGM). While these two approaches have similar goals, they are motivated by different assumptions about confounding. In this paper, we explore the connection between these two approaches and propose a new method, which combines the strengths of these two approaches. We prove the consistency and convergence rate for the PCA-based method and use these results to provide guidance about when to use each method. We demonstrate the effectiveness of our methodology using both simulations and in two real-world applications.


Extending Models Via Gradient Boosting: An Application to Mendelian Models

arXiv.org Machine Learning

Improving existing widely-adopted prediction models is often a more efficient and robust way towards progress than training new models from scratch. Existing models may (a) incorporate complex mechanistic knowledge, (b) leverage proprietary information and, (c) have surmounted barriers to adoption. Compared to model training, model improvement and modification receive little attention. In this paper we propose a general approach to model improvement: we combine gradient boosting with any previously developed model to improve model performance while retaining important existing characteristics. To exemplify, we consider the context of Mendelian models, which estimate the probability of carrying genetic mutations that confer susceptibility to disease by using family pedigrees and health histories of family members. Via simulations we show that integration of gradient boosting with an existing Mendelian model can produce an improved model that outperforms both that model and the model built using gradient boosting alone. We illustrate the approach on genetic testing data from the USC-Stanford Cancer Genetics Hereditary Cancer Panel (HCP) study.


Bias, Fairness, and Accountability with AI and ML Algorithms

arXiv.org Machine Learning

Artificial intelligence (AI) techniques are used increasingly in many areas of applications, including banking and finance. They have several advantages over traditional statistical methods: i) ability to handle new data types such as text, audio, and images; ii) flexible models that yield excellent predictive performance; and iii) ability to automate many of the routine, and time-consuming, tasks in model development. However, the use of these algorithms also raise several challenges. A well-known problem is the opaqueness of ML models and the difficulties in understanding and interpreting the model results. In this paper, we focus on a related and equally important challenge: potential for bias and lack of fairness when using AI/ML techniques.


Improved Algorithms for Agnostic Pool-based Active Classification

arXiv.org Machine Learning

We consider active learning for binary classification in the agnostic pool-based setting. The vast majority of works in active learning in the agnostic setting are inspired by the CAL algorithm where each query is uniformly sampled from the disagreement region of the current version space. The sample complexity of such algorithms is described by a quantity known as the disagreement coefficient which captures both the geometry of the hypothesis space as well as the underlying probability space. To date, the disagreement coefficient has been justified by minimax lower bounds only, leaving the door open for superior instance dependent sample complexities. In this work we propose an algorithm that, in contrast to uniform sampling over the disagreement region, solves an experimental design problem to determine a distribution over examples from which to request labels. We show that the new approach achieves sample complexity bounds that are never worse than the best disagreement coefficient-based bounds, but in specific cases can be dramatically smaller. From a practical perspective, the proposed algorithm requires no hyperparameters to tune (e.g., to control the aggressiveness of sampling), and is computationally efficient by means of assuming access to an empirical risk minimization oracle (without any constraints). Empirically, we demonstrate that our algorithm is superior to state of the art agnostic active learning algorithms on image classification datasets.


Policy Optimization in Bayesian Network Hybrid Models of Biomanufacturing Processes

arXiv.org Artificial Intelligence

Biopharmaceutical manufacturing is a rapidly growing industry with impact in virtually all branches of medicine. Biomanufacturing processes require close monitoring and control, in the presence of complex bioprocess dynamics with many interdependent factors, as well as extremely limited data due to the high cost and long duration of experiments. We develop a novel model-based reinforcement learning framework that can achieve human-level control in low-data environments. The model uses a probabilistic knowledge graph to capture causal interdependencies between factors in the underlying stochastic decision process, leveraging information from existing kinetic models from different unit operations while incorporating real-world experimental data. We then present a computationally efficient, provably convergent stochastic gradient method for policy optimization. Validation is conducted on a realistic application with a multi-dimensional, continuous state variable.


A rigorous introduction for linear models

arXiv.org Machine Learning

This note is meant to provide an introduction to linear models and the theories behind them. Our goal is to give a rigorous introduction to the readers with prior exposure to ordinary least squares. In machine learning, the output is usually a nonlinear function of the input. Deep learning even aims to find a nonlinear dependence with many layers which require a large amount of computation. However, most of these algorithms build upon simple linear models. We then describe linear models from different views and find the properties and theories behind the models. The linear model is the main technique in regression problems and the primary tool for it is the least squares approximation which minimizes a sum of squared errors. This is a natural choice when we're interested in finding the regression function which minimizes the corresponding expected squared error. We first describe ordinary least squares from three different points of view upon which we disturb the model with random noise and Gaussian noise. By Gaussian noise, the model gives rise to the likelihood so that we introduce a maximum likelihood estimator. It also develops some distribution theories for it via this Gaussian disturbance. The distribution theory of least squares will help us answer various questions and introduce related applications. We then prove least squares is the best unbiased linear model in the sense of mean squared error and most importantly, it actually approaches the theoretical limit. We end up with linear models with the Bayesian approach and beyond.