Statistical Learning
A determinantal point process for column subset selection
Belhadji, Ayoub, Bardenet, Rémi, Chainais, Pierre
Dimensionality reduction is a first step of many machine learning pipelines. Two popular approaches are principal component analysis, which projects onto a small number of well chosen but non-interpretable directions, and feature selection, which selects a small number of the original features. Feature selection can be abstracted as a numerical linear algebra problem called the column subset selection problem (CSSP). CSSP corresponds to selecting the best subset of columns of a matrix $X \in \mathbb{R}^{N \times d}$, where \emph{best} is often meant in the sense of minimizing the approximation error, i.e., the norm of the residual after projection of $X$ onto the space spanned by the selected columns. Such an optimization over subsets of $\{1,\dots,d\}$ is usually impractical. One workaround that has been vastly explored is to resort to polynomial-cost, random subset selection algorithms that favor small values of this approximation error. We propose such a randomized algorithm, based on sampling from a projection determinantal point process (DPP), a repulsive distribution over a fixed number $k$ of indices $\{1,\dots,d\}$ that favors diversity among the selected columns. We give bounds on the ratio of the expected approximation error for this DPP over the optimal error of PCA. These bounds improve over the state-of-the-art bounds of \emph{volume sampling} when some realistic structural assumptions are satisfied for $X$. Numerical experiments suggest that our bounds are tight, and that our algorithms have comparable performance with the \emph{double phase} algorithm, often considered to be the practical state-of-the-art. Column subset selection with DPPs thus inherits the best of both worlds: good empirical performance and tight error bounds.
Detecting British Columbia Coastal Rainfall Patterns by Clustering Gaussian Processes
Paton, Forrest, McNicholas, Paul D.
Functional data analysis is a statistical framework where data are assumed to follow some functional form. This method of analysis is commonly applied to time series data, where time, measured continuously or in discrete intervals, serves as the location for a function's value. Gaussian processes are a generalization of the multivariate normal distribution to function space and, in this paper, they are used to shed light on coastal rainfall patterns in British Columbia (BC). Specifically, this work addressed the question over how one should carry out an exploratory cluster analysis for the BC, or any similar, coastal rainfall data. An approach is developed for clustering multiple processes observed on a comparable interval, based on how similar their underlying covariance kernel is. This approach provides significant insights into the BC data, and these insights can be described in terms of El Nino and La Nina; however, the result is not simply one cluster representing El Nino years and another for La Nina years. From one perspective, the results show that clustering annual rainfall can potentially be used to identify extreme weather patterns.
Let Me Not Lie: Learning MultiNomial Logit
Sifringer, Brian, Lurkin, Virginie, Alahi, Alexandre
Discrete choice models generally assume that model specification is known a priori. In practice, determiningthe utility specification for a particular application remains a difficult task and model misspecification may lead to biased parameter estimates. In this paper, we propose anew mathematical framework for estimating choice models in which the systematic part of the utility specification is divided into an interpretable part and a learning representation partthat aims at automatically discovering a good utility specification from available data. We show the effectiveness of our framework by augmenting the utility specification of the Multinomial Logit Model (MNL) with a new nonlinear representation arising from a Neural Network (NN). This leads to a new choice model referred to as the Learning Multinomial Logit(L-MNL) model. Our experiments show that our L-MNL model outperformed the traditional MNL models and existing hybrid neural network models both in terms of predictive performance and accuracy in parameter estimation. Keywords: Discrete choice models, Neural networks, Utility specification 1. Introduction Discrete Choice Models (DCM) have emerged as a powerful theoretical framework for analyzing individual travel behavior. The goal of these models is to predict the choice among a given set of discrete alternatives (e.g., choice of walking as transportation mode rather than taking the car or the bus), while understanding the behavioral process that led to the specific choice. For many years, the Multinomial Logit Model (MNL) based on a linear utility specification has provided the foundation for the analysis of discrete choice. Despite its oversimplified assumptions regarding the actual decision-making, this model is still commonly used in practice because it enables a high level of interpretability. Interpretability is critical for researchers and practitioners to get insights into the complex human decision-making process. For instance, linear specifications allow for straightforward derivation of the value-of-time (VOT), i.e., the marginal rate of substitution between time and cost, that constitutes a highly relevant measure in a wide range of public transport policy.
23 Best Data Science Courses Online for Data Scientists JA Directives
Are you looking for Best Data Science Degree Online? This Online Data Science Course list will help you to become a top Data Scientist. Data science or data-driven science is one of today's fastest-growing fields. Do you want to become a Data Scientist in 2019? The list of the Data Science Degree will give you a clear idea from data science definition to expert's levels. Also, this Data Science training will give you an idea about data science, python, data scientist, big data, analytics, machine learning, deep learning and Artificial Intelligence (AI) which are the most booming topics now. You can be a data science master in a short period of time. All big companies, publishers, advertisers, and other industries are now highly depended on data science or machine learning. So, it is high time to learn some skills in data science, for example, get the high demanded Data Science online certifications. How does it work at the present time, why data scientist's career and data science jobs are in top position? If you like a trendy career, you have that opportunity right now and get hired by the big industries. At the same time, online entrepreneurs and business personals also need to update themselves with the fundamental machine learning skills to compete with the fast-moving industry. Below are few best Data Science online courses that might assist you to jump-start the knowledge of data science sector. Best Data Science online tutorial and programs listing displays the'Best Course,' 'Product Description,' 'Rating,' 'Students Enrolled' 'Product's Image' and as well as an Enroll button to purchase the Courses from respective learning platforms for your convenience. Description: If you want to learn machine learning then this is the perfect course for you. Two professional data scientists designed this course so that you can learn the theory and algorithms behind the machine learning.
Uncertainty Quantification for Kernel Methods
Csáji, Balázs Csanád, Kis, Krisztián Balázs
We propose a data-driven approach to quantify the uncertainty of models constructed by kernel methods. Our approach minimizes the needed distributional assumptions, hence, instead of working with, for example, Gaussian processes or exponential families, it only requires knowledge about some mild regularity of the measurement noise, such as it is being symmetric or exchangeable. We show, by building on recent results from finite-sample system identification, that by perturbing the residuals in the gradient of the objective function, information can be extracted about the amount of uncertainty our model has. Particularly, we provide an algorithm to build exact, non-asymptotically guaranteed, distribution-free confidence regions for ideal, noise-free representations of the function we try to estimate. For the typical convex quadratic problems and symmetric noises, the regions are star convex centered around a given nominal estimate, and have efficient ellipsoidal outer approximations. Finally, we illustrate the ideas on typical kernel methods, such as LS-SVC, KRR, kernelized LASSO and $\varepsilon$-SVR.
Neural networks versus Logistic regression for 30 days all-cause readmission prediction
Allam, Ahmed, Nagy, Mate, Thoma, George, Krauthammer, Michael
Heart failure (HF) is one of the leading causes of hospital admissions in the US. Readmission within 30 days after a HF hospitalization is both a recognized indicator for disease progression and a source of considerable financial burden to the healthcare system. Consequently, the identification of patients at risk for readmission is a key step in improving disease management and patient outcome. In this work, we used a large administrative claims dataset to (1)explore the systematic application of neural network-based models versus logistic regression for predicting 30 days all-cause readmission after discharge from a HF admission, and (2)to examine the additive value of patients' hospitalization timelines on prediction performance. Based on data from 272,778 (49% female) patients with a mean (SD) age of 73 years (14) and 343,328 HF admissions (67% of total admissions), we trained and tested our predictive readmission models following a stratified 5-fold cross-validation scheme. Among the deep learning approaches, a recurrent neural network (RNN) combined with conditional random fields (CRF) model (RNNCRF) achieved the best performance in readmission prediction with 0.642 AUC (95% CI, 0.640-0.645). Other models, such as those based on RNN, convolutional neural networks and CRF alone had lower performance, with a non-timeline based model (MLP) performing worst. A competitive model based on logistic regression with LASSO achieved a performance of 0.643 AUC (95%CI, 0.640-0.646). We conclude that data from patient timelines improve 30 day readmission prediction for neural network-based models, that a logistic regression with LASSO has equal performance to the best neural network model and that the use of administrative data result in competitive performance compared to published approaches based on richer clinical datasets.
Modified Causal Forests for Estimating Heterogeneous Causal Effects
Although science and the public celebrated the amazing predictive power of the new machine learningmethods, many researchers are left with some unease, simply because prediction does not imply causation. The ability to uncover causal relations is, however, at the core of most questions concerning the effects of particular policies, medical treatments, marketing campaigns, businessdecisions, etc. (see Athey, 2017, for a recent discussion). The recently rapidly expanding causal machine learning literature holds great promise for the improved estimation of causal effects by merging the statistics and econometrics literature oncausality with the supervised statistical and machine learning (ML) literature focussing on prediction. The classical causality literature clarifies the conditions needed for being able to estimate causal effects. It also shows how to transform a counterfactual causal problem into specific prediction problems (e.g., Imbens and Wooldridge, 2009). The latter literature on ML provides tools that can be highly effective in solving prediction problems (e.g.
Classification of load forecasting studies by forecasting problem to select load forecasting techniques and methodologies
Dumas, Jonathan, Cornélusse, Bertrand
This article proposes a two-dimensional classification methodology to select the relevant forecasting tools developed by the scientific community based on a classification of load forecasting studies. The inputs of the classifier are the articles of the literature and the outputs are articles classified into categories. The classification process relies on two couple of parameters that defines a forecasting problem. The temporal couple is the forecasting horizon and the forecasting resolution. The system couple is the system size and the load resolution. Each article is classified with key information about the dataset used and the forecasting tools implemented: the forecasting techniques (probabilistic or deterministic) and methodologies, the cleansing data techniques and the error metrics. This process is illustrated by reviewing and classifying thirty-four articles.
Dynamic Graph Representation Learning via Self-Attention Networks
Sankar, Aravind, Wu, Yanhong, Gou, Liang, Zhang, Wei, Yang, Hao
Learning latent representations of nodes in graphs is an important and ubiquitous task with widespread applications such as link prediction, node classification, and graph visualization. Previous methods on graph representation learning mainly focus on static graphs, however, many real-world graphs are dynamic and evolve over time. In this paper, we present Dynamic Self-Attention Network (DySAT), a novel neural architecture that operates on dynamic graphs and learns node representations that capture both structural properties and temporal evolutionary patterns. Specifically, DySAT computes node representations by jointly employing self-attention layers along two dimensions: structural neighborhood and temporal dynamics. We conduct link prediction experiments on two classes of graphs: communication networks and bipartite rating networks. Our experimental results show that DySAT has a significant performance gain over several different state-of-the-art graph embedding baselines.
Distributed sequential method for analyzing massive data
Wang, Zhanfeng, Chang, Yuan-chin Ivan
To analyse a very large data set containing lengthy variables, we adopt a sequential estimation idea and propose a parallel divide-and-conquer method. We conduct several conventional sequential estimation procedures separately, and properly integrate their results while maintaining the desired statistical properties. Additionally, using a criterion from the statistical experiment design, we adopt an adaptive sample selection, together with an adaptive shrinkage estimation method, to simultaneously accelerate the estimation procedure and identify the effective variables. We confirm the cogency of our methods through theoretical justifications and numerical results derived from synthesized data sets. We then apply the proposed method to three real data sets, including those pertaining to appliance energy use and particulate matter concentration.