Statistical Learning
Top 10 Challenges to Practicing Data Science at Work
A recent survey of over 16,000 data professionals showed that the most common challenges to data science included dirty data (36%), lack of data science talent (30%) and lack of management support (27%). Also, data professionals reported experiencing around three challenges in the previous year. A principal component analysis of the 20 challenges studied showed that challenges can be grouped into five categories. Data science is about finding useful insights and putting them to use. Data science, however, doesn't occur in a vacuum.
Mathematics for Machine Learning: PCA Coursera
About this course: This course introduces the mathematical foundations to derive Principal Component Analysis (PCA), a fundamental dimensionality reduction technique. We'll cover some basic statistics of data sets, such as mean values and variances, we'll compute distances and angles between vectors using inner products and derive orthogonal projections of data onto lower-dimensional subspaces. Using all these tools, we'll then derive PCA as a method that minimizes the average squared reconstruction error between data points and their reconstruction. At the end of this course, you'll be familiar with important mathematical concepts and you can implement PCA all by yourself. If you're struggling, you'll find a set of jupyter notebooks that will allow you to explore properties of the techniques and walk you through what you need to do to get on track.
Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression
Kuchibhotla, Arun Kumar, Chakrabortty, Abhishek
Concentration inequalities form an essential toolkit in the study of high-dimensional statistical methods. Most of the relevant statistics literature is based on the assumptions of sub-Gaussian/sub-exponential random vectors. In this paper, we bring together various probability inequalities for sums of independent random variables under much weaker exponential type (sub-Weibull) tail assumptions. These results extract a part sub-Gaussian tail behavior in finite samples, matching the asymptotics governed by the central limit theorem, and are compactly represented in terms of a new Orlicz quasi-norm - the Generalized Bernstein-Orlicz norm - that typifies such tail behaviors. We illustrate the usefulness of these inequalities through the analysis of four fundamental problems in high-dimensional statistics. In the first two problems, we study the rate of convergence of the sample covariance matrix in terms of the maximum elementwise norm and the maximum k-sub-matrix operator norm which are key quantities of interest in bootstrap procedures and high-dimensional structured covariance matrix estimation. The third example concerns the restricted eigenvalue condition, required in high dimensional linear regression, which we verify for all sub-Weibull random vectors under only marginal (not joint) tail assumptions on the covariates. To our knowledge, this is the first unified result obtained in such generality. In the final example, we consider the Lasso estimator for linear regression and establish its rate of convergence under much weaker tail assumptions (on the errors as well as the covariates) than those in the existing literature. The common feature in all our results is that the convergence rates under most exponential tails match the usual ones under sub-Gaussian assumptions. Finally, we also establish a high-dimensional CLT and tail bounds for empirical processes for sub-Weibulls.
Supervised vs. Unsupervised Learning
Within the field of machine learning, there are two main types of tasks: supervised, and unsupervised. The main difference between the two types is that supervised learning is done using a ground truth, or in other words, we have prior knowledge of what the output values for our samples should be. Therefore, the goal of supervised learning is to learn a function that, given a sample of data and desired outputs, best approximates the relationship between input and output observable in the data. Unsupervised learning, on the other hand, does not have labeled outputs, so its goal is to infer the natural structure present within a set of data points. Supervised learning is typically done in the context of classification, when we want to map input to output labels, or regression, when we want to map input to a continuous output.
Large Scale Online Brand Networks to Study Brand Effects
Malhotra, Pankhuri (University of Illinois at Chicago) | Bhattacharyya, Siddhartha (University of Illinois at Chicago)
Mining consumer perceptions of brands has been a dominant research area in marketing. The marketing literature provides a well-developed rationale for proposing brands as intangible assets that significantly contribute to firm performance. Consumer-brand perceptions typically collected through surveys or focus groups, require recruitment and interaction with a large set of participants; leading to cost, feasibility and validity issues. The advent of web 2.0 opens the door to the application of a wide range of data-centric approaches which can automate and scale beyond the traditional methods used in marketing science. We address this knowledge area by exploiting social media based brand communities to generate a brand network, incorporating consumer perceptions across a broad ecosystem of brands. A brand network is one in which individual nodes represent brands, and a weighted link between two nodes represents the strength of consumer co-interest in these two brands. The implicit brand-brand network is used to examine two branding effects, in particular, positioning and performance. We use hard and soft clustering algorithms, Walktrap Clustering and Stochastic Block Modelling respectively, to identify subsets of closely related brands; and this provides the basis for examining brand positioning. We also examine how a focal brand’s location in the brand network relates to performance, measured in terms of relative market share. For this, a hierarchical regression analysis is conducted between brand network variables and brand performance. While the size of brand community on Twitter does relate to brand performance, the brand network variables like degree, eigenvector centrality and between-industry links help improve the model fit considerably.
Visual Listening In: Extracting Brand Image Portrayed on Social Media
Liu, Liu (New York University) | Dzyabura, Daria (New York University) | Mizik, Natalie (University of Washington)
Marketing academics and practitioners recognize the importance of monitoring consumer online conversations about brands. The focus so far has been on text content. However, images are on their way to surpassing text as the medium of choice for social conversations. In these images, consumers often tag brands. We propose a ``visual listening in" approach to measuring how brands are portrayed on social media (Instagram) by mining visual content posted by users, and show what insights brand managers can gather from social media by using this approach. We first use two supervised machine learning methods, traditional support vector machine classifiers and deep convolutional neural networks, to measure brand attributes (glamorous, rugged, healthy, fun) from images. We then apply the classifiers to brand-related images posted on social media. We study 56 brands in the apparel and beverages categories, and compare their portrayal in consumer-created images with images on the firm's official Instagram account, as well as with consumer brand perceptions measured in a national brand survey. Although the three measures exhibit convergent validity, we find key differences between how consumers and firms portray the brands on visual social media, and how the average consumer perceives the brands.
A Framework for Utilizing Lab Test Results for Clinical Prediction of ICU Patients
Masud, Mohammad Mehedy (United Arab Emirates University) | Cheratta, Muhsin (United Arab Emirates University)
Clinical decision support has gained significant attention in recent years, especially with the advancement of data analytics techniques. One active research area in this domain is survival prediction or deterioration prediction of critical care patients, such as intensive care unit (ICU) patients. Usually, ICUs are equipped with continuous monitoring devices, which monitor vital signs such as heart rate, blood pressure, Oxygen saturation and so on. In addition to this, ICU patients also undergo different pathological (i.e., lab) tests. Recent studies claim that vital signs can be used to predict the near future status of a patient, with the help of predictive analytics. However, in this work, we investigate the usefulness of lab test results in patient survival prediction, which have been rarely used for this purpose. We propose a framework for utilizing the lab test data for this clinical prediction task. We encounter several challenges associated with this task, including variable-length feature vector, longitudinal features, missing data, class imbalance and high dimensionality. The proposed work addresses most of these challenges under this single framework. In this framework we propose a novel orthogonal clustering technique to reduce data dimensions as well as missing data. We also propose a systematic approach to inject informative background knowledge into the data and increase the prediction performance. The proposed technique has been evaluated on a real ICU patients database, achieving notable success in reducing 66% of the data dimensions without discarding any feature, while improving the weighted average F1-score 5% on average and achieving about 3 times speedup. We believe that the proposed technique will provide a powerful framework in the field of clinical and healthcare data analytics and healthcare decision support.
Using Digital Purchasing Data to Generate Public Health Evidence: Learning Unhealthy Beverage Demand from Grocery Transaction Data
Mamiya, Hiroshi (McGill University) | Lu, Xing Han (McGill University) | Ma, Yu (McGill University) | Buckeridge, David L. (McGill University)
Unhealthy diet plays a major role in driving chronic disease incidence and prevalence. Taxation of unhealthy food has been proposed to improve population-level dietary patterns, and its effectiveness can be estimated by the prediction of the change in unhealthy food purchasing upon increase of food price. Recent availability of grocery transaction data from scanner technologies enables an accurate prediction of food sales. However, the very large number of product at-tributes in these data prohibits the application of conventional statistical learning algorithms. In this study, we explored the predictive performance of learning algorithms adapted for high-dimensional data, namely the Least Absolute Shrinkage and Selection Operator (LASSO) and Decision Tree Regressor with Adaptive Boosting (DTR-Ada-Boost), in comparison with a conventional statistical learning based on Ordinary Least Square (OLS). LASSO demonstrated superior predictive accuracy to OLS, possibly due to its ability to reduce over fitting and collinearity across predictive features of food sales. DTR-AdaBoost showed the best predictive accuracy, suggesting the presence of extensive non-linearity between the predictive features in the transaction data and sales.
Multiple-Implementation Testing of Supervised Learning Software
Srisakaokul, Siwakorn (University of Illinois at Urbana-Champaign) | Wu, Zhengkai (University of Illinois at Urbana-Champaign) | Astorga, Angello (University of Illinois at Urbana-Champaign) | Alebiosu, Oreoluwa (University of Illinois at Urbana-Champaign) | Xie, Tao (University of Illinois at Urbana-Champaign)
Machine Learning (ML) algorithms are now used in a wide range of application domains in society. Naturally, software implementations of these algorithms have become ubiquitous. Faults in ML software can cause substantial losses in these application domains. Thus, it is very critical to conduct effective testing of ML software to detect and eliminate its faults. However, testing ML software is difficult, partly because producing test oracles used for checking behavior correctness (such as using expected properties or expected test outputs) is challenging. In this paper, we propose an approach of multiple-implementation testing to test supervised learning software, a major type of ML software. In particular, our approach derives a test input's proxy oracle from the majority-voted output running the test input of multiple implementations of the same algorithm (based on a pre-defined percentage threshold). Our approach reports likely those test inputs whose outputs (produced by an implementation under test) are different from the majority-voted outputs as failing tests. We evaluate our approach on two highly popular supervised learning algorithms: k-Nearest Neighbor (kNN) and Naive Bayes (NB). Our results show that our approach is highly effective in detecting faults in real-world supervised learning software. In particular, our approach detects 13 real faults and 1 potential fault from 19 kNN implementations and 16 real faults from 7 NB implementations. Our approach can even detect 7 real faults and 1 potential fault among the three popularly used open-source ML projects (Weka, RapidMiner, and KNIME).
Adequacy of the Gradient-Descent Method for Classifier Evasion Attacks
Han, Yi (The University of Melbourne) | Rubinstein, Benjamin (The University of Melbourne)
Despite the widespread use of machine learning in adversarial settings such as computer security, recent studies have demonstrated vulnerabilities to evasion attacks---carefully crafted adversarial samples that closely resemble legitimate instances, but cause misclassification. In this paper, we examine the adequacy of the leading approach to generating adversarial samples---the gradient-descent approach. In particular (1) we perform extensive experiments on three datasets, MNIST, USPS and Spambase, in order to analyse the effectiveness of the gradient-descent method against non-linear support vector machines, and conclude that carefully reduced kernel smoothness can significantly increase robustness to the attack; (2) we demonstrate that separated inter-class support vectors lead to more secure models, and propose a quantity similar to margin that can efficiently predict potential susceptibility to gradient-descent attacks, before the attack is launched; and (3) we design a new adversarial sample construction algorithm based on optimising the multiplicative ratio of class decision functions.