Statistical Learning
Data science cookbook style code reference in Python for beginners
Here is another resource I use for teaching my students at AI for Edge computing course. I like this resource because I like the cookbook style of learning to code. The resource is based on the book Machine Learning With Python Cookbook. I like the approach of using a simple simulated dataset like we see in LDA for dimensionality reduction and pandas functions. This link contains others such as linux, postgres etc which I have not tried.
Large-scale Gender/Age Prediction of Tumblr Users
Zhan, Yao, Hu, Changwei, Hu, Yifan, Kasturi, Tejaswi, Ramasamy, Shanmugam, Gillingham, Matt, Yamamoto, Keith
Abstract--T umblr, as a leading content provider and social media, attracts 371 million monthly visits, 280 million blo gs and 53.3 million daily posts However, it is a challenging task t o target specific demographic groups for ads, since T umblr doe s not require user information like gender and ages during the ir registration. Hence, to promote ad targeting, it is essenti al to predict user's demography using rich content such as posts, images and social connections. In this paper, we propose gra ph based and deep learning models for age and gender prediction s, which take into account user activities and content feature s. For graph based models, we come up with two approaches, network embedding and label propagation, to generate connection fe atures as well as directly infer user's demography. Experimental results on real T umblr daily dataset, with hun dreds of millions of active users and billions of following relati ons, demonstrate that our approaches significantly outperform t he baseline model, by improving the accuracy relatively by 81% for age, and the AUC and accuracy by 5% for gender . Online social media has become a ubiquitous part of our daily life, which allows us to easily share ideas/contents w ith other users, discuss social events/activities, and get con nected with friends. The rich content including text, images, and videos, provide great opportunities for advertisers to champion th eir products to specific groups. In particular, Tumblr offers "n ative advertisement" that allows advertisers to present their sp on-sored posts on the users" interface. Native advertising has gained over 3 billion paid ad impressions in 2015 since it was started in 2012 [1].
Hydrological time series forecasting using simple combinations: Big data testing and investigations on one-year ahead river flow predictability
Papacharalampous, Georgia, Tyralis, Hristos
Delivering useful hydrological forecasts is critical for urban and agricultural water management, hydropower generation, flood protection and management, drought mitigation and alleviation, and river basin planning and management, among others. In this work, we present and appraise a new methodology for hydrological time series forecasting. This methodology is based on simple combinations. The appraisal is made by using a big dataset consisted of 90-year-long mean annual river flow time series from approximately 600 stations. Covering large parts of North America and Europe, these stations represent various climate and catchment characteristics, and thus can collectively support benchmarking. Five individual forecasting methods and 26 variants of the introduced methodology are applied to each time series. The application is made in one-step ahead forecasting mode. The individual methods are the last-observation benchmark, simple exponential smoothing, complex exponential smoothing, automatic autoregressive fractionally integrated moving average (ARFIMA) and Facebook's Prophet, while the 26 variants are defined by all the possible combinations (per two, three, four or five) of the five afore-mentioned methods. The findings have both practical and theoretical implications. The simple methodology of the study is identified as well-performing in the long run. Our large-scale results are additionally exploited for finding an interpretable relationship between predictive performance and temporal dependence in the river flow time series, and for examining one-year ahead river flow predictability.
Bayesian task embedding for few-shot Bayesian optimization
Atkinson, Steven, Ghosh, Sayan, Chennimalai-Kumar, Natarajan, Khan, Genghis, Wang, Liping
We describe a method for Bayesian optimization by which one may incorporate data from multiple systems whose quantitative interrelationships are unknown a priori. All general (nonreal-valued) features of the systems are associated with continuous latent variables that enter as inputs into a single metamodel that simultaneously learns the response surfaces of all of the systems. Bayesian inference is used to determine appropriate beliefs regarding the latent variables. We explain how the resulting probabilistic metamodel may be used for Bayesian optimization tasks and demonstrate its implementation on a variety of synthetic and real-world examples, comparing its performance under zero-, one-, and few-shot settings against traditional Bayesian optimization, which usually requires substantially more data from the system of interest.
On Large-Scale Dynamic Topic Modeling with Nonnegative CP Tensor Decomposition
Ahn, Miju, Eikmeier, Nicole, Haddock, Jamie, Kassab, Lara, Kryshchenko, Alona, Leonard, Kathryn, Needell, Deanna, Madushani, R. W. M. A., Sizikova, Elena, Wang, Chuntian
There is currently an unprecedented demand for large-scale temporal data analysis due to the explosive growth of data. Dynamic topic modeling has been widely used in social and data sciences with the goal of learning latent topics that emerge, evolve, and fade over time. Previous work on dynamic topic modeling primarily employ the method of nonnegative matrix factorization (NMF), where slices of the data tensor are each factorized into the product of lower-dimensional nonnegative matrices. With this approach, however, information contained in the temporal dimension of the data is often neglected or underutilized. To overcome this issue, we propose instead adopting the method of nonnegative CANDECOMP/PARAPAC (CP) tensor decomposition (NNCPD), where the data tensor is directly decomposed into a minimal sum of outer products of nonnegative vectors, thereby preserving the temporal information. The viability of NNCPD is demonstrated through application to both synthetic and real data, where significantly improved results are obtained compared to those of typical NMF-based methods. The advantages of NNCPD over such approaches are studied and discussed. To the best of our knowledge, this is the first time that NNCPD has been utilized for the purpose of dynamic topic modeling, and our findings will be transformative for both applications and further developments.
A Loss-Function for Causal Machine-Learning
Causal machine-learning is about predicting the net-effect (true-lift) of treatments. Given the data of a treatment group and a control group, it is similar to a standard supervised-learning problem. Unfortunately, there is no similarly well-defined loss function due to the lack of point-wise true values in the data. Many advances in modern machine-learning are not directly applicable due to the absence of such loss function. We propose a novel method to define a loss function in this context, which is equal to mean-square-error (MSE) in a standard regression problem. Our loss function is universally applicable, thus providing a general standard to evaluate the quality of any model/strategy that predicts the true-lift. We demonstrate that despite its novel definition, one can still perform gradient descent directly on this loss function to find the best fit. This leads to a new way to train any parameter-based model, such as deep neural networks, to solve causal machine-learning problems without going through the meta-learner strategy.
Accelerating Smooth Games by Manipulating Spectral Shapes
Azizian, Waรฏss, Scieur, Damien, Mitliagkas, Ioannis, Lacoste-Julien, Simon, Gidel, Gauthier
We use matrix iteration theory to characterize acceleration in smooth games. We define the spectral shape of a family of games as the set containing all eigenvalues of the Jacobians of standard gradient dynamics in the family. Shapes restricted to the real line represent well-understood classes of problems, like minimization. Shapes spanning the complex plane capture the added numerical challenges in solving smooth games. In this framework, we describe gradient-based methods, such as extragradient, as transformations on the spectral shape. Using this perspective, we propose an optimal algorithm for bilinear games. For smooth and strongly monotone operators, we identify a continuum between convex minimization, where acceleration is possible using Polyak's momentum, and the worst case where gradient descent is optimal. Finally, going beyond first-order methods, we propose an accelerated version of consensus optimization.
Modeling Historical AIS Data For Vessel Path Prediction: A Comprehensive Treatment
Tu, Enmei, Zhang, Guanghao, Mao, Shangbo, Rachmawati, Lily, Huang, Guang-Bin
The prosperity of artificial intelligence has aroused intensive interests in intelligent/autonomous navigation, in which path prediction is a key functionality for decision supports, e.g. route planning, collision warning, and traffic regulation. For maritime intelligence, Automatic Identification System (AIS) plays an important role because it recently has been made compulsory for large international commercial vessels and is able to provide nearly real-time information of the vessel. Therefore AIS data based vessel path prediction is a promising way in future maritime intelligence. However, real-world AIS data collected online are just highly irregular trajectory segments (AIS message sequences) from different types of vessels and geographical regions, with possibly very low data quality. So even there are some works studying how to build a path prediction model using historical AIS data, but still, it is a very challenging problem. In this paper, we propose a comprehensive framework to model massive historical AIS trajectory segments for accurate vessel path prediction. Experimental comparisons with existing popular methods are made to validate the proposed approach and results show that our approach could outperform the baseline methods by a wide margin.
Auditing and Debugging Deep Learning Models via Decision Boundaries: Individual-level and Group-level Analysis
Yousefzadeh, Roozbeh, O'Leary, Dianne P.
Deep learning models have been criticized for their lack of easy interpretation, which undermines confidence in their use for important applications. Nevertheless, they are consistently utilized in many applications, consequential to humans' lives, mostly because of their better performance. Therefore, there is a great need for computational methods that can explain, audit, and debug such models. Here, we use flip points to accomplish these goals for deep learning models with continuous output scores (e.g., computed by softmax), used in social applications. A flip point is any point that lies on the boundary between two output classes: e.g. for a model with a binary yes/no output, a flip point is any input that generates equal scores for "yes" and "no". The flip point closest to a given input is of particular importance because it reveals the least changes in the input that would change a model's classification, and we show that it is the solution to a well-posed optimization problem. Flip points also enable us to systematically study the decision boundaries of a deep learning classifier. The resulting insight into the decision boundaries of a deep model can clearly explain the model's output on the individual-level, via an explanation report that is understandable by non-experts. We also develop a procedure to understand and audit model behavior towards groups of people. Flip points can also be used to alter the decision boundaries in order to improve undesirable behaviors. We demonstrate our methods by investigating several models trained on standard datasets used in social applications of machine learning. We also identify the features that are most responsible for particular classifications and misclassifications.