Statistical Learning
Screening for Sparse Online Learning
Sparsity promoting regularizers are widely used to impose low-complexity structure (e.g. l1-norm for sparsity) to the regression coefficients of supervised learning. In the realm of deterministic optimization, the sequence generated by iterative algorithms (such as proximal gradient descent) exhibit "finite activity identification", namely, they can identify the low-complexity structure in a finite number of iterations. However, most online algorithms (such as proximal stochastic gradient descent) do not have the property owing to the vanishing step-size and non-vanishing variance. In this paper, by combining with a screening rule, we show how to eliminate useless features of the iterates generated by online algorithms, and thereby enforce finite activity identification. One consequence is that when combined with any convergent online algorithm, sparsity properties imposed by the regularizer can be exploited for computational gains. Numerically, significant acceleration can be obtained.
Top 8 Data Science Use Cases in Marketing
In this article, we will discuss some notable data science use cases in marketing. As far as the primary goal of data science is to extract actionable insights from data, the marketing sphere cannot exclude the application of these insights for its advantage. Big data in marketing offers a chance to understand the target audience a lot better. Data science is mostly employed in marketing areas of search engine optimization (SEO), profiling, customer engagement, responsiveness, and real-time marketing campaigns. Furthermore, new ways to employ data science and analytics for marketing emerge every day.
Machine Learning Algorithms: Supervised Learning Tip to Tail
This course takes you from understanding the fundamentals of a machine learning project. Learners will understand and implement supervised learning techniques on real case studies to analyze business case scenarios where decision trees, k-nearest neighbours and support vector machines are optimally used. Learners will also gain skills to contrast the practical consequences of different data preparation steps and describe common production issues in applied ML. To be successful, you should have at least beginner-level background in Python programming (e.g., be able to read and code trace existing code, be comfortable with conditionals, loops, variables, lists, dictionaries and arrays). You should have a basic understanding of linear algebra (vector notation) and statistics (probability distributions and mean/median/mode).
Reconstruction and analysis of negatively buoyant jets with interpretable machine learning
Alvir, Marta, Grbฤiฤ, Luka, Sikirica, Ante, Kranjฤeviฤ, Lado
Such phenomena mostly occur during the disposal of wastewater from desalination plants, power plants, industrial factories, and cooling water discharge from liquefied natural gas (LNG) plants. Nowadays, numerous arid and semiarid coastal regions are encountering freshwater scarcity due to population growth and deficiency of portable water, hence the number of desalination plants is significantly increased. During the desalination process, using brackish or seawater, drinkable water is obtained, while the by-product is high-salinity concentrated effluents, socalled desalination brine, which is discharged back into the coastal waters using submerged outfalls. Besides elevated salt concentrations, desalination brine contains traces of chemicals, such as antiscalants, flocculants and coagulants ([43], which can lead to environmental degradation ([4])). To minimize harmful environmental effects and maximize dilution, the brine is predominantly discharged from diffusers directed upwards at an angle ([41]), producing negatively inclined buoyant jets. The design of the discharge systems and condition parameters are based on detailed experimental, mathematical, and numerical investigation in order to determine the characteristics of the jet. Experimental investigation of inclined buoyant jets was done by numerous researchers: [15], [31], [18], [45], [46], [2], [6]. Mathematical models such as VISJET ([14]) and CORJET ([16]) were implemented and compared with experimental data ([27], [32], [42]).
On the Privacy Risks of Algorithmic Recourse
Pawelczyk, Martin, Lakkaraju, Himabindu, Neel, Seth
As predictive models are increasingly being employed to make consequential decisions, there is a growing emphasis on developing techniques that can provide algorithmic recourse to affected individuals. While such recourses can be immensely beneficial to affected individuals, potential adversaries could also exploit these recourses to compromise privacy. In this work, we make the first attempt at investigating if and how an adversary can leverage recourses to infer private information about the underlying model's training data. To this end, we propose a series of novel membership inference attacks which leverage algorithmic recourse. More specifically, we extend the prior literature on membership inference attacks to the recourse setting by leveraging the distances between data instances and their corresponding counterfactuals output by state-of-the-art recourse methods. Extensive experimentation with real world and synthetic datasets demonstrates significant privacy leakage through recourses. Our work establishes unintended privacy leakage as an important risk in the widespread adoption of recourse methods.
New Interpretable Patterns and Discriminative Features from Brain Functional Network Connectivity Using Dictionary Learning
Ghayem, Fateme, Yang, Hanlu, Kantar, Furkan, Kim, Seung-Jun, Calhoun, Vince D., Adali, Tulay
Independent component analysis (ICA) of multi-subject functional magnetic resonance imaging (fMRI) data has proven useful in providing a fully multivariate summary that can be used for multiple purposes. ICA can identify patterns that can discriminate between healthy controls (HC) and patients with various mental disorders such as schizophrenia (Sz). Temporal functional network connectivity (tFNC) obtained from ICA can effectively explain the interactions between brain networks. On the other hand, dictionary learning (DL) enables the discovery of hidden information in data using learnable basis signals through the use of sparsity. In this paper, we present a new method that leverages ICA and DL for the identification of directly interpretable patterns to discriminate between the HC and Sz groups. We use multi-subject resting-state fMRI data from $358$ subjects and form subject-specific tFNC feature vectors from ICA results. Then, we learn sparse representations of the tFNCs and introduce a new set of sparse features as well as new interpretable patterns from the learned atoms. Our experimental results show that the new representation not only leads to effective classification between HC and Sz groups using sparse features, but can also identify new interpretable patterns from the learned atoms that can help understand the complexities of mental diseases such as schizophrenia.
A Randomised Subspace Gauss-Newton Method for Nonlinear Least-Squares
Cartis, Coralia, Fowkes, Jaroslav, Shao, Zhen
We propose a Randomised Subspace Gauss-Newton (R-SGN) algorithm for solving nonlinear least-squares optimization problems, that uses a sketched Jacobian of the residual in the variable domain and solves a reduced linear least-squares on each iteration. A sublinear global rate of convergence result is presented for a trust-region variant of R-SGN, with high probability, which matches deterministic counterpart results in the order of the accuracy tolerance. Promising preliminary numerical results are presented for R-SGN on logistic regression and on nonlinear regression problems from the CUTEst collection.
Deep equilibrium models as estimators for continuous latent variables
Tsuchida, Russell, Ong, Cheng Soon
Principal Component Analysis (PCA) and its exponential family extensions have three components: observations, latents and parameters of a linear transformation. We consider a generalised setting where the canonical parameters of the exponential family are a nonlinear transformation of the latents. We show explicit relationships between particular neural network architectures and the corresponding statistical models. We find that deep equilibrium models -- a recently introduced class of implicit neural networks -- solve maximum a-posteriori (MAP) estimates for the latents and parameters of the transformation. Our analysis provides a systematic way to relate activation functions, dropout, and layer structure, to statistical assumptions about the observations, thus providing foundational principles for unsupervised DEQs. For hierarchical latents, individual neurons can be interpreted as nodes in a deep graphical model. Our DEQ feature maps are end-to-end differentiable, enabling fine-tuning for downstream tasks.
Robust Model Selection of Non Tree-Structured Gaussian Graphical Models
Zahin, Abrar, Anguluri, Rajasekhar, Kosut, Oliver, Sankar, Lalitha, Dasarathy, Gautam
We consider the problem of learning the structure underlying a Gaussian graphical model when the variables (or subsets thereof) are corrupted by independent noise. A recent line of work establishes that even for tree-structured graphical models, only partial structure recovery is possible and goes on to devise algorithms to identify the structure up to an (unavoidable) equivalence class of trees. We extend these results beyond trees and consider the model selection problem under noise for non tree-structured graphs, as tree graphs cannot model several real-world scenarios. Although unidentifiable, we show that, like the tree-structured graphs, the ambiguity is limited to an equivalence class. This limited ambiguity can help provide meaningful clustering information (even with noise), which is helpful in computer and social networks, protein-protein interaction networks, and power networks. Furthermore, we devise an algorithm based on a novel ancestral testing method for recovering the equivalence class. We complement these results with finite sample guarantees for the algorithm in the high-dimensional regime.