Statistical Learning
CETransformer: Casual Effect Estimation via Transformer Based Representation Learning
Guo, Zhenyu, Zheng, Shuai, Liu, Zhizhe, Yan, Kun, Zhu, Zhenfeng
Treatment effect estimation, which refers to the estimation of causal effects and aims to measure the strength of the causal relationship, is of great importance in many fields but is a challenging problem in practice. As present, data-driven causal effect estimation faces two main challenges, i.e., selection bias and the missing of counterfactual. To address these two issues, most of the existing approaches tend to reduce the selection bias by learning a balanced representation, and then to estimate the counterfactual through the representation. However, they heavily rely on the finely hand-crafted metric functions when learning balanced representations, which generally doesn't work well for the situations where the original distribution is complicated. In this paper, we propose a CETransformer model for casual effect estimation via transformer based representation learning. To learn the representation of covariates(features) robustly, a self-supervised transformer is proposed, by which the correlation between covariates can be well exploited through self-attention mechanism. In addition, an adversarial network is adopted to balance the distribution of the treated and control groups in the representation space. Experimental results on three real-world datasets demonstrate the advantages of the proposed CETransformer, compared with the state-of-the-art treatment effect estimation methods.
2 Hours of ML a day -- Series
I am starting this blog, mainly to be accountable, disciplined and share my journey (both ups and downs) in my learning process. I tend to slack, after a hectic work day, and end up watching on OTT. I try to start learning (do it rightly for 2 days) and then end up slacking for 10โ15 days and start again. I am in a vicious cycle of wanting to learn and not being able to achieve it. I have a basic understanding of ML concepts like Decision trees, Linear Regression, Logistic Regression etc. Basic for me is -- knowing the algorithm, without any in-depth knowledge.
A Walk-through Artificial Intelligence
Hello everyone, I'm Tejaswini and in this blog I will be answering the following questions: Artificial Intelligence is all about training machines to mimic human behavior, specifically, the human brain and its thinking abilities. Similar to the human brain, AI systems develop the ability to rationalize and perform actions that have the best chance of achieving a specific goal. Artificial Intelligence focuses on performing 3 cognitive skills just like a human -- learning, reasoning, and self-correction. Let's have a quick look at the three broad categories of Artificial Intelligence and how we are rapidly evolving in these areas! Artificial Narrow Intelligence systems are designed and trained to complete one specific task and are often termed as Weak AI / Narrow AI.
Gentle Introduction to Gradient Descent and Momentum
In this article, we will talk about a fundamental concept in machine learning called the Gradient Descent. The gradient descent is one of the most popular algorithms that tends to reduce the error in prediction i.e minimizing your cost function. This might have been confusing but that's okay, before we jump into more details I'll give a very small gist of where it is mostly used. In deep learning, we have a concept called backpropagation. Wikipedia says " backpropagation computes the gradient of the loss function with respect to the weights of the network for a single inputโoutput example, and does so efficiently, unlike a naรฏve direct computation of the gradient with respect to each weight individually" I had a brain-freeze when I read this, so let me give you an intuitive example to help you understand better.
Support vector machines for learning reactive islands
Naik, Shibabrat, Krajลรกk, Vladimรญr, Wiggins, Stephen
We develop a machine learning framework that can be applied to data sets derived from the trajectories of Hamilton's equations. The goal is to learn the phase space structures that play the governing role for phase space transport relevant to particular applications. Our focus is on learning reactive islands in two degrees-of-freedom Hamiltonian systems. Reactive islands are constructed from the stable and unstable manifolds of unstable periodic orbits and play the role of quantifying transition dynamics. We show that support vector machines (SVM) is an appropriate machine learning framework for this purpose as it provides an approach for finding the boundaries between qualitatively distinct dynamical behaviors, which is in the spirit of the phase space transport framework. We show how our method allows us to find reactive islands directly in the sense that we do not have to first compute unstable periodic orbits and their stable and unstable manifolds. We apply our approach to the H\'enon-Heiles Hamiltonian system, which is a benchmark system in the dynamical systems community. We discuss different sampling and learning approaches and their advantages and disadvantages.
Decoupling Shrinkage and Selection for the Bayesian Quantile Regression
While modern day economics, and broadly social science research, is often faced with high dimensional estimation problems in which the number of potential explanatory variables is large, often larger than the number of sample observations, the extant literature for high dimensional methods has focused developments mainly on for conditional mean models. Moving beyond the conditional mean, by estimating quantile regression on the other hand, allows to gauge potentially heterogeneous effects of variables directly across the conditional response distribution. While highly influential in the risk-management and finance literature in calculating risk measures such as VaR (i.e., the loss a portfolio's value incurs at a specific probability level), quantile regression has experienced a recent surge in popularity within the macroeconomic literature to quantify risks and vulnerabilities of output growth in response to summary measures of financial health, aptly named growth-at-risk (GaR) (Adrian et al., 2019; Figueres and Jarociลski, 2020; Adams et al., 2020). As an important distinction to literature that focuses on forecasting crisis periods directly such as through Markov-switching models (Hubrich and Tetlow, 2015; Guรฉrin and Marcellino, 2013) or probit models (McCracken et al., 2021), GaR instead gives information about the accumulation of risks facing an economy. Since sources of risk can be numerous, high dimensional quantile problems are becoming ever more pertinent to policy makers and practitioners alike which has spurned methods that deal with variable selection and shrinkage for the quantile regression problem (Chernozhukov et al., 2010; Kohns and Szendrei, 2020; Hasenzagl et al., 2020).
Compressed particle methods for expensive models with application in Astronomy and Remote Sensing
Martino, Luca, Elvira, Vรญctor, Lรณpez-Santiago, Javier, Camps-Valls, Gustau
In many inference problems, the evaluation of complex and costly models is often required. In this context, Bayesian methods have become very popular in several fields over the last years, in order to obtain parameter inversion, model selection or uncertainty quantification. Bayesian inference requires the approximation of complicated integrals involving (often costly) posterior distributions. Generally, this approximation is obtained by means of Monte Carlo (MC) methods. In order to reduce the computational cost of the corresponding technique, surrogate models (also called emulators) are often employed. Another alternative approach is the so-called Approximate Bayesian Computation (ABC) scheme. ABC does not require the evaluation of the costly model but the ability to simulate artificial data according to that model. Moreover, in ABC, the choice of a suitable distance between real and artificial data is also required. In this work, we introduce a novel approach where the expensive model is evaluated only in some well-chosen samples. The selection of these nodes is based on the so-called compressed Monte Carlo (CMC) scheme. We provide theoretical results supporting the novel algorithms and give empirical evidence of the performance of the proposed method in several numerical experiments. Two of them are real-world applications in astronomy and satellite remote sensing.
Compressed Monte Carlo with application in particle filtering
Martino, Luca, Elvira, Vรญctor
Bayesian models have become very popular over the last years in several fields such as signal processing, statistics, and machine learning. Bayesian inference requires the approximation of complicated integrals involving posterior distributions. For this purpose, Monte Carlo (MC) methods, such as Markov Chain Monte Carlo and importance sampling algorithms, are often employed. In this work, we introduce the theory and practice of a Compressed MC (C-MC) scheme to compress the statistical information contained in a set of random samples. In its basic version, C-MC is strictly related to the stratification technique, a well-known method used for variance reduction purposes. Deterministic C-MC schemes are also presented, which provide very good performance. The compression problem is strictly related to the moment matching approach applied in different filtering techniques, usually called as Gaussian quadrature rules or sigma-point methods. C-MC can be employed in a distributed Bayesian inference framework when cheap and fast communications with a central processor are required. Furthermore, C-MC is useful within particle filtering and adaptive IS algorithms, as shown by three novel schemes introduced in this work. Six numerical results confirm the benefits of the introduced schemes, outperforming the corresponding benchmark methods. A related code is also provided.
A Survey on Role-Oriented Network Embedding
Jiao, Pengfei, Guo, Xuan, Pan, Ting, Zhang, Wang, Pei, Yulong
Recently, Network Embedding (NE) has become one of the most attractive research topics in machine learning and data mining. NE approaches have achieved promising performance in various of graph mining tasks including link prediction and node clustering and classification. A wide variety of NE methods focus on the proximity of networks. They learn community-oriented embedding for each node, where the corresponding representations are similar if two nodes are closer to each other in the network. Meanwhile, there is another type of structural similarity, i.e., role-based similarity, which is usually complementary and completely different from the proximity. In order to preserve the role-based structural similarity, the problem of role-oriented NE is raised. However, compared to community-oriented NE problem, there are only a few role-oriented embedding approaches proposed recently. Although less explored, considering the importance of roles in analyzing networks and many applications that role-oriented NE can shed light on, it is necessary and timely to provide a comprehensive overview of existing role-oriented NE methods. In this review, we first clarify the differences between community-oriented and role-oriented network embedding. Afterwards, we propose a general framework for understanding role-oriented NE and a two-level categorization to better classify existing methods. Then, we select some representative methods according to the proposed categorization and briefly introduce them by discussing their motivation, development and differences. Moreover, we conduct comprehensive experiments to empirically evaluate these methods on a variety of role-related tasks including node classification and clustering (role discovery), top-k similarity search and visualization using some widely used synthetic and real-world datasets...
Math and Data Science: What Do You Need to Know?
Mathematics is an integral part of data science. Any practicing data scientist or person interested in building a career in data science will need to have a strong background in specific mathematical fields. Depending on your career choice as a data scientist, you will need at least a B.A., M.A., or Ph.D. degree to qualify for hire at most organizations. A significant portion of your ability to translate your data science skills into real-world scenarios depends on your success and understanding of mathematics. Data science careers require mathematical study because machine learning algorithms, and performing analyses and discovering insights from data require math.