Statistical Learning
Fundamental Tradeoffs in Distributionally Adversarial Training
Mehrabi, Mohammad, Javanmard, Adel, Rossi, Ryan A., Rao, Anup, Mai, Tung
Adversarial training is among the most effective techniques to improve the robustness of models against adversarial perturbations. However, the full effect of this approach on models is not well understood. For example, while adversarial training can reduce the adversarial risk (prediction error against an adversary), it sometimes increase standard risk (generalization error when there is no adversary). Even more, such behavior is impacted by various elements of the learning problem, including the size and quality of training data, specific forms of adversarial perturbations in the input, model overparameterization, and adversary's power, among others. In this paper, we focus on \emph{distribution perturbing} adversary framework wherein the adversary can change the test distribution within a neighborhood of the training data distribution. The neighborhood is defined via Wasserstein distance between distributions and the radius of the neighborhood is a measure of adversary's manipulative power. We study the tradeoff between standard risk and adversarial risk and derive the Pareto-optimal tradeoff, achievable over specific classes of models, in the infinite data limit with features dimension kept fixed. We consider three learning settings: 1) Regression with the class of linear models; 2) Binary classification under the Gaussian mixtures data model, with the class of linear classifiers; 3) Regression with the class of random features model (which can be equivalently represented as two-layer neural network with random first-layer weights). We show that a tradeoff between standard and adversarial risk is manifested in all three settings. We further characterize the Pareto-optimal tradeoff curves and discuss how a variety of factors, such as features correlation, adversary's power or the width of two-layer neural network would affect this tradeoff.
Sensitivity Prewarping for Local Surrogate Modeling
Wycoff, Nathan, Binois, Mickaรซl, Gramacy, Robert B.
In the continual effort to improve product quality and decrease operations costs, computational modeling is increasingly being deployed to determine feasibility of product designs or configurations. Surrogate modeling of these computer experiments via local models, which induce sparsity by only considering short range interactions, can tackle huge analyses of complicated input-output relationships. However, narrowing focus to local scale means that global trends must be re-learned over and over again. In this article, we propose a framework for incorporating information from a global sensitivity analysis into the surrogate model as an input rotation and rescaling preprocessing step. We discuss the relationship between several sensitivity analysis methods based on kernel regression before describing how they give rise to a transformation of the input variables. Specifically, we perform an input warping such that the "warped simulator" is equally sensitive to all input directions, freeing local models to focus on local dynamics. Numerical experiments on observational data and benchmark test functions, including a high-dimensional computer simulator from the automotive industry, provide empirical validation.
Causal Gradient Boosting: Boosted Instrumental Variable Regression
Bakhitov, Edvard, Singh, Amandeep
Recent advances in the literature have demonstrated that standard supervised learning algorithms are ill-suited for problems with endogenous explanatory variables. To correct for the endogeneity bias, many variants of nonparameteric instrumental variable regression methods have been developed. In this paper, we propose an alternative algorithm called boostIV that builds on the traditional gradient boosting algorithm and corrects for the endogeneity bias. The algorithm is very intuitive and resembles an iterative version of the standard 2SLS estimator. Moreover, our approach is data driven, meaning that the researcher does not have to make a stance on neither the form of the target function approximation nor the choice of instruments. We demonstrate that our estimator is consistent under mild conditions. We carry out extensive Monte Carlo simulations to demonstrate the finite sample performance of our algorithm compared to other recently developed methods. We show that boostIV is at worst on par with the existing methods and on average significantly outperforms them.
Heating up decision boundaries: isocapacitory saturation, adversarial scenarios and generalization bounds
Georgiev, Bogdan, Franken, Lukas, Mukherjee, Mayukh
In the present work we study classifiers' decision boundaries via Brownian motion processes in ambient data space and associated probabilistic techniques. Intuitively, our ideas correspond to placing a heat source at the decision boundary and observing how effectively the sample points warm up. We are largely motivated by the search for a soft measure that sheds further light on the decision boundary's geometry. En route, we bridge aspects of potential theory and geometric analysis (Mazya, 2011, Grigoryan-Saloff-Coste, 2002) with active fields of ML research such as adversarial examples and generalization bounds. First, we focus on the geometric behavior of decision boundaries in the light of adversarial attack/defense mechanisms. Experimentally, we observe a certain capacitory trend over different adversarial defense strategies: decision boundaries locally become flatter as measured by isoperimetric inequalities (Ford et al, 2019); however, our more sensitive heat-diffusion metrics extend this analysis and further reveal that some non-trivial geometry invisible to plain distance-based methods is still preserved. Intuitively, we provide evidence that the decision boundaries nevertheless retain many persistent "wiggly and fuzzy" regions on a finer scale. Second, we show how Brownian hitting probabilities translate to soft generalization bounds which are in turn connected to compression and noise stability (Arora et al, 2018), and these bounds are significantly stronger if the decision boundary has controlled geometric features.
Likelihood Ratio Exponential Families
Brekelmans, Rob, Nielsen, Frank, Makhzani, Alireza, Galstyan, Aram, Steeg, Greg Ver
The exponential family is well known in machine learning and statistical physics as the maximum entropy distribution subject to a set of observed constraints, while the geometric mixture path is common in MCMC methods such as annealed importance sampling. Linking these two ideas, recent work has interpreted the geometric mixture path as an exponential family of distributions to analyze the thermodynamic variational objective (TVO). We extend these likelihood ratio exponential families to include solutions to rate-distortion (RD) optimization, the information bottleneck (IB) method, and recent rate-distortion-classification approaches which combine RD and IB. This provides a common mathematical framework for understanding these methods via the conjugate duality of exponential families and hypothesis testing. Further, we collect existing results to provide a variational representation of intermediate RD or TVO distributions as a minimizing an expectation of KL divergences. This solution also corresponds to a size-power tradeoff using the likelihood ratio test and the Neyman Pearson lemma. In thermodynamic integration bounds such as the TVO, we identify the intermediate distribution whose expected sufficient statistics match the log partition function.
Score Matched Conditional Exponential Families for Likelihood-Free Inference
Pacchiardi, Lorenzo, Dutta, Ritabrata
To perform Bayesian inference for stochastic simulator models for which the likelihood is not accessible, Likelihood-Free Inference (LFI) relies on simulations from the model. Standard LFI methods can be split according to how these simulations are used: to build an explicit Surrogate Likelihood, or to accept/reject parameter values according to a measure of distance from the observations (Approximate Bayesian Computation (ABC)). In both cases, simulations are adaptively tailored to the value of the observation. Here, we generate parameter-simulation pairs from the model independently on the observation, and use them to learn a conditional exponential family likelihood approximation; to parametrize it, we use Neural Networks whose weights are tuned with Score Matching. With our likelihood approximation, we can employ MCMC for doubly intractable distributions to draw samples from the posterior for any number of observations without additional model simulations, with performance competitive to comparable approaches. Further, the sufficient statistics of the exponential family can be used as summaries in ABC, outperforming the state-of-the-art method in five different models with known likelihood. Finally, we apply our method to a challenging model from meteorology.
Transmission heterogeneities, kinetics, and controllability of SARS-CoV-2
A minority of people infected with severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) transmit most infections. How does this happen? Sun et al. reconstructed transmission in Hunan, China, up to April 2020. Such detailed data can be used to separate out the relative contribution of transmission control measures aimed at isolating individuals relative to population-level distancing measures. The authors found that most of the secondary transmissions could be traced back to a minority of infected individuals, and well over half of transmission occurred in the presymptomatic phase. Furthermore, the duration of exposure to an infected person combined with closeness and number of household contacts constituted the greatest risks for transmission, particularly when lockdown conditions prevailed. These findings could help in the design of infection control policies that have the potential to minimize both virus transmission and economic strain. Science , this issue p. [eabe2424][1] ### INTRODUCTION The role of transmission heterogeneities in severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) dynamics remains unclear, particularly those heterogeneities driven by demography, behavior, and interventions. To understand individual heterogeneities and their effect on disease control, we analyze detailed contact-tracing data from Hunan, a province in China adjacent to Hubei and one of the first regions to experience a SARS-CoV-2 outbreak in January to March 2020. The Hunan outbreak was swiftly brought under control by March 2020 through a combination of nonpharmaceutical interventions including population-level mobility restriction (i.e., lockdown), traveler screening, case isolation, contact tracing, and quarantine. In parallel, highly detailed epidemiological information on SARS-CoV-2โinfected individuals and their close contacts was collected by the Hunan Provincial Center for Disease Control and Prevention. ### RATIONALE Contact-tracing data provide information to reconstruct transmission chains and understand outbreak dynamics. These data can in turn generate valuable intelligence on key epidemiological parameters and risk factors for transmission, which paves the way for more-targeted and cost-effective interventions. ### RESULTS On the basis of epidemiological information and exposure diaries on 1178 SARS-CoV-2โinfected individuals and their 15,648 close contacts, we developed a series of statistical and computational models to stochastically reconstruct transmission chains, identify risk factors for transmission, and infer the infectiousness profile over the course of a typical infection. We observe overdispersion in the distribution of secondary infections, with 80% of secondary cases traced back to 15% of infections, which indicates substantial transmission heterogeneities. We find that SARS-CoV-2 transmission risk scales positively with the duration of exposure and the closeness of social interactions, with the highest per-contact risk estimated in the household. Lockdown interventions increase transmission risk in families and households, whereas the timely isolation of infected individuals reduces risk across all types of contacts. There is a gradient of increasing susceptibility with age but no significant difference in infectivity by age or clinical severity. Early isolation of SARS-CoV-2โinfected individuals drastically alters transmission kinetics, leading to shorter generation and serial intervals and a higher fraction of presymptomatic transmission. After adjusting for the censoring effects of isolation, we find that the infectiousness profile of a typical SARS-CoV-2 patient peaks just before symptom onset, with 53% of transmission occurring in the presymptomatic phase in an uncontrolled setting. We then use these results to evaluate the effectiveness of individual-based strategies (case isolation and contact quarantine) both alone and in combination with population-level contact reductions. We find that a plausible parameter space for SARS-CoV-2 control is restricted to scenarios where interventions are synergistically combined, owing to the particular transmission kinetics of this virus. ### CONCLUSION There is considerable heterogeneity in SARS-CoV-2 transmission owing to individual differences in biology and contacts that is modulated by the effects of interventions. We estimate that about half of secondary transmission events occur in the presymptomatic phase of a primary case in uncontrolled outbreaks. Achieving epidemic control requires that isolation and contact-tracing interventions are layered with population-level approaches, such as mask wearing, increased teleworking, and restrictions on large gatherings. Our study also demonstrates the value of conducting high-quality contact-tracing investigations to advance our understanding of the transmission dynamics of an emerging pathogen. ![Figure][2] Transmission chains, contact patterns, and transmission kinetics of SARS-CoV-2 in Hunan, China, based on case and contact-tracing data from Hunan, China. (Top left) One realization of the reconstructed transmission chains, with a histogram representing overdispersion in the distribution of secondary infections. (Top right) Contact matrices of community, social, extended family, and household contacts reveal distinct age profiles. (Bottom) Earlier isolation of primary infections shortens the generation and serial intervals while increasing the relative contribution of transmission in the presymptomatic phase. A long-standing question in infectious disease dynamics concerns the role of transmission heterogeneities, which are driven by demography, behavior, and interventions. On the basis of detailed patient and contact-tracing data in Hunan, China, we find that 80% of secondary infections traced back to 15% of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) primary infections, which indicates substantial transmission heterogeneities. Transmission risk scales positively with the duration of exposure and the closeness of social interactions and is modulated by demographic and clinical factors. The lockdown period increases transmission risk in the family and households, whereas isolation and quarantine reduce risks across all types of contacts. The reconstructed infectiousness profile of a typical SARS-CoV-2 patient peaks just before symptom presentation. Modeling indicates that SARS-CoV-2 control requires the synergistic efforts of case isolation, contact quarantine, and population-level interventions because of the specific transmission kinetics of this virus. [1]: /lookup/doi/10.1126/science.abe2424 [2]: pending:yes
Code Adam Gradient Descent Optimization From Scratch
Gradient descent is an optimization algorithm that follows the negative gradient of an objective function in order to locate the minimum of the function. A limitation of gradient descent is that a single step size (learning rate) is used for all input variables. Extensions to gradient descent like AdaGrad and RMSProp update the algorithm to use a separate step size for each input variable but may result in a step size that rapidly decreases to very small values. The Adaptive Movement Estimation algorithm, or Adam for short, is an extension to gradient descent and a natural successor to techniques like AdaGrad and RMSProp that automatically adapts a learning rate for each input variable for the objective function and further smooths the search process by using an exponentially decreasing moving average of the gradient to make updates to variables. In this tutorial, you will discover how to develop gradient descent with Adam optimization algorithm from scratch.
A Tensor-Based Formulation of Hetero-functional Graph Theory
Farid, Amro M., Thompson, Dakota, Hegde, Prabhat, Schoonenberg, Wester
Recently, hetero-functional graph theory (HFGT) has developed as a means to mathematically model the structure of large flexible engineering systems. In that regard, it intellectually resembles a fusion of network science and model-based systems engineering. With respect to the former, it relies on multiple graphs as data structures so as to support matrix-based quantitative analysis. In the meantime, HFGT explicitly embodies the heterogeneity of conceptual and ontological constructs found in model-based systems engineering including system form, system function, and system concept. At their foundation, these disparate conceptual constructs suggest multi-dimensional rather than two-dimensional relationships. This paper provides the first tensor-based treatment of some of the most important parts of hetero-functional graph theory. In particular, it addresses the "system concept", the hetero-functional adjacency matrix, and the hetero-functional incidence tensor. The tensor-based formulation described in this work makes a stronger tie between HFGT and its ontological foundations in MBSE. Finally, the tensor-based formulation facilitates an understanding of the relationships between HFGT and multi-layer networks.
Hostility Detection and Covid-19 Fake News Detection in Social Media
Gupta, Ayush, Sukumaran, Rohan, John, Kevin, Teki, Sundeep
Withtheadventofsocialmedia,therehasbeenanextremely rapid increase in the content shared online. Consequently, the propagation of fake news and hostile messages on social media platforms has also skyrocketed. In this paper, we address the problem of detecting hostile and fake content in the Devanagari (Hindi) script as a multi-class, multi-label problem. Using NLP techniques, we build a model that makes use of an abusive language detector coupled with features extracted via Hindi BERT and Hindi FastText models and metadata. Our model achieves a 0.97 F1 score on coarse grain evaluation on Hostility detection task. Additionally, we built models to identify fake news related to Covid-19 in English tweets. We leverage entity information extracted from the tweets along with textual representations learned from word embeddings and achieve a 0.93 F1 score on the English fake news detection task.