Statistical Learning
Improved Bound for Mixing Time of Parallel Tempering
A key problem in statistics, computer science, and statistical physics is to draw samples given access to the probability density function, up to a constant of proportionality. Because it is often hard to draw independent samples from the target distribution directly, Markov Chain Monte Carlo(MCMC) methods are often used instead. However, a common difficulty for typical MCMC methods is that for strongly multimodal distributions, MCMC methods take unreasonably long time to reach stationarity. Parallel tempering is an MCMC algorithm that is widely used in sampling from multimodal distributions. Though highly effective in practice, theoretical guarantees on its performance are limited. Since large spectral gap implies fast mixing, a common way to obtain an upper bound on mixing time is to obtain a lower bound on spectral gap.
Learning from data with structured missingness
Mitra, Robin, McGough, Sarah F., Chakraborti, Tapabrata, Holmes, Chris, Copping, Ryan, Hagenbuch, Niels, Biedermann, Stefanie, Noonan, Jack, Lehmann, Brieuc, Shenvi, Aditi, Doan, Xuan Vinh, Leslie, David, Bianconi, Ginestra, Sanchez-Garcia, Ruben, Davies, Alisha, Mackintosh, Maxine, Andrinopoulou, Eleni-Rosalina, Basiri, Anahid, Harbron, Chris, MacArthur, Ben D.
Missing data are an unavoidable complication in many machine learning tasks. When data are `missing at random' there exist a range of tools and techniques to deal with the issue. However, as machine learning studies become more ambitious, and seek to learn from ever-larger volumes of heterogeneous data, an increasingly encountered problem arises in which missing values exhibit an association or structure, either explicitly or implicitly. Such `structured missingness' raises a range of challenges that have not yet been systematically addressed, and presents a fundamental hindrance to machine learning at scale. Here, we outline the current literature and propose a set of grand challenges in learning from data with structured missingness.
Clustering Social Touch Gestures for Human-Robot Interaction
Chahine, Ramzi Abou, Vasquez, Steven, Fazli, Pooyan, Seifi, Hasti
Social touch provides a rich non-verbal communication channel between humans and robots. Prior work has identified a set of touch gestures for human-robot interaction and described them with natural language labels (e.g., stroking, patting). Yet, no data exists on the semantic relationships between the touch gestures in users' minds. To endow robots with touch intelligence, we investigated how people perceive the similarities of social touch labels from the literature. In an online study, 45 participants grouped 36 social touch labels based on their perceived similarities and annotated their groupings with descriptive names. We derived quantitative similarities of the gestures from these groupings and analyzed the similarities using hierarchical clustering. The analysis resulted in 9 clusters of touch gestures formed around the social, emotional, and contact characteristics of the gestures. We discuss the implications of our results for designing and evaluating touch sensing and interactions with social robots.
Lidar based 3D Tracking and State Estimation of Dynamic Objects
Suresh, Patil Shubham, Narasimhan, Gautham Narayan
Generally these values are under constrained and rely 3D point cloud data obtained using Lidar sensor has been on high co-variance for doing trajectory prediction. Prediction very crucial in 3D localization and mapping of static objects is currently one of the hardest problem in autonomous within a scene. For Autonomous vehicles, Lidar point cloud vehicles due to lack of state information like position, velocity, is fused with camera to help in detecting objects like cars acceleration, yaw, yaw rate to correctly model oncoming and pedestrians. However, it hasn't been used to determine vehicle future trajectory as these cannot be determined accurately the dynamic states like velocity, yaw, yaw rate, etc of nonego using just visual camera data or constant velocity objects. The rich positional data obtained using Lidar can be models.
Adaptive Defective Area Identification in Material Surface Using Active Transfer Learning-based Level Set Estimation
Hozumi, Shota, Kutsukake, Kentaro, Matsui, Kota, Kusakawa, Syunya, Ujihara, Toru, Takeuchi, Ichiro
In material characterization, identifying defective areas on a material surface is fundamental. The conventional approach involves measuring the relevant physical properties point-by-point at the predetermined mesh grid points on the surface and determining the area at which the property does not reach the desired level. To identify defective areas more efficiently, we propose adaptive mapping methods in which measurement resources are used preferentially to detect the boundaries of defective areas. We interpret this problem as an active-learning (AL) of the level set estimation (LSE) problem. The goal of AL-based LSE is to determine the level set of the physical property function defined on the surface with as small number of measurements as possible. Furthermore, to handle the situations in which materials with similar specifications are repeatedly produced, we introduce a transfer learning approach so that the information of previously produced materials can be effectively utilized. As a proof-of-concept, we applied the proposed methods to the red-zone estimation problem of silicon wafers and demonstrated that we could identify the defective areas with significantly lower measurement costs than those of conventional methods.
Matched Machine Learning: A Generalized Framework for Treatment Effect Inference With Learned Metrics
Morucci, Marco, Rudin, Cynthia, Volfovsky, Alexander
We introduce Matched Machine Learning, a framework that combines the flexibility of machine learning black boxes with the interpretability of matching, a longstanding tool in observational causal inference. Interpretability is paramount in many high-stakes application of causal inference. Current tools for nonparametric estimation of both average and individualized treatment effects are black-boxes that do not allow for human auditing of estimates. Our framework uses machine learning to learn an optimal metric for matching units and estimating outcomes, thus achieving the performance of machine learning black-boxes, while being interpretable. Our general framework encompasses several published works as special cases. We provide asymptotic inference theory for our proposed framework, enabling users to construct approximate confidence intervals around estimates of both individualized and average treatment effects. We show empirically that instances of Matched Machine Learning perform on par with black-box machine learning methods and better than existing matching methods for similar problems. Finally, in our application we show how Matched Machine Learning can be used to perform causal inference even when covariate data are highly complex: we study an image dataset, and produce high quality matches and estimates of treatment effects.
Artificial neural networks and time series of counts: A class of nonlinear INGARCH models
Time series of counts are frequently analyzed using generalized integer-valued autoregressive models with conditional heteroskedasticity (INGARCH). These models employ response functions to map a vector of past observations and past conditional expectations to the conditional expectation of the present observation. In this paper, it is shown how INGARCH models can be combined with artificial neural network (ANN) response functions to obtain a class of nonlinear INGARCH models. The ANN framework allows for the interpretation of many existing INGARCH models as a degenerate version of a corresponding neural model. Details on maximum likelihood estimation, marginal effects and confidence intervals are given. The empirical analysis of time series of bounded and unbounded counts reveals that the neural INGARCH models are able to outperform reasonable degenerate competitor models in terms of the information loss.
Online stochastic Newton methods for estimating the geometric median and applications
Godichon-Baggioni, Antoine, Lu, Wei
In the context of large samples, a small number of individuals might spoil basic statistical indicators like the mean. It is difficult to detect automatically these atypical individuals, and an alternative strategy is using robust approaches. This paper focuses on estimating the geometric median of a random variable, which is a robust indicator of central tendency. In order to deal with large samples of data arriving sequentially, online stochastic Newton algorithms for estimating the geometric median are introduced and we give their rates of convergence. Since estimates of the median and those of the Hessian matrix can be recursively updated, we also determine confidences intervals of the median in any designated direction and perform online statistical tests.
Negativity Spreads Faster: A Large-Scale Multilingual Twitter Analysis on the Role of Sentiment in Political Communication
Antypas, Dimosthenis, Preece, Alun, Camacho-Collados, Jose
Social media has become extremely influential when it comes to policy making in modern societies, especially in the western world, where platforms such as Twitter allow users to follow politicians, thus making citizens more involved in political discussion. In the same vein, politicians use Twitter to express their opinions, debate among others on current topics and promote their political agendas aiming to influence voter behaviour. In this paper, we attempt to analyse tweets of politicians from three European countries and explore the virality of their tweets. Previous studies have shown that tweets conveying negative sentiment are likely to be retweeted more frequently. By utilising state-of-the-art pre-trained language models, we performed sentiment analysis on hundreds of thousands of tweets collected from members of parliament in Greece, Spain and the United Kingdom, including devolved administrations. We achieved this by systematically exploring and analysing the differences between influential and less popular tweets. Our analysis indicates that politicians' negatively charged tweets spread more widely, especially in more recent times, and highlights interesting differences between political parties as well as between politicians and the general population.