Statistical Learning
On making optimal transport robust to all outliers
Optimal transport (OT) is known to be sensitive against outliers because of its marginal constraints. Outlier robust OT variants have been proposed based on the definition that outliers are samples which are expensive to move. In this paper, we show that this definition is restricted by considering the case where outliers are closer to the target measure than clean samples. We show that outlier robust OT fully transports these outliers leading to poor performances in practice. To tackle these outliers, we propose to detect them by relying on a classifier trained with adversarial training to classify source and target samples. A sample is then considered as an outlier if the prediction from the classifier is different from its assigned label. To decrease the influence of these outliers in the transport problem, we propose to either remove them from the problem or to increase the cost of moving them by using the classifier prediction. We show that we successfully detect these outliers and that they do not influence the transport problem on several experiments such as gradient flows, generative models and label propagation.
Affinity-Aware Graph Networks
Velingker, Ameya, Sinop, Ali Kemal, Ktena, Ira, Veliฤkoviฤ, Petar, Gollapudi, Sreenivas
Graph Neural Networks (GNNs) have emerged as a powerful technique for learning on relational data. Owing to the relatively limited number of message passing steps they perform -- and hence a smaller receptive field -- there has been significant interest in improving their expressivity by incorporating structural aspects of the underlying graph. In this paper, we explore the use of affinity measures as features in graph neural networks, in particular measures arising from random walks, including effective resistance, hitting and commute times. We propose message passing networks based on these features and evaluate their performance on a variety of node and graph property prediction tasks. Our architecture has lower computational complexity, while our features are invariant to the permutations of the underlying graph. The measures we compute allow the network to exploit the connectivity properties of the graph, thereby allowing us to outperform relevant benchmarks for a wide variety of tasks, often with significantly fewer message passing steps. On one of the largest publicly available graph regression datasets, OGB-LSC-PCQM4Mv1, we obtain the best known single-model validation MAE at the time of writing.
User Engagement in Mobile Health Applications
Olaniyi, Babaniyi Yusuf, del Rรญo, Ana Fernรกndez, Periรกรฑez, รfrica, Bellhouse, Lauren
Mobile health apps are revolutionizing the healthcare ecosystem by improving communication, efficiency, and quality of service. In low- and middle-income countries, they also play a unique role as a source of information about health outcomes and behaviors of patients and healthcare workers, while providing a suitable channel to deliver both personalized and collective policy interventions. We propose a framework to study user engagement with mobile health, focusing on healthcare workers and digital health apps designed to support them in resource-poor settings. The behavioral logs produced by these apps can be transformed into daily time series characterizing each user's activity. We use probabilistic and survival analysis to build multiple personalized measures of meaningful engagement, which could serve to tailor content and digital interventions suiting each health worker's specific needs. Special attention is given to the problem of detecting churn, understood as a marker of complete disengagement. We discuss the application of our methods to the Indian and Ethiopian users of the Safe Delivery App, a capacity-building tool for skilled birth attendants. This work represents an important step towards a full characterization of user engagement in mobile health applications, which can significantly enhance the abilities of health workers and, ultimately, save lives.
Wasserstein t-SNE
Bachmann, Fynn, Hennig, Philipp, Kobak, Dmitry
Scientific datasets often have hierarchical structure: for example, in surveys, individual participants (samples) might be grouped at a higher level (units) such as their geographical region. In these settings, the interest is often in exploring the structure on the unit level rather than on the sample level. Units can be compared based on the distance between their means, however this ignores the within-unit distribution of samples. Here we develop an approach for exploratory analysis of hierarchical datasets using the Wasserstein distance metric that takes into account the shapes of within-unit distributions. We use t-SNE to construct 2D embeddings of the units, based on the matrix of pairwise Wasserstein distances between them. The distance matrix can be efficiently computed by approximating each unit with a Gaussian distribution, but we also provide a scalable method to compute exact Wasserstein distances. We use synthetic data to demonstrate the effectiveness of our Wasserstein t-SNE, and apply it to data from the 2017 German parliamentary election, considering polling stations as samples and voting districts as units.
Backward baselines: Is your model predicting the past?
Hardt, Moritz, Kim, Michael P.
Proponents of predictive technologies for consequential decision-making emphasize the seeming ability of statistical models to anticipate future outcomes. The ability to predict the future, so the argument goes, creates a rationale for adopting machine learning as policy: if a risk score charted the future trajectory of individuals, then intervening in a person's life on the basis of the score would be justified [KLMO15, OE16]. At the same time, critical scholars caution that predictive technologies reproduce historical patterns of injustice and social stratification. In this account, rather than predicting future outcomes, statistical risk assessment tools punish individuals for factors predating their own agency [Eub18, Ben19]. Does a statistical model predict the future or recite the past? The answer to the question is often not obvious. Consider the problem of loan default prediction, one of many tasks often framed as predicting future outcomes. A forward-looking predictor might identify individual behavior detrimental to loan repayment and adjust the predicted likelihood of default accordingly. Alternatively, a backward-looking predictor might take note of historical associations between repayment and demographic factors, then predict based solely on the historical factors.
Invariant Causal Mechanisms through Distribution Matching
Chevalley, Mathieu, Bunne, Charlotte, Krause, Andreas, Bauer, Stefan
Learning representations that capture the underlying data generating process is a key problem for data efficient and robust use of neural networks. One key property for robustness which the learned representation should capture and which recently received a lot of attention is described by the notion of invariance. In this work we provide a causal perspective and new algorithm for learning invariant representations. Empirically we show that this algorithm works well on a diverse set of tasks and in particular we observe state-of-the-art performance on domain generalization, where we are able to significantly boost the score of existing models.
Physics-Informed Statistical Modeling for Wildfire Aerosols Process Using Multi-Source Geostationary Satellite Remote-Sensing Data Streams
Wei, Guanzhou, Krishnan, Venkat, Xie, Yu, Sengupta, Manajit, Zhang, Yingchen, Liao, Haitao, Liu, Xiao
Increasingly frequent wildfires significantly affect solar energy production as the atmospheric aerosols generated by wildfires diminish the incoming solar radiation to the earth. Atmospheric aerosols are measured by Aerosol Optical Depth (AOD), and AOD data streams can be retrieved and monitored by geostationary satellites. However, multi-source remote-sensing data streams often present heterogeneous characteristics, including different data missing rates, measurement errors, systematic biases, and so on. To accurately estimate and predict the underlying AOD propagation process, there exist practical needs and theoretical interests to propose a physics-informed statistical approach for modeling wildfire AOD propagation by simultaneously utilizing, or fusing, multi-source heterogeneous satellite remote-sensing data streams. Leveraging a spectral approach, the proposed approach integrates multi-source satellite data streams with a fundamental advection-diffusion equation that governs the AOD propagation process. A bias correction process is included in the statistical model to account for the bias of the physics model and the truncation error of the Fourier series. The proposed approach is applied to California wildfires AOD data streams obtained from the National Oceanic and Atmospheric Administration. Comprehensive numerical examples are provided to demonstrate the predictive capabilities and model interpretability of the proposed approach. Computer code has been made available on GitHub.
Optimization paper production through digitalization by developing an assistance system for machine operators including quality forecast: a concept
Schroth, Moritz, Hake, Felix, Merker, Konstantin, Becher, Alexander, Klaeger, Tilman, Huesmann, Robin, Eichhorn, Detlef, Oehm, Lukas
Nowadays cross-industry ranging challenges include the reduction of greenhouse gas emission and enabling a circular economy. However, the production of paper from waste paper is still a highly resource intensive task, especially in terms of energy consumption. While paper machines produce a lot of data, we have identified a lack of utilization of it and implement a concept using an operator assistance system and state-of-the-art machine learning techniques, e.g., classification, forecasting and alarm flood handling algorithms, to support daily operator tasks. Our main objective is to provide situation-specific knowledge to machine operators utilizing available data. We expect this will result in better adjusted parameters and therefore a lower footprint of the paper machines.
On the Tractability of SHAP Explanations
Van den Broeck, Guy, Lykov, Anton, Schleich, Maximilian, Suciu, Dan
Shap explanations are a popular feature-attribution mechanism for explainable AI. They use game-theoretic notions to measure the influence of individual features on the prediction of a machine learning model. Despite a lot of recent interest from both academia and industry, it is not known whether Shap explanations of common machine learning models can be computed efficiently. In this paper, we establish the complexity of computing the Shap explanation in three important settings. First, we consider fully-factorized data distributions, and show that the complexity of computing the Shap explanation is the same as the complexity of computing the expected value of the model. This fully-factorized setting is often used to simplify the Shap computation, yet our results show that the computation can be intractable for commonly used models such as logistic regression. Going beyond fully-factorized distributions, we show that computing Shap explanations is already intractable for a very simple setting: computing Shap explanations of trivial classifiers over naive Bayes distributions. Finally, we show that even computing Shap over the empirical distribution is #P-hard.
What is quantum artificial intelligence? - Dataconomy
Quantum artificial intelligence is here to pave the way for the next chapter of our digital intellect pursuit. Artificial intelligence is a transformative technology, and it needs quantum computing to achieve significant improvement. Although artificial intelligence may be used with conventional computers, it is restricted by conventional computational capabilities. Artificial intelligence's capacity to tackle more complex issues can be enhanced by quantum computing, allowing it to solve much more complicated problems. Quantum artificial intelligence allows quantum computing to be used with machine learning algorithms.