Statistical Learning
Learning Effective SDEs from Brownian Dynamics Simulations of Colloidal Particles
Evangelou, Nikolaos, Dietrich, Felix, Bello-Rivas, Juan M., Yeh, Alex, Stein, Rachel, Bevan, Michael A., Kevrekidis, Ioannis G.
The identification of nonlinear dynamical systems from experimental time series and image series data became an important research theme in the early 1990s [25, 37, 36]. After lapsing for almost two decades, it is now experiencing a spectacular rebirth. A key element of the older work was the use of neural architectures [14, 37] (recurrent, convolutional, ResNet) motivated by traditional numerical analysis algorithms. Importantly, such architectures allow researchers to identify effective, coarse-grained, mean-field type evolution models from fine-scale (atomistic, molecular, agent-based) data [29, 5]. In this paper, we identify coarse-grained, effective stochastic differential equations (eSDE) for colloidal particle selfassembly based onfine-grained, Brownian dynamics simulations under the influence of electric fields [51, 11]. We demonstrate that the identified eSDE encodes accurately the physics of the Brownian Dynamic simulations and captures the dynamics of corresponding experimental data. Those experiments have previously been shown to quantitatively match to BD simulations at equilibrium in terms of time-averaged distribution functions [11, 18, 20]. Figure 1 shows a sample path of a latent space trajectory {t, ฯ(t)}
Looped Transformers as Programmable Computers
Giannou, Angeliki, Rajput, Shashank, Sohn, Jy-yong, Lee, Kangwook, Lee, Jason D., Papailiopoulos, Dimitris
We present a framework for using transformer networks as universal computers by programming them with specific weights and placing them in a loop. Our input sequence acts as a punchcard, consisting of instructions and memory for data read/writes. We demonstrate that a constant number of encoder layers can emulate basic computing blocks, including embedding edit operations, non-linear functions, function calls, program counters, and conditional branches. Using these building blocks, we emulate a small instruction-set computer. This allows us to map iterative algorithms to programs that can be executed by a looped, 13-layer transformer. We show how this transformer, instructed by its input, can emulate a basic calculator, a basic linear algebra library, and in-context learning algorithms that employ backpropagation. Our work highlights the versatility of the attention mechanism, and demonstrates that even shallow transformers can execute full-fledged, general-purpose programs.
Generating Synthetic Mixed-type Longitudinal Electronic Health Records for Artificial Intelligent Applications
Li, Jin, Cairns, Benjamin J., Li, Jingsong, Zhu, Tingting
The recent availability of electronic health records (EHRs) have provided enormous opportunities to develop artificial intelligence (AI) algorithms. However, patient privacy has become a major concern that limits data sharing across hospital settings and subsequently hinders the advances in AI. Synthetic data, which benefits from the development and proliferation of generative models, has served as a promising substitute for real patient EHR data. However, the current generative models are limited as they only generate single type of clinical data for a synthetic patient, i.e., either continuous-valued or discrete-valued. To mimic the nature of clinical decision-making which encompasses various data types/sources, in this study, we propose a generative adversarial network (GAN) entitled EHR-M-GAN which simultaneously synthesizes mixed-type timeseries EHR data. EHR-M-GAN is capable of capturing the multidimensional, heterogeneous, and correlated temporal dynamics in patient trajectories. We have validated EHR-M-GAN on three publicly-available intensive care unit databases with records from a total of 141,488 unique patients, and performed privacy risk evaluation of the proposed model. EHR-M-GAN has demonstrated its superiority over state-of-the-art benchmarks for synthesizing clinical timeseries with high fidelity, while addressing the limitations regarding data types and dimensionality in the current generative models. Notably, prediction models for outcomes of intensive care performed significantly better when training data was augmented with the addition of EHR-M-GAN-generated timeseries. EHR-M-GAN may have use in developing AI algorithms in resource-limited settings, lowering the barrier for data acquisition while preserving patient privacy.
Hierarchical learning, forecasting coherent spatio-temporal individual and aggregated building loads
Leprince, Julien, Madsen, Henrik, Mรธller, Jan Kloppenborg, Zeiler, Wim
Optimal decision-making compels us to anticipate the future at different horizons. However, in many domains connecting together predictions from multiple time horizons and abstractions levels across their organization becomes all the more important, else decision-makers would be planning using separate and possibly conflicting views of the future. This notably applies to smart grid operation. To optimally manage energy flows in such systems, accurate and coherent predictions must be made across varying aggregation levels and horizons. With this work, we propose a novel multi-dimensional hierarchical forecasting method built upon structurally-informed machine-learning regressors and established hierarchical reconciliation taxonomy. A generic formulation of multi-dimensional hierarchies, reconciling spatial and temporal hierarchies under a common frame is initially defined. Next, a coherency-informed hierarchical learner is developed built upon a custom loss function leveraging optimal reconciliation methods. Coherency of the produced hierarchical forecasts is then secured using similar reconciliation technics. The outcome is a unified and coherent forecast across all examined dimensions. The method is evaluated on two different case studies to predict building electrical loads across spatial, temporal, and spatio-temporal hierarchies. Although the regressor natively profits from computationally efficient learning, results displayed disparate performances, demonstrating the value of hierarchical-coherent learning in only one setting. Yet, supported by a comprehensive result analysis, existing obstacles were clearly delineated, presenting distinct pathways for future work. Overall, the paper expands and unites traditionally disjointed hierarchical forecasting methods providing a fertile route toward a novel generation of forecasting regressors.
Probabilistic Neural Data Fusion for Learning from an Arbitrary Number of Multi-fidelity Data Sets
Mora, Carlos, Eweis-Labolle, Jonathan Tammer, Johnson, Tyler, Gadde, Likith, Bostanabad, Ramin
In many applications in engineering and sciences analysts have simultaneous access to multiple data sources. In such cases, the overall cost of acquiring information can be reduced via data fusion or multi-fidelity (MF) modeling where one leverages inexpensive low-fidelity (LF) sources to reduce the reliance on expensive high-fidelity (HF) data. In this paper, we employ neural networks (NNs) for data fusion in scenarios where data is very scarce and obtained from an arbitrary number of sources with varying levels of fidelity and cost. We introduce a unique NN architecture that converts MF modeling into a nonlinear manifold learning problem. Our NN architecture inversely learns non-trivial (e.g., non-additive and non-hierarchical) biases of the LF sources in an interpretable and visualizable manifold where each data source is encoded via a low-dimensional distribution. This probabilistic manifold quantifies model form uncertainties such that LF sources with small bias are encoded close to the HF source. Additionally, we endow the output of our NN with a parametric distribution not only to quantify aleatoric uncertainties, but also to reformulate the network's loss function based on strictly proper scoring rules which improve robustness and accuracy on unseen HF data. Through a set of analytic and engineering examples, we demonstrate that our approach provides a high predictive power while quantifying various sources uncertainties.
ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency
Ren, Pengzhen, Li, Changlin, Xu, Hang, Zhu, Yi, Wang, Guangrun, Liu, Jianzhuang, Chang, Xiaojun, Liang, Xiaodan
Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic alignment, while ignoring the correspondence among multiple augmented views of the same image. To overcome such limitation, we propose multi-\textbf{View} \textbf{Co}nsistent learning (ViewCo) for text-supervised semantic segmentation. Specifically, we first propose text-to-views consistency modeling to learn correspondence for multiple views of the same input image. Additionally, we propose cross-view segmentation consistency modeling to address the ambiguity issue of text supervision by contrasting the segment features of Siamese visual encoders. The text-to-views consistency benefits the dense assignment of the visual features by encouraging different crops to align with the same text, while the cross-view segmentation consistency modeling provides additional self-supervision, overcoming the limitation of ambiguous text supervision for segmentation masks. Trained with large-scale image-text data, our model can directly segment objects of arbitrary categories in a zero-shot manner. Extensive experiments show that ViewCo outperforms state-of-the-art methods on average by up to 2.9\%, 1.6\%, and 2.4\% mIoU on PASCAL VOC2012, PASCAL Context, and COCO, respectively.
Improved machine learning algorithm for predicting ground state properties
Lewis, Laura, Huang, Hsin-Yuan, Tran, Viet T., Lehner, Sebastian, Kueng, Richard, Preskill, John
Finding the ground state of a quantum many-body system is a fundamental problem in quantum physics. In this work, we give a classical machine learning (ML) algorithm for predicting ground state properties with an inductive bias encoding geometric locality. The proposed ML model can efficiently predict ground state properties of an $n$-qubit gapped local Hamiltonian after learning from only $\mathcal{O}(\log(n))$ data about other Hamiltonians in the same quantum phase of matter. This improves substantially upon previous results that require $\mathcal{O}(n^c)$ data for a large constant $c$. Furthermore, the training and prediction time of the proposed ML model scale as $\mathcal{O}(n \log n)$ in the number of qubits $n$. Numerical experiments on physical systems with up to 45 qubits confirm the favorable scaling in predicting ground state properties using a small training dataset.
EDSA-Ensemble: an Event Detection Sentiment Analysis Ensemble Architecture
Petrescu, Alexandru, Truicฤ, Ciprian-Octavian, Apostol, Elena-Simona, Paschke, Adrian
As social media platforms grow more and more each day, it also increases the need to analyze and understand certain aspects, such as the impact of important or spiking topics over the network[49]. Event Detection techniques are used to automatically identify important or spiking topics by analysing social media data. In this paper, we use the angle of the positive emotion generated by these topics for the users and the magnitude, both reach and time span, in order to better understand what is happening on social media platforms, mainly Twitter. Sentiment Analysis is a field in Natural Language Processing that analyzes user opinions and emotions from written language [38, 66], while Event Detection deals with analyzing information diffusion in graph networks [24]. Although there is a large volume of work done on Event Detection using social media data and on Sentiment Analysis of this type of content, in the current literature, there is a shortcoming of the approaches that combine the two domains. There are multiple communities that are involved in mining, gathering, and giving some meaning to the vast amount of content generated daily by the users of those platforms, namely the Network Analysis and Natural Language Processing communities. The two communities are using different types of approaches since they have different purposes: For the Network Analysis community, the main purpose is developing methods to deal with the spread and mitigation of harmful content using Event Detection. Event Detection is used to detect the impact and spread of topics on Social Networks using multiple types of approaches such as sliding windows, topic detection, etc.
Do Gradient Inversion Attacks Make Federated Learning Unsafe?
Hatamizadeh, Ali, Yin, Hongxu, Molchanov, Pavlo, Myronenko, Andriy, Li, Wenqi, Dogra, Prerna, Feng, Andrew, Flores, Mona G., Kautz, Jan, Xu, Daguang, Roth, Holger R.
Federated learning (FL) allows the collaborative training of AI models without needing to share raw data. This capability makes it especially interesting for healthcare applications where patient and data privacy is of utmost concern. However, recent works on the inversion of deep neural networks from model gradients raised concerns about the security of FL in preventing the leakage of training data. In this work, we show that these attacks presented in the literature are impractical in FL use-cases where the clients' training involves updating the Batch Normalization (BN) statistics and provide a new baseline attack that works for such scenarios. Furthermore, we present new ways to measure and visualize potential data leakage in FL. Our work is a step towards establishing reproducible methods of measuring data leakage in FL and could help determine the optimal tradeoffs between privacy-preserving techniques, such as differential privacy, and model accuracy based on quantifiable metrics. Code is available at https://nvidia.github.io/NVFlare/research/quantifying-data-leakage.
Don't Explain Noise: Robust Counterfactuals for Randomized Ensembles
Forel, Alexandre, Parmentier, Axel, Vidal, Thibaut
Counterfactual explanations describe how to modify a feature vector in order to flip the outcome of a trained classifier. Obtaining robust counterfactual explanations is essential to provide valid algorithmic recourse and meaningful explanations. We study the robustness of explanations of randomized ensembles, which are always subject to algorithmic uncertainty even when the training data is fixed. We formalize the generation of robust counterfactual explanations as a probabilistic problem and show the link between the robustness of ensemble models and the robustness of base learners. We develop a practical method with good empirical performance and support it with theoretical guarantees for ensembles of convex base learners. Our results show that existing methods give surprisingly low robustness: the validity of naive counterfactuals is below $50\%$ on most data sets and can fall to $20\%$ on problems with many features. In contrast, our method achieves high robustness with only a small increase in the distance from counterfactual explanations to their initial observations.