Goto

Collaborating Authors

 Uncertainty


Stochastic Thermodynamics of Learning Parametric Probabilistic Models

arXiv.org Artificial Intelligence

Starting from nearly half a century ago, physicists began to learn that information is a physical entity [1, 2, 3]. Today, the information-theoretic perspective has significantly impacted various fields of physics, including quantum computing [4], cosmology [5], and thermodynamics [6]. Simultaneously, recent years have witnessed the remarkable success of an algorithmic approach known as machine learning, which is adept at learning information from data. This paper is propelled by a straightforward proposition: if "information is physical", then the process of learning information must inherently be a physical process. The concepts of memory, prediction, and information exchange between subsystems have undergone extensive exploration within the realms of Thermodynamics of Information [6] and Stochastic Thermodynamics [7]. For instance, Still et al. [8] delved into the thermodynamics of prediction. And, the role of information exchange between thermodynamic subsystems has been studied by Sagawa and Ueda [9], and Esposito et al. [10]. This rich toolbox of thermodynamic of information is our main venue to study physics of machine learning process, with motivation to assess the information content of the learning process. The type of machine learning problems we consider in this study encompasses any algorithmic approach that evolves a Parametric Probabilistic Model (PPM), or simply the model, towards a desirable target distribution through gradientbased loss function minimization.


Improved DDIM Sampling with Moment Matching Gaussian Mixtures

arXiv.org Artificial Intelligence

W e propose using a Gaussian Mixture Model (GMM) as reverse tr ansition operator (kernel) within the Denoising Diffusion Implicit Model s (DDIM) framework, which is one of the most widely used approaches for accelerat ed sampling from pre-trained Denoising Diffusion Probabilistic Models (DD PM). Specifically we match the first and second order central moments of the DDPM fo rward marginals by constraining the parameters of the GMM. W e see that moment matching is sufficient to obtain samples with equal or better quality than th e original DDIM with Gaussian kernels. W e provide experimental results with unc onditional models trained on CelebAHQ and FFHQ and class-conditional models t rained on ImageNet datasets respectively. Our results suggest that usin g the GMM kernel leads to significant improvements in the quality of the generated s amples when the number of sampling steps is small, as measured by FID and IS metri cs. For example on ImageNet 256x256, using 10 sampling steps, we achieve a FI D of 6.94 and IS of 207.85 with a GMM kernel compared to 10.15 and 196.73 respe ctively with a Gaussian kernel. In spite of their success, the main bottleneck to their adoption is th e slow sampling speed, usually requiring hundreds to thousands of denoising steps to generat e a sample. Denoising Diffusion Implicit Models (DDIM) (Song et al., 20 21) accelerate sampling from Denois-ing Diffusion Probabilistic Models (DDPM) (Ho et al., 2020) by hypothesizing a family of non-Markovian forward processes, whose reverse process (Marko vian) estimators can be trained with the same surrogate objective as DDPMs, assuming the same par ameterization for reverse estimators. In other words, one can sample with a pretrained DDPM denoiser by designing a dif ferent forward/backward process than the original DDPM given that the forward marginals are t he same.


Score-based Source Separation with Applications to Digital Communication Signals

arXiv.org Artificial Intelligence

We propose a new method for separating superimposed sources using diffusion-based generative models. Our method relies only on separately trained statistical priors of independent sources to establish a new objective function guided by maximum a posteriori estimation with an $\alpha$-posterior, across multiple levels of Gaussian smoothing. Motivated by applications in radio-frequency (RF) systems, we are interested in sources with underlying discrete nature and the recovery of encoded bits from a signal of interest, as measured by the bit error rate (BER). Experimental results with RF mixtures demonstrate that our method results in a BER reduction of 95% over classical and existing learning-based methods. Our analysis demonstrates that our proposed method yields solutions that asymptotically approach the modes of an underlying discrete distribution. Furthermore, our method can be viewed as a multi-source extension to the recently proposed score distillation sampling scheme, shedding additional light on its use beyond conditional sampling. The project webpage is available at https://alpha-rgs.github.io


Learning from Sparse Offline Datasets via Conservative Density Estimation

arXiv.org Artificial Intelligence

Offline reinforcement learning (RL) offers a promising direction for learning policies from pre-collected datasets without requiring further interactions with the environment. However, existing methods struggle to handle out-of-distribution (OOD) extrapolation errors, especially in sparse reward or scarce data settings. In this paper, we propose a novel training algorithm called Conservative Density Estimation (CDE), which addresses this challenge by explicitly imposing constraints on the state-action occupancy stationary distribution. CDE overcomes the limitations of existing approaches, such as the stationary distribution correction method, by addressing the support mismatch issue in marginal importance sampling. Our method achieves state-of-the-art performance on the D4RL benchmark. Notably, CDE consistently outperforms baselines in challenging tasks with sparse rewards or insufficient data, demonstrating the advantages of our approach in addressing the extrapolation error problem in offline RL.


The Impact of Differential Feature Under-reporting on Algorithmic Fairness

arXiv.org Artificial Intelligence

Predictive risk models in the public sector are commonly developed using administrative data that is more complete for subpopulations that more greatly rely on public services. In the United States, for instance, information on health care utilization is routinely available to government agencies for individuals supported by Medicaid and Medicare, but not for the privately insured. Critiques of public sector algorithms have identified such differential feature under-reporting as a driver of disparities in algorithmic decision-making. Yet this form of data bias remains understudied from a technical viewpoint. While prior work has examined the fairness impacts of additive feature noise and features that are clearly marked as missing, the setting of data missingness absent indicators (i.e. differential feature under-reporting) has been lacking in research attention. In this work, we present an analytically tractable model of differential feature under-reporting which we then use to characterize the impact of this kind of data bias on algorithmic fairness. We demonstrate how standard missing data methods typically fail to mitigate bias in this setting, and propose a new set of methods specifically tailored to differential feature under-reporting. Our results show that, in real world data settings, under-reporting typically leads to increasing disparities. The proposed solution methods show success in mitigating increases in unfairness.


Enhancing Dynamical System Modeling through Interpretable Machine Learning Augmentations: A Case Study in Cathodic Electrophoretic Deposition

arXiv.org Artificial Intelligence

We introduce a comprehensive data-driven framework aimed at enhancing the modeling of physical systems, employing inference techniques and machine learning enhancements. As a demonstrative application, we pursue the modeling of cathodic electrophoretic deposition (EPD), commonly known as e-coating. Our approach illustrates a systematic procedure for enhancing physical models by identifying their limitations through inference on experimental data and introducing adaptable model enhancements to address these shortcomings. We begin by tackling the issue of model parameter identifiability, which reveals aspects of the model that require improvement. To address generalizability , we introduce modifications which also enhance identifiability. However, these modifications do not fully capture essential experimental behaviors. To overcome this limitation, we incorporate interpretable yet flexible augmentations into the baseline model. These augmentations are parameterized by simple fully-connected neural networks (FNNs), and we leverage machine learning tools, particularly Neural Ordinary Differential Equations (Neural ODEs), to learn these augmentations. Our simulations demonstrate that the machine learning-augmented model more accurately captures observed behaviors and improves predictive accuracy. Nevertheless, we contend that while the model updates offer superior performance and capture the relevant physics, we can reduce off-line computational costs by eliminating certain dynamics without compromising accuracy or interpretability in downstream predictions of quantities of interest, particularly film thickness predictions. The entire process outlined here provides a structured approach to leverage data-driven methods. Firstly, it helps us comprehend the root causes of model inaccuracies, and secondly, it offers a principled method for enhancing model performance.


Personalized Federated Learning of Probabilistic Models: A PAC-Bayesian Approach

arXiv.org Artificial Intelligence

Federated learning aims to infer a shared model from private and decentralized data stored locally by multiple clients. Personalized federated learning (PFL) goes one step further by adapting the global model to each client, enhancing the model's fit for different clients. A significant level of personalization is required for highly heterogeneous clients, but can be challenging to achieve especially when they have small datasets. To address this problem, we propose a PFL algorithm named PAC-PFL for learning probabilistic models within a PAC-Bayesian framework that utilizes differential privacy to handle data-dependent priors. Our algorithm collaboratively learns a shared hyper-posterior and regards each client's posterior inference as the personalization step. By establishing and minimizing a generalization bound on the average true risk of clients, PAC-PFL effectively combats over-fitting. PACPFL achieves accurate and well-calibrated predictions, supported by experiments on a dataset of photovoltaic panel power generation, FEMNIST dataset (Caldas et al., 2019), and Dirichlet-partitioned EMNIST dataset (Cohen et al., 2017).


Machine Learning on Dynamic Graphs: A Survey on Applications

arXiv.org Artificial Intelligence

Dynamic graph learning has gained significant attention as it offers a powerful means to model intricate interactions among entities across various real-world and scientific domains. Notably, graphs serve as effective representations for diverse networks such as transportation, brain, social, and internet networks. Furthermore, the rapid advancements in machine learning have expanded the scope of dynamic graph applications beyond the aforementioned domains. In this paper, we present a review of lesser-explored applications of dynamic graph learning. This study revealed the potential of machine learning on dynamic graphs in addressing challenges across diverse domains, including those with limited levels of association with the field.


How to Turn Your Knowledge Graph Embeddings into Generative Models

arXiv.org Artificial Intelligence

Some of the most successful knowledge graph embedding (KGE) models for link prediction -- CP, RESCAL, TuckER, ComplEx -- can be interpreted as energy-based models. Under this perspective they are not amenable for exact maximum-likelihood estimation (MLE), sampling and struggle to integrate logical constraints. This work re-interprets the score functions of these KGEs as circuits -- constrained computational graphs allowing efficient marginalisation. Then, we design two recipes to obtain efficient generative circuit models by either restricting their activations to be non-negative or squaring their outputs. Our interpretation comes with little or no loss of performance for link prediction, while the circuits framework unlocks exact learning by MLE, efficient sampling of new triples, and guarantee that logical constraints are satisfied by design. Furthermore, our models scale more gracefully than the original KGEs on graphs with millions of entities.


Explainable Predictive Maintenance: A Survey of Current Methods, Challenges and Opportunities

arXiv.org Artificial Intelligence

Predictive maintenance is a well studied collection of techniques that aims to prolong the life of a mechanical system by using artificial intelligence and machine learning to predict the optimal time to perform maintenance. The methods allow maintainers of systems and hardware to reduce financial and time costs of upkeep. As these methods are adopted for more serious and potentially life-threatening applications, the human operators need trust the predictive system. This attracts the field of Explainable AI (XAI) to introduce explainability and interpretability into the predictive system. XAI brings methods to the field of predictive maintenance that can amplify trust in the users while maintaining well-performing systems. This survey on explainable predictive maintenance (XPM) discusses and presents the current methods of XAI as applied to predictive maintenance while following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. We categorize the different XPM methods into groups that follow the XAI literature. Additionally, we include current challenges and a discussion on future research directions in XPM.