Goto

Collaborating Authors

 Model-Based Reasoning


Multifidelity Kolmogorov-Arnold Networks

arXiv.org Artificial Intelligence

In recent years, scientific machine learning (SciML) has emerged as a paradigm for modeling physical systems [1, 2, 3]. Typically using the theory of multilayer perceptrons (MLPs), SciML has shown great success in modeling a wide range of applications, however, data-informed training struggles when high-quality data is not available. Kolmogorov-Arnold networks (KANs) have recently been developed as an alternative to MLPs [4, 5]. KANs use the Kolmogorov-Arnold Theorem as inspiration and can offer advantages over MLPs in some cases, such as for discovering interpretable models. However, KANs have been shown to struggle to reach the accuracy of MLPs, particularly without modifications [6, 7, 8, 9]. In the short time since the publication of [4], many variations of KANs have been developed, including physics-informed KANs (PIKANs)[9], KAN-informed neural networks (KINNs)[10], temporal KANs [11], wavelet KANs [12], graph KANs [13, 14, 15], Chebyshev KANs (cKANs) [16], convolutional KANs [17], ReLU-KANs [18], Higher-order-ReLU-KANs (HRKANs) [19], fractional KANs [20], finite basis KANs [21], deep operator KANs [22], and others.


Physics-Informed Learning for the Friction Modeling of High-Ratio Harmonic Drives

arXiv.org Artificial Intelligence

This paper presents a scalable method for friction identification in robots equipped with electric motors and high-ratio harmonic drives, utilizing Physics-Informed Neural Networks (PINN). This approach eliminates the need for dedicated setups and joint torque sensors by leveraging the robo\v{t}s intrinsic model and state data. We present a comprehensive pipeline that includes data acquisition, preprocessing, ground truth generation, and model identification. The effectiveness of the PINN-based friction identification is validated through extensive testing on two different joints of the humanoid robot ergoCub, comparing its performance against traditional static friction models like the Coulomb-viscous and Stribeck-Coulomb-viscous models. Integrating the identified PINN-based friction models into a two-layer torque control architecture enhances real-time friction compensation. The results demonstrate significant improvements in control performance and reductions in energy losses, highlighting the scalability and robustness of the proposed method, also for application across a large number of joints as in the case of humanoid robots.


A model learning framework for inferring the dynamics of transmission rate depending on exogenous variables for epidemic forecasts

arXiv.org Artificial Intelligence

In this work, we aim to formalize a novel scientific machine learning framework to reconstruct the hidden dynamics of the transmission rate, whose inaccurate extrapolation can significantly impair the quality of the epidemic forecasts, by incorporating the influence of exogenous variables (such as environmental conditions and strain-specific characteristics). We propose an hybrid model that blends a data-driven layer with a physics-based one. The data-driven layer is based on a neural ordinary differential equation that learns the dynamics of the transmission rate, conditioned on the meteorological data and wave-specific latent parameters. The physics-based layer, instead, consists of a standard SEIR compartmental model, wherein the transmission rate represents an input. The learning strategy follows an end-to-end approach: the loss function quantifies the mismatch between the actual numbers of infections and its numerical prediction obtained from the SEIR model incorporating as an input the transmission rate predicted by the neural ordinary differential equation. We validate this original approach using both a synthetic test case and a realistic test case based on meteorological data (temperature and humidity) and influenza data from Italy between 2010 and 2020. In both scenarios, we achieve low generalization error on the test set and observe strong alignment between the reconstructed model and established findings on the influence of meteorological factors on epidemic spread. Finally, we implement a data assimilation strategy to adapt the neural equation to the specific characteristics of an epidemic wave under investigation, and we conduct sensitivity tests on the network hyperparameters.


Analysis of Variance of Multiple Causal Networks

Neural Information Processing Systems

Constructing a directed cyclic graph (DCG) is challenged by both algorithmic difficulty and computational burden. Comparing multiple DCGs is even more difficult, compounded by the need to identify dynamic causalities across graphs. We propose to unify multiple DCGs with a single structural model and develop a limited-information-based method to simultaneously construct multiple networks and infer their disparities, which can be visualized by appropriate correspondence analysis. The algorithm provides DCGs with robust non-asymptotic theoretical properties. It is designed with two sequential stages, each of which involves parallel computation tasks that are scalable to the network complexity.


PDEBench: An Extensive Benchmark for Scientific Machine Learning

Neural Information Processing Systems

Machine learning-based modeling of physical systems has experienced increased interest in recent years. Despite some impressive progress, there is still a lack of benchmarks for Scientific ML that are easy to use but still challenging and repre- sentative of a wide range of problems. We introduce PDEBENCH, a benchmark suite of time-dependent simulation tasks based on Partial Differential Equations (PDEs). PDEBENCH comprises both code and data to benchmark the performance of novel machine learning models against both classical numerical simulations and machine learning baselines. Our proposed set of benchmark problems con- tribute the following unique features: (1) A much wider range of PDEs compared to existing benchmarks, ranging from relatively common examples to more real- istic and difficult problems; (2) much larger ready-to-use datasets compared to prior work, comprising multiple simulation runs across a larger number of ini- tial and boundary conditions and PDE parameters; (3) more extensible source codes with user-friendly APIs for data generation and baseline results with popular machine learning models (FNO, U-Net, PINN, Gradient-Based Inverse Method).


On the Robustness of Mechanism Design under Total Variation Distance

Neural Information Processing Systems

We study the problem of designing mechanisms when agents' valuation functions are drawn from unknown and correlated prior distributions. In particular, we are given a prior distribution D, and we are interested in designing a (truthful) mechanism that has good performance for all "true distributions" that are close to D in Total Variation (TV) distance. We show that DSIC and BIC mechanisms in this setting are strongly robust with respect to TV distance, for any bounded objective function \mathcal{O}, extending a recent result of Brustle et al. ([BCD20], EC 2020). At the heart of our result is a fundamental duality property of total variation distance. As direct applications of our result, we (i) demonstrate how to find approximately revenue-optimal and approximately BIC mechanisms for weakly dependent prior distributions; (ii) show how to find correlation-robust mechanisms when only noisy'' versions of marginals are accessible, extending recent results of Bei et.


Causal Effect Identification in Uncertain Causal Networks

Neural Information Processing Systems

Causal identification is at the core of the causal inference literature, where complete algorithms have been proposed to identify causal queries of interest. The validity of these algorithms hinges on the restrictive assumption of having access to a correctly specified causal structure. In this work, we study the setting where a probabilistic model of the causal structure is available. Specifically, the edges in a causal graph exist with uncertainties which may, for example, represent degree of belief from domain experts. Alternatively, the uncertainty about an edge may reflect the confidence of a particular statistical test. The question that naturally arises in this setting is: Given such a probabilistic graph and a specific causal effect of interest, what is the subgraph which has the highest plausibility and for which the causal effect is identifiable?


Reviews: Avoiding Discrimination through Causal Reasoning

Neural Information Processing Systems

This paper formulates fairness for protected attributes as a causal inference problem. It is a well-written paper that makes an important connection that discrimination, in itself, is a causal concept that can benefit from using causal inference explicitly. The major technical contribution of the paper is to show a method using causal graphical models for ensuring fairness in a predictive model. In line with social science research on discrimination, they steer away from thinking about a gender or race counterfactual which can be an unwieldy or even absurd counterfactual to construct---what would happen if someone's race is changed---and instead focus on using proxy variables and controlling causal pathways from proxy variables to an outcome. This is an important distinction, and allows us to declare proxy variables such as home location or name and use them for minimizing discrimination.


Reviews: Inference Aided Reinforcement Learning for Incentive Mechanism Design in Crowdsourcing

Neural Information Processing Systems

Summary: In this paper, the authors explore the problem of data collecting using crowdsourcing. In the setting of the paper, each task is a labeling task with binary labels, and workers are strategic in choosing effort levels and reporting strategies that maximize their utility. The true label for each task and workers' parameters are all unknown to the requester. The requester's goal is to learn how to decide the payment and how to aggregate the collected labels by learning from workers' past answers. The authors' proposed approach is a combination of incentive design, Bayesian inference, and reinforcement learning.


Reviews: Data center cooling using model-predictive control

Neural Information Processing Systems

This paper addresses the problem of temperature and airflow regulation for a large-scale data center and considers how a data-driven, model-based approach using Reinforcement Learning (RL) might improve operational efficiency relative to the existing approach of hand-crafted PID controllers. Existing controllers in large-scale data centers tend to be simple, conservative and hand-tuned to physical equipment layouts and configurations. Safety constraints and a low tolerance for performance degradation and equipment damage impose additional constraints. The authors use model-predictive control (MPC) to learn a linear model of the data center dynamics (a LQ controller) using safe, random exploration, starting with little or no prior knowledge. They then determine the control actions at each time step by optimizing the cost of the model-predicted trajectories, ensuring to re-optimize at each time step.