Markov Models
Neural Markov Prolog
Thomson, Alexander, Page, David
Neural network performance has made great strides in recent years by incorporating key assumptions, often referred to as inductive biases, about data domains into specialized model structures. The designs of popular neural network architectures such as recurrent neural networks, convolutional neural networks, graph neural networks, and transformers all incorporate aspects of their respective task-specific domains into the operations, weight sharing, and connections of their underlying network structure [1, 3, 4, 9, 12]. That specialization, has, in turn, yielded improved efficiency and performance over the more general, fully connected design. Note, however, when implemented, these neural networks tend to be treated as entirely separate architectures, with no explicit connections between them, despite their similar underlying assumptions. Not only does this practice obscures some of the core theoretical similarities between these models, but it can also make modifying the architecture cumbersome when any of those original assumptions about the task domain change even slightly. There exist several well-established methods for describing and reasoning from logical knowledge bases that could trivially describe both the assumptions made on a task's domain and the graphical structure of the neural network itself. Nonetheless, simply using deterministic logic on its own to define that structure, through any given logical programming language, does not immediately align with the constrained structure of the neural network and the uncertainty present in said network's predictions.
Learning Multimodal Latent Dynamics for Human-Robot Interaction
Prasad, Vignesh, Heitlinger, Lea, Koert, Dorothea, Stock-Homburg, Ruth, Peters, Jan, Chalvatzaki, Georgia
This article presents a method for learning well-coordinated Human-Robot Interaction (HRI) from Human-Human Interactions (HHI). We devise a hybrid approach using Hidden Markov Models (HMMs) as the latent space priors for a Variational Autoencoder to model a joint distribution over the interacting agents. We leverage the interaction dynamics learned from HHI to learn HRI and incorporate the conditional generation of robot motions from human observations into the training, thereby predicting more accurate robot trajectories. The generated robot motions are further adapted with Inverse Kinematics to ensure the desired physical proximity with a human, combining the ease of joint space learning and accurate task space reachability. For contact-rich interactions, we modulate the robot's stiffness using HMM segmentation for a compliant interaction. We verify the effectiveness of our approach deployed on a Humanoid robot via a user study. Our method generalizes well to various humans despite being trained on data from just two humans. We find that Users perceive our method as more human-like, timely, and accurate and rank our method with a higher degree of preference over other baselines.
Interactive Autonomous Navigation with Internal State Inference and Interactivity Estimation
Li, Jiachen, Isele, David, Lee, Kanghoon, Park, Jinkyoo, Fujimura, Kikuo, Kochenderfer, Mykel J.
Deep reinforcement learning (DRL) provides a promising way for intelligent agents (e.g., autonomous vehicles) to learn to navigate complex scenarios. However, DRL with neural networks as function approximators is typically considered a black box with little explainability and often suffers from suboptimal performance, especially for autonomous navigation in highly interactive multi-agent environments. To address these issues, we propose three auxiliary tasks with spatio-temporal relational reasoning and integrate them into the standard DRL framework, which improves the decision making performance and provides explainable intermediate indicators. We propose to explicitly infer the internal states (i.e., traits and intentions) of surrounding agents (e.g., human drivers) as well as to predict their future trajectories in the situations with and without the ego agent through counterfactual reasoning. These auxiliary tasks provide additional supervision signals to infer the behavior patterns of other interactive agents. Multiple variants of framework integration strategies are compared. We also employ a spatio-temporal graph neural network to encode relations between dynamic entities, which enhances both internal state inference and decision making of the ego agent. Moreover, we propose an interactivity estimation mechanism based on the difference between predicted trajectories in these two situations, which indicates the degree of influence of the ego agent on other agents. To validate the proposed method, we design an intersection driving simulator based on the Intelligent Intersection Driver Model (IIDM) that simulates vehicles and pedestrians. Our approach achieves robust and state-of-the-art performance in terms of standard evaluation metrics and provides explainable intermediate indicators (i.e., internal states, and interactivity scores) for decision making.
Evaluating the Impact of Personalized Value Alignment in Human-Robot Interaction: Insights into Trust and Team Performance Outcomes
Bhat, Shreyas, Lyons, Joseph B., Shi, Cong, Yang, X. Jessie
This paper examines the effect of real-time, personalized alignment of a robot's reward function to the human's values on trust and team performance. We present and compare three distinct robot interaction strategies: a non-learner strategy where the robot presumes the human's reward function mirrors its own, a non-adaptive-learner strategy in which the robot learns the human's reward function for trust estimation and human behavior modeling, but still optimizes its own reward function, and an adaptive-learner strategy in which the robot learns the human's reward function and adopts it as its own. Two human-subject experiments with a total number of 54 participants were conducted. In both experiments, the human-robot team searches for potential threats in a town. The team sequentially goes through search sites to look for threats. We model the interaction between the human and the robot as a trust-aware Markov Decision Process (trust-aware MDP) and use Bayesian Inverse Reinforcement Learning (IRL) to estimate the reward weights of the human as they interact with the robot. In Experiment 1, we start our learning algorithm with an informed prior of the human's values/goals. In Experiment 2, we start the learning algorithm with an uninformed prior. Results indicate that when starting with a good informed prior, personalized value alignment does not seem to benefit trust or team performance. On the other hand, when an informed prior is unavailable, alignment to the human's values leads to high trust and higher perceived performance while maintaining the same objective team performance.
Towards Transfer Learning for Large-Scale Image Classification Using Annealing-based Quantum Boltzmann Machines
Schuman, Daniëlle, Sünkel, Leo, Altmann, Philipp, Stein, Jonas, Roch, Christoph, Gabor, Thomas, Linnhoff-Popien, Claudia
Quantum Transfer Learning (QTL) recently gained popularity as a hybrid quantum-classical approach for image classification tasks by efficiently combining the feature extraction capabilities of large Convolutional Neural Networks with the potential benefits of Quantum Machine Learning (QML). Existing approaches, however, only utilize gate-based Variational Quantum Circuits for the quantum part of these procedures. In this work we present an approach to employ Quantum Annealing (QA) in QTL-based image classification. Specifically, we propose using annealing-based Quantum Boltzmann Machines as part of a hybrid quantum-classical pipeline to learn the classification of real-world, large-scale data such as medical images through supervised training. We demonstrate our approach by applying it to the three-class COVID-CT-MD dataset, a collection of lung Computed Tomography (CT) scan slices. Using Simulated Annealing as a stand-in for actual QA, we compare our method to classical transfer learning, using a neural network of the same order of magnitude, to display its improved classification performance. We find that our approach consistently outperforms its classical baseline in terms of test accuracy and AUC-ROC-Score and needs less training epochs to do this.
Multi-Agent Reinforcement Learning for Power Control in Wireless Networks via Adaptive Graphs
Amorosa, Lorenzo Mario, Skocaj, Marco, Verdone, Roberto, Gündüz, Deniz
Wireless communication networks constitute complex systems This interest can be attributed to the innate characteristics of demanding careful optimization of network procedures GNNs, which enable a scalable solution and exhibit inductive to attain predefined performance objectives. Multi-agent deep capability and, thanks to the permutation equivariance property, reinforcement learning (MADRL), owing to its inherent advantages, increased generalization. Notably, these properties find has emerged as a promising strategy for the optimization practical application in works such as [5], where GNNs are of a variety of network problems. Nevertheless, the practical harnessed to capture the dynamic structure of fading channel implementation of MADRL in real systems is hindered by states for the purpose of learning optimal resource allocation challenges related to convergence, which continue to constitute policies in wireless networks. Another domain that has witnessed an active area of research. These challenges encompass the substantial utilization of GNNs is channel management non-stationarity of the environment, the partial observability of within wireless local area networks (WLANs), as evidenced the state, as well as the coordination and cooperation among by works such as [6] and [7]. A notable insight derived from agents [1, 2]. To this end, this paper elucidates the role of the study by Gao et al. [6] is the inherent property of GNNs to leveraging graph structures as an effective means to account provide decentralized inference, rendering them a viable and for non-stationarity in MADRL systems by introducing a promising approach for the practical implementation of overthe-air relational inductive bias in the collective decision-making MADRL systems.
Energy Discrepancies: A Score-Independent Loss for Energy-Based Models
Schröder, Tobias, Ou, Zijing, Lim, Jen Ning, Li, Yingzhen, Vollmer, Sebastian J., Duncan, Andrew B.
Energy-based models are a simple yet powerful class of probabilistic models, but their widespread adoption has been limited by the computational burden of training them. We propose a novel loss function called Energy Discrepancy (ED) which does not rely on the computation of scores or expensive Markov chain Monte Carlo. We show that ED approaches the explicit score matching and negative log-likelihood loss under different limits, effectively interpolating between both. Consequently, minimum ED estimation overcomes the problem of nearsightedness encountered in score-based estimation methods, while also enjoying theoretical guarantees. Through numerical experiments, we demonstrate that ED learns low-dimensional data distributions faster and more accurately than explicit score matching or contrastive divergence. For high-dimensional image data, we describe how the manifold hypothesis puts limitations on our approach and demonstrate the effectiveness of energy discrepancy by training the energy-based model as a prior of a variational decoder model.
A Foundational Framework and Methodology for Personalized Early and Timely Diagnosis
Schubert, Tim, Peck, Richard W, Gimson, Alexander, Davtyan, Camelia, van der Schaar, Mihaela
Early diagnosis of diseases holds the potential for deep transformation in healthcare by enabling better treatment options, improving long-term survival and quality of life, and reducing overall cost. With the advent of medical big data, advances in diagnostic tests as well as in machine learning and statistics, early or timely diagnosis seems within reach. Early diagnosis research often neglects the potential for optimizing individual diagnostic paths. To enable personalized early diagnosis, a foundational framework is needed that delineates the diagnosis process and systematically identifies the time-dependent value of various diagnostic tests for an individual patient given their unique characteristics. Here, we propose the first foundational framework for early and timely diagnosis. It builds on decision-theoretic approaches to outline the diagnosis process and integrates machine learning and statistical methodology for estimating the optimal personalized diagnostic path. To describe the proposed framework as well as possibly other frameworks, we provide essential definitions. The development of a foundational framework is necessary for several reasons: 1) formalism provides clarity for the development of decision support tools; 2) observed information can be complemented with estimates of the future patient trajectory; 3) the net benefit of counterfactual diagnostic paths and associated uncertainties can be modeled for individuals 4) 'early' and 'timely' diagnosis can be clearly defined; 5) a mechanism emerges for assessing the value of technologies in terms of their impact on personalized early diagnosis, resulting health outcomes and incurred costs. Finally, we hope that this foundational framework will unlock the long-awaited potential of timely diagnosis and intervention, leading to improved outcomes for patients and higher cost-effectiveness for healthcare systems.
Sensor Allocation and Online-Learning-based Path Planning for Maritime Situational Awareness Enhancement: A Multi-Agent Approach
Nguyen, Bach Long, Doan, Anh-Dzung, Chin, Tat-Jun, Guettier, Christophe, Gupta, Surabhi, Parra, Estelle, Reid, Ian, Wagner, Markus
Countries with access to large bodies of water often aim to protect their maritime transport by employing maritime surveillance systems. However, the number of available sensors (e.g., cameras) is typically small compared to the to-be-monitored targets, and their Field of View (FOV) and range are often limited. This makes improving the situational awareness of maritime transports challenging. To this end, we propose a method that not only distributes multiple sensors but also plans paths for them to observe multiple targets, while minimizing the time needed to achieve situational awareness. In particular, we provide a formulation of this sensor allocation and path planning problem which considers the partial awareness of the targets' state, as well as the unawareness of the targets' trajectories. To solve the problem we present two algorithms: 1) a greedy algorithm for assigning sensors to targets, and 2) a distributed multi-agent path planning algorithm based on regret-matching learning. Because a quick convergence is a requirement for algorithms developed for high mobility environments, we employ a forgetting factor to quickly converge to correlated equilibrium solutions. Experimental results show that our combined approach achieves situational awareness more quickly than related work.
Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs
Lew, Alexander K., Zhi-Xuan, Tan, Grand, Gabriel, Mansinghka, Vikash K.
Even after fine-tuning and reinforcement learning, large language models (LLMs) can be difficult, if not impossible, to control reliably with prompts alone. We propose a new inference-time approach to enforcing syntactic and semantic constraints on the outputs of LLMs, called sequential Monte Carlo (SMC) steering. The key idea is to specify language generation tasks as posterior inference problems in a class of discrete probabilistic sequence models, and replace standard decoding with sequential Monte Carlo inference. For a computational cost similar to that of beam search, SMC can steer LLMs to solve diverse tasks, including infilling, generation under syntactic constraints, and prompt intersection. To facilitate experimentation with SMC steering, we present a probabilistic programming library, LLaMPPL, for concisely specifying new generation tasks as language model probabilistic programs, and automating steering of LLaMA-family Transformers.