Goto

Collaborating Authors

 Learning Graphical Models


Partially Observable Task and Motion Planning with Uncertainty and Risk Awareness

arXiv.org Artificial Intelligence

Integrated task and motion planning (TAMP) has proven to be a valuable approach to generalizable long-horizon robotic manipulation and navigation problems. However, the typical TAMP problem formulation assumes full observability and deterministic action effects. These assumptions limit the ability of the planner to gather information and make decisions that are risk-aware. We propose a strategy for TAMP with Uncertainty and Risk Awareness (TAMPURA) that is capable of efficiently solving long-horizon planning problems with initial-state and action outcome uncertainty, including problems that require information gathering and avoiding undesirable and irreversible outcomes. Our planner reasons under uncertainty at both the abstract task level and continuous controller level. Given a set of closed-loop goal-conditioned controllers operating in the primitive action space and a description of their preconditions and potential capabilities, we learn a high-level abstraction that can be solved efficiently and then refined to continuous actions for execution. We demonstrate our approach on several robotics problems where uncertainty is a crucial factor and show that reasoning under uncertainty in these problems outperforms previously proposed determinized planning, direct search, and reinforcement learning strategies. Lastly, we demonstrate our planner on two real-world robotics problems using recent advancements in probabilistic perception.


Belief Aided Navigation using Bayesian Reinforcement Learning for Avoiding Humans in Blind Spots

arXiv.org Artificial Intelligence

Recent research on mobile robot navigation has focused on socially aware navigation in crowded environments. However, existing methods do not adequately account for human robot interactions and demand accurate location information from omnidirectional sensors, rendering them unsuitable for practical applications. In response to this need, this study introduces a novel algorithm, BNBRL+, predicated on the partially observable Markov decision process framework to assess risks in unobservable areas and formulate movement strategies under uncertainty. BNBRL+ consolidates belief algorithms with Bayesian neural networks to probabilistically infer beliefs based on the positional data of humans. It further integrates the dynamics between the robot, humans, and inferred beliefs to determine the navigation paths and embeds social norms within the reward function, thereby facilitating socially aware navigation. Through experiments in various risk laden scenarios, this study validates the effectiveness of BNBRL+ in navigating crowded environments with blind spots. The model's ability to navigate effectively in spaces with limited visibility and avoid obstacles dynamically can significantly improve the safety and reliability of autonomous vehicles.


HeR-DRL:Heterogeneous Relational Deep Reinforcement Learning for Decentralized Multi-Robot Crowd Navigation

arXiv.org Artificial Intelligence

Crowd navigation has received significant research attention in recent years, especially DRL-based methods. While single-robot crowd scenarios have dominated research, they offer limited applicability to real-world complexities. The heterogeneity of interaction among multiple agent categories, like in decentralized multi-robot pedestrian scenarios, are frequently disregarded. This "interaction blind spot" hinders generalizability and restricts progress towards robust navigation algorithms. In this paper, we propose a heterogeneous relational deep reinforcement learning(HeR-DRL), based on customised heterogeneous GNN, in order to improve navigation strategies in decentralized multi-robot crowd navigation. Firstly, we devised a method for constructing robot-crowd heterogenous relation graph that effectively simulates the heterogeneous pair-wise interaction relationships. We proposed a new heterogeneous graph neural network for transferring and aggregating the heterogeneous state information. Finally, we incorporate the encoded information into deep reinforcement learning to explore the optimal policy. HeR-DRL are rigorously evaluated through comparing it to state-of-the-art algorithms in both single-robot and multi-robot circle crowssing scenario. The experimental results demonstrate that HeR-DRL surpasses the state-of-the-art approaches in overall performance, particularly excelling in safety and comfort metrics. This underscores the significance of interaction heterogeneity for crowd navigation. The source code will be publicly released in https://github.com/Zhouxy-Debugging-Den/HeR-DRL.


A Multilingual Perspective on Probing Gender Bias

arXiv.org Artificial Intelligence

Gender bias represents a form of systematic negative treatment that targets individuals based on their gender. This discrimination can range from subtle sexist remarks and gendered stereotypes to outright hate speech. Prior research has revealed that ignoring online abuse not only affects the individuals targeted but also has broader societal implications. These consequences extend to the discouragement of women's engagement and visibility within public spheres, thereby reinforcing gender inequality. This thesis investigates the nuances of how gender bias is expressed through language and within language technologies. Significantly, this thesis expands research on gender bias to multilingual contexts, emphasising the importance of a multilingual and multicultural perspective in understanding societal biases. In this thesis, I adopt an interdisciplinary approach, bridging natural language processing with other disciplines such as political science and history, to probe gender bias in natural language and language models.


Horizon-Free Regret for Linear Markov Decision Processes

arXiv.org Artificial Intelligence

A recent line of works showed regret bounds in reinforcement learning (RL) can be (nearly) independent of planning horizon, a.k.a. the horizon-free bounds. However, these regret bounds only apply to settings where a polynomial dependency on the size of transition model is allowed, such as tabular Markov Decision Process (MDP) and linear mixture MDP. We give the first horizon-free bound for the popular linear MDP setting where the size of the transition model can be exponentially large or even uncountable. In contrast to prior works which explicitly estimate the transition model and compute the inhomogeneous value functions at different time steps, we directly estimate the value functions and confidence sets. We obtain the horizon-free bound by: (1) maintaining multiple weighted least square estimators for the value functions; and (2) a structural lemma which shows the maximal total variation of the inhomogeneous value functions is bounded by a polynomial factor of the feature dimension.


Diffusion-Reinforcement Learning Hierarchical Motion Planning in Adversarial Multi-agent Games

arXiv.org Artificial Intelligence

Reinforcement Learning- (RL-)based motion planning has recently shown the potential to outperform traditional approaches from autonomous navigation to robot manipulation. In this work, we focus on a motion planning task for an evasive target in a partially observable multi-agent adversarial pursuit-evasion games (PEG). These pursuit-evasion problems are relevant to various applications, such as search and rescue operations and surveillance robots, where robots must effectively plan their actions to gather intelligence or accomplish mission tasks while avoiding detection or capture themselves. We propose a hierarchical architecture that integrates a high-level diffusion model to plan global paths responsive to environment data while a low-level RL algorithm reasons about evasive versus global path-following behavior. Our approach outperforms baselines by 51.2% by leveraging the diffusion model to guide the RL algorithm for more efficient exploration and improves the explanability and predictability.


An Improved Strategy for Blood Glucose Control Using Multi-Step Deep Reinforcement Learning

arXiv.org Artificial Intelligence

Diabetes profoundly affects human life and health, regardless of country, age, or gender, and is one of the leading causes of death and disability worldwide [1]. From 1990 to 2021, the age-standardized prevalence of diabetes increased by 90.5 % globally, with increases of more than 100 % in several regions, and it is projected that by 2050, there will be 1.31 billion people with diabetes worldwide [1]. Furthermore, people with diabetes have more than twice the normal risk of early death, resulting in an estimated 150-500 million deaths around the world each year, while generating approximately 12% of health expenditure ($966 billion) [2, 3]. The rising prevalence and serious health and economic hazards have attracted the attention of scientists around the globe, and as a result, the number of studies on diabetes is increasing. The pancreas of a diabetic does not produce or produces very little insulin, or the insulin produced is not used efficiently, leading to high BG and a variety of life-threatening complications such as cardiovascular disease, nerve damage, kidney damage, lower limb amputations, and eye disease leading to decreased vision and even blindness [3]. BG control is their basic treatment, as well as the basis for preventing and treating diabetic complications. Patients mainly maintain the stability of BG by injecting insulin. However, this traditional self-management is usually cumbersome and challenging, as it requires patients to measure their BG levels several times a day, while they suffer from many of the aforementioned complications [2].


Single- and Multi-Agent Private Active Sensing: A Deep Neuroevolution Approach

arXiv.org Artificial Intelligence

The problem of single-agent Evasive AHT (EAHT), Active Hypothesis Testing (AHT) refers to the family of where a passive Eavesdropper (Eve) collects noisy estimates problems where one legitimate agent or decision maker, or a of the legit observations and tries to infer the underlying group of collaborating agents or decision makers, adaptively hypothesis, was studied in [24], focusing however explicitly select(s) sensing actions and collect(s) observations in order on the asymptotical case. In that work, the authors formulated to infer the underlying true hypothesis in a fast and reliable single-agent EAHT as a constrained optimization problem manner [1], [2]. AHT and related problems, such as active including the legitimate agent's and the Eavesdropper's (Eve) parameter estimation [3] and active change point detection [4], error exponent. However, near-optimal or optimal action selection [5], find numerous applications in wireless communications, policies were not presented. In this paper, motivated including anomaly detection over sensor networks [6], strong by the lack of explicit policies for EAHT, we present novel or weak radar models for target detection [7], cyber-intrusion single-and multi-agent EAHT approaches for wireless sensor detection, target search, and adaptive beamforming [8], as well networks that are based on a deep NeuroEvolution (NE) as, very recently, RIS-enabled localization [9] and channel framework. Our contributions are summarized as follows: estimation [10]. In addition, AHT is closely related to the 1) We formulate the single-agent EAHT problem studied feedback channel coding problem [11].


Generative Modelling of Stochastic Rotating Shallow Water Noise

arXiv.org Machine Learning

In recent work, the authors have developed a generic methodology for calibrating the noise in fluid dynamics stochastic partial differential equations where the stochasticity was introduced to parametrize subgrid-scale processes. The stochastic parameterization of sub-grid scale processes is required in the estimation of uncertainty in weather and climate predictions, to represent systematic model errors arising from subgrid-scale fluctuations. The previous methodology used a principal component analysis (PCA) technique based on the ansatz that the increments of the stochastic parametrization are normally distributed. In this paper, the PCA technique is replaced by a generative model technique. This enables us to avoid imposing additional constraints on the increments. The methodology is tested on a stochastic rotating shallow water model with the elevation variable of the model used as input data. The numerical simulations show that the noise is indeed non-Gaussian. The generative modelling technology gives good RMSE, CRPS score and forecast rank histogram results.


Hessian-Free Laplace in Bayesian Deep Learning

arXiv.org Machine Learning

The Laplace approximation (LA) of the Bayesian posterior is a Gaussian distribution centered at the maximum a posteriori estimate. Its appeal in Bayesian deep learning stems from the ability to quantify uncertainty post-hoc (i.e., after standard network parameter optimization), the ease of sampling from the approximate posterior, and the analytic form of model evidence. However, an important computational bottleneck of LA is the necessary step of calculating and inverting the Hessian matrix of the log posterior. The Hessian may be approximated in a variety of ways, with quality varying with a number of factors including the network, dataset, and inference task. In this paper, we propose an alternative framework that sidesteps Hessian calculation and inversion. The Hessian-free Laplace (HFL) approximation uses curvature of both the log posterior and network prediction to estimate its variance. Only two point estimates are needed: the standard maximum a posteriori parameter and the optimal parameter under a loss regularized by the network prediction. We show that, under standard assumptions of LA in Bayesian deep learning, HFL targets the same variance as LA, and can be efficiently amortized in a pre-trained network. Experiments demonstrate comparable performance to that of exact and approximate Hessians, with excellent coverage for in-between uncertainty.