Energy
Digital Twin System for Home Service Robot Based on Motion Simulation
Jiang, Zhengsong, Tian, Guohui, Cui, Yongcheng, Liu, Tiantian, Gu, Yu, Wang, Yifei
In order to improve the task execution capability of home service robot, and to cope with the problem that purely physical robot platforms cannot sense the environment and make decisions online, a method for building digital twin system for home service robot based on motion simulation is proposed. A reliable mapping of the home service robot and its working environment from physical space to digital space is achieved in three dimensions: geometric, physical and functional. In this system, a digital space-oriented URDF file parser is designed and implemented for the automatic construction of the robot geometric model. Next, the physical model is constructed from the kinematic equations of the robot and an improved particle swarm optimization algorithm is proposed for the inverse kinematic solution. In addition, to adapt to the home environment, functional attributes are used to describe household objects, thus improving the semantic description of the digital space for the real home environment. Finally, through geometric model consistency verification, physical model validity verification and virtual-reality consistency verification, it shows that the digital twin system designed in this paper can construct the robot geometric model accurately and completely, complete the operation of household objects successfully, and the digital twin system is effective and practical.
Safe Reinforcement Learning for Strategic Bidding of Virtual Power Plants in Day-Ahead Markets
Stanojev, Ognjen, Mitridati, Lesia, di Prata, Riccardo de Nardis, Hug, Gabriela
For this reason, their applicability in practice is limited. Growing environmental concerns and advancements in communication The above-mentioned scalability issues can be addressed and monitoring technologies have led to the increased by employing deep RL methods like the Deep Deterministic deployment of Distributed Energy Resources (DERs) Policy Gradient (DDPG) algorithm [10], which utilizes neural in power networks [1], comprising renewable energy sources networks to extend the Q-learning capabilities to continuous and prosumers. The market integration of these units is facilitated state and action spaces. The authors in [11]-[13] propose deep by their large-scale aggregation under financial entities, RL methods for the economic dispatch and market participation commonly known as Virtual Power Plants (VPPs), which have of DERs aggregated in a VPP. The main limitation the capacity for trading in wholesale electricity markets [2], of these works is that they fail to account for the complex [3]. As a self-interested market participant, a VPP aims at internal physical constraints of large-scale VPPs, such as maximizing its own profit generated by its market participation power generation limits and power flow constraints, in order to and the fulfillment of contractual obligations towards its ensure a safe operation.
Computationally Efficient Reinforcement Learning: Targeted Exploration leveraging Simple Rules
Di Natale, Loris, Svetozarevic, Bratislav, Heer, Philipp, Jones, Colin N.
Model-free Reinforcement Learning (RL) generally suffers from poor sample complexity, mostly due to the need to exhaustively explore the state-action space to find well-performing policies. On the other hand, we postulate that expert knowledge of the system often allows us to design simple rules we expect good policies to follow at all times. In this work, we hence propose a simple yet effective modification of continuous actor-critic frameworks to incorporate such rules and avoid regions of the state-action space that are known to be suboptimal, thereby significantly accelerating the convergence of RL agents. Concretely, we saturate the actions chosen by the agent if they do not comply with our intuition and, critically, modify the gradient update step of the policy to ensure the learning process is not affected by the saturation step. On a room temperature control case study, it allows agents to converge to well-performing policies up to 6-7x faster than classical agents without computational overhead and while retaining good final performance.
Nonlinear MPC for Quadrotors in Close-Proximity Flight with Neural Network Downwash Prediction
Li, Jinjie, Han, Liang, Yu, Haoyang, Lin, Yuheng, Li, Qingdong, Ren, Zhang
Swarm aerial robots are required to maintain close proximity to successfully traverse narrow areas in cluttered environments. However, this movement is affected by the downwash effect generated from other quadrotors in the swarm. This aerodynamic effect is highly nonlinear and hard to describe through mathematical modeling. Additionally, the existence of the downwash disturbance can be predicted based on the states of neighboring quadrotors. If this prediction is considered, the control loop can proactively handle the disturbance, resulting in improved performance. To address these challenges, we propose an approach that integrates a Neural network Downwash Predictor with Nonlinear Model Predictive Control (NDP-NMPC). The neural network is trained with spectral normalization to ensure robustness and safety in uncollected cases. The predicted disturbances are then incorporated into the optimization scheme in NMPC, which enforces constraints to ensure that states and inputs remain within safe limits. We also design a quadrotor system, identify its parameters, and implement the proposed method on board. Finally, we conduct a prediction experiment to validate the safety and effectiveness of the network. In addition, a real-time trajectory tracking experiment is performed with the entire system, demonstrating a 75.37% reduction in tracking error in height under the downwash effect.
A DRL-based Reflection Enhancement Method for RIS-assisted Multi-receiver Communications
Wang, Wei, Li, Peizheng, Doufexi, Angela, Beach, Mark A
In reconfigurable intelligent surface (RIS)-assisted wireless communication systems, the pointing accuracy and intensity of reflections depend crucially on the 'profile,' representing the amplitude/phase state information of all elements in a RIS array. The superposition of multiple single-reflection profiles enables multi-reflection for distributed users. However, the optimization challenges from periodic element arrangements in single-reflection and multi-reflection profiles are understudied. The combination of periodical single-reflection profiles leads to amplitude/phase counteractions, affecting the performance of each reflection beam. This paper focuses on a dual-reflection optimization scenario and investigates the far-field performance deterioration caused by the misalignment of overlapped profiles. To address this issue, we introduce a novel deep reinforcement learning (DRL)-based optimization method. Comparative experiments against random and exhaustive searches demonstrate that our proposed DRL method outperforms both alternatives, achieving the shortest optimization time. Remarkably, our approach achieves a 1.2 dB gain in the reflection peak gain and a broader beam without any hardware modifications.
Data-Driven Model Reduction and Nonlinear Model Predictive Control of an Air Separation Unit by Applied Koopman Theory
Schulze, Jan C., Doncevic, Danimir T., Erwes, Nils, Mitsos, Alexander
Model reduction using Koopman theory as well as the related dynamic mode decomposition (Schmid, 2010), Computationally tractable models are a main requirement build on a lift-and-project concept and aim to construct linear for real-time NMPC (Marquardt, 2002). Data-driven nonintrusive representations of nonlinear dynamics through (nonlinear) model reduction comprises a class of model-free coordinate transformation. Applied Koopman theory has methods for producing low-order representations of highorder a system-theoretic foundation and naturally combines simple dynamical systems from data, e.g., Antoulas et al. dynamic forms with data-driven identification of coordinate (2017). Similar to classical model reduction approaches transformations, e.g., through Kernel methods (Williams (Marquardt, 2002), these data-driven methods project a highorder et al., 2015), deep learning (Lusch et al., 2018), or sparse regression system from the full state space to a lower dimensional techniques (Brunton et al., 2016).
Learning noise-induced transitions by multi-scaling reservoir computing
Lin, Zequn, Lu, Zhaofan, Di, Zengru, Tang, Ying
Noise is usually regarded as adversarial to extract the effective dynamics from time series, such that the conventional data-driven approaches usually aim at learning the dynamics by mitigating the noisy effect. However, noise can have a functional role of driving transitions between stable states underlying many natural and engineered stochastic dynamics. To capture such stochastic transitions from data, we find that leveraging a machine learning model, reservoir computing as a type of recurrent neural network, can learn noise-induced transitions. We develop a concise training protocol for tuning hyperparameters, with a focus on a pivotal hyperparameter controlling the time scale of the reservoir dynamics. The trained model generates accurate statistics of transition time and the number of transitions. The approach is applicable to a wide class of systems, including a bistable system under a double-well potential, with either white noise or colored noise. It is also aware of the asymmetry of the double-well potential, the rotational dynamics caused by non-detailed balance, and transitions in multi-stable systems. For the experimental data of protein folding, it learns the transition time between folded states, providing a possibility of predicting transition statistics from a small dataset. The results demonstrate the capability of machine-learning methods in capturing noise-induced phenomena.
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
Parcollet, Titouan, Nguyen, Ha, Evain, Solene, Boito, Marcely Zanon, Pupier, Adrien, Mdhaffar, Salima, Le, Hang, Alisamir, Sina, Tomashenko, Natalia, Dinarelli, Marco, Zhang, Shucong, Allauzen, Alexandre, Coavoux, Maximin, Esteve, Yannick, Rouvier, Mickael, Goulian, Jerome, Lecouteux, Benjamin, Portet, Francois, Rossato, Solange, Ringeval, Fabien, Schwab, Didier, Besacier, Laurent
Self-supervised learning (SSL) is at the origin of unprecedented improvements in many different domains including computer vision and natural language processing. Speech processing drastically benefitted from SSL as most of the current domain-related tasks are now being approached with pre-trained models. This work introduces LeBenchmark 2.0 an open-source framework for assessing and building SSL-equipped French speech technologies. It includes documented, large-scale and heterogeneous corpora with up to 14,000 hours of heterogeneous speech, ten pre-trained SSL wav2vec 2.0 models containing from 26 million to one billion learnable parameters shared with the community, and an evaluation protocol made of six downstream tasks to complement existing benchmarks. LeBenchmark 2.0 also presents unique perspectives on pre-trained SSL models for speech with the investigation of frozen versus fine-tuned downstream models, task-agnostic versus task-specific pre-trained models as well as a discussion on the carbon footprint of large-scale model training.
Mind the Uncertainty: Risk-Aware and Actively Exploring Model-Based Reinforcement Learning
Vlastelica, Marin, Blaes, Sebastian, Pineri, Cristina, Martius, Georg
We introduce a simple but effective method for managing risk in model-based reinforcement learning with trajectory sampling that involves probabilistic safety constraints and balancing of optimism in the face of epistemic uncertainty and pessimism in the face of aleatoric uncertainty of an ensemble of stochastic neural networks. Various experiments indicate that the separation of uncertainties is essential to performing well with data-driven MPC approaches in uncertain and safety-critical control environments.