Learning Graphical Models
DeePLT: Personalized Lighting Facilitates by Trajectory Prediction of Recognized Residents in the Smart Home
Safaei, Danial, Sobhani, Ali, Kiaei, Ali Akbar
In recent years, the intelligence of various parts of the home has become one of the essential features of any modern home. One of these parts is the intelligence lighting system that personalizes the light for each person. This paper proposes an intelligent system based on machine learning that personalizes lighting in the instant future location of a recognized user, inferred by trajectory prediction. Our proposed system consists of the following modules: (I) human detection to detect and localize the person in each given video frame, (II) face recognition to identify the detected person, (III) human tracking to track the person in the sequence of video frames and (IV) trajectory prediction to forecast the future location of the user in the environment using Inverse Reinforcement Learning. The proposed method provides a unique profile for each person, including specifications, face images, and custom lighting settings. This profile is used in the lighting adjustment process. Unlike other methods that consider constant lighting for every person, our system can apply each 'person's desired lighting in terms of color and light intensity without direct user intervention. Therefore, the lighting is adjusted with higher speed and better efficiency. In addition, the predicted trajectory path makes the proposed system apply the desired lighting, creating more pleasant and comfortable conditions for the home residents. In the experimental results, the system applied the desired lighting in an average time of 1.4 seconds from the moment of entry, as well as a performance of 22.1mAp in human detection, 95.12% accuracy in face recognition, 93.3% MDP in human tracking, and 10.80 MinADE20, 18.55 MinFDE20, 15.8 MinADE5 and 30.50 MinFDE5 in trajectory prediction.
On information captured by neural networks: connections with memorization and generalization
Despite the popularity and success of deep learning, there is limited understanding of when, how, and why neural networks generalize to unseen examples. Since learning can be seen as extracting information from data, we formally study information captured by neural networks during training. Specifically, we start with viewing learning in presence of noisy labels from an information-theoretic perspective and derive a learning algorithm that limits label noise information in weights. We then define a notion of unique information that an individual sample provides to the training of a deep network, shedding some light on the behavior of neural networks on examples that are atypical, ambiguous, or belong to underrepresented subpopulations. We relate example informativeness to generalization by deriving nonvacuous generalization gap bounds. Finally, by studying knowledge distillation, we highlight the important role of data and label complexity in generalization. Overall, our findings contribute to a deeper understanding of the mechanisms underlying neural network generalization.
Rosenthal-type inequalities for linear statistics of Markov chains
Durmus, Alain, Moulines, Eric, Naumov, Alexey, Samsonov, Sergey, Sheshukova, Marina
Probability and moment inequalities for sums of random variables are of paramount importance in the complexity analysis of numerous stochastic approximation algorithms or finite-time analysis of Monte Carlo estimators; see [20], [10], and references therein. The main focus in this area has been on concentration inequalities for independent random variable sums or martingale difference sequences; see e.g. in [4, 36]. However, the study of concentration inequalities for additive Markov chain functions is still relatively underdeveloped. For the technically simple case of uniformly ergodic Markov chains, there is extensive work on Hoeffding-and Bernstein-like inequalities as found in [23, 34, 20, 38]. Nevertheless, the application of these results may be difficult due to a lack of quantitative data or the substitution of asymptotic variance of the chain by surrogates; see Section 2.1 for relevant definitions. The present work aims to fill this gap by extending Rosenthal-and Bernstein-type inequalities to Markov chains which converge geometrically fast to a unique invariant distribution, with an explicit emphasis on the mixing time of the underlying Markov chain. An important tool for establishing deviation bounds for sums of random variables is based on moment inequalities.
Stochastic Methods in Variational Inequalities: Ergodicity, Bias and Refinements
Vlatakis-Gkaragkounis, Emmanouil-Vasileios, Giannou, Angeliki, Chen, Yudong, Xie, Qiaomin
For min-max optimization and variational inequalities problems (VIP) encountered in diverse machine learning tasks, Stochastic Extragradient (SEG) and Stochastic Gradient Descent Ascent (SGDA) have emerged as preeminent algorithms. Constant step-size variants of SEG/SGDA have gained popularity, with appealing benefits such as easy tuning and rapid forgiveness of initial conditions, but their convergence behaviors are more complicated even in rudimentary bilinear models. Our work endeavors to elucidate and quantify the probabilistic structures intrinsic to these algorithms. By recasting the constant step-size SEG/SGDA as time-homogeneous Markov Chains, we establish a first-of-its-kind Law of Large Numbers and a Central Limit Theorem, demonstrating that the average iterate is asymptotically normal with a unique invariant distribution for an extensive range of monotone and non-monotone VIPs. Specializing to convex-concave min-max optimization, we characterize the relationship between the step-size and the induced bias with respect to the Von-Neumann's value. Finally, we establish that Richardson-Romberg extrapolation can improve proximity of the average iterate to the global solution for VIPs. Our probabilistic analysis, underpinned by experiments corroborating our theoretical discoveries, harnesses techniques from optimization, Markov chains, and operator theory.
Sharper Model-free Reinforcement Learning for Average-reward Markov Decision Processes
We develop several provably efficient model-free reinforcement learning (RL) algorithms for infinite-horizon average-reward Markov Decision Processes (MDPs). We consider both online setting and the setting with access to a simulator. In the online setting, we propose model-free RL algorithms based on reference-advantage decomposition. Our algorithm achieves $\widetilde{O}(S^5A^2\mathrm{sp}(h^*)\sqrt{T})$ regret after $T$ steps, where $S\times A$ is the size of state-action space, and $\mathrm{sp}(h^*)$ the span of the optimal bias function. Our results are the first to achieve optimal dependence in $T$ for weakly communicating MDPs. In the simulator setting, we propose a model-free RL algorithm that finds an $\epsilon$-optimal policy using $\widetilde{O} \left(\frac{SA\mathrm{sp}^2(h^*)}{\epsilon^2}+\frac{S^2A\mathrm{sp}(h^*)}{\epsilon} \right)$ samples, whereas the minimax lower bound is $\Omega\left(\frac{SA\mathrm{sp}(h^*)}{\epsilon^2}\right)$. Our results are based on two new techniques that are unique in the average-reward setting: 1) better discounted approximation by value-difference estimation; 2) efficient construction of confidence region for the optimal bias function with space complexity $O(SA)$.
Representation Learning via Variational Bayesian Networks
Barkan, Oren, Caciularu, Avi, Rejwan, Idan, Katz, Ori, Weill, Jonathan, Malkiel, Itzik, Koenigstein, Noam
In the recommender system community, this situation is known as the "cold-start" problem [7, 9], where rare ('cold') entities (e.g., We present Variational Bayesian Network (VBN) - a novel Bayesian unpopular items or new items that are introduced to the catalog) are entity representation learning model that utilizes hierarchical and often poorly represented due to insufficient statistics. In the natural relational side information and is particularly useful for modeling language processing community, where the focus is on learning entities in the "long-tail", where the data is scarce. VBN provides representations for words and phrases, a common mitigation is to better modeling for long-tail entities via two complementary mechanisms: increase the training set size by utilizing increasingly larger corpus First, VBN employs informative hierarchical priors that e.g., BERT [20, 39]. However, it was shown that even when enable information propagation between entities sharing common increasing the amount of co-occurrence data, the existence of rare, ancestors. Additionally, VBN models explicit relations between entities out-of-vocabulary entities persists [26, 50, 52, 53].
Recent Advances in Optimal Transport for Machine Learning
Montesuma, Eduardo Fernandes, Mboula, Fred Ngolè, Souloumiac, Antoine
Recently, Optimal Transport has been proposed as a probabilistic framework in Machine Learning for comparing and manipulating probability distributions. This is rooted in its rich history and theory, and has offered new solutions to different problems in machine learning, such as generative modeling and transfer learning. In this survey we explore contributions of Optimal Transport for Machine Learning over the period 2012 -- 2022, focusing on four sub-fields of Machine Learning: supervised, unsupervised, transfer and reinforcement learning. We further highlight the recent development in computational Optimal Transport, and its interplay with Machine Learning practice.
Sparse Representations, Inference and Learning
Lauditi, Clarissa, Troiani, Emanuele, Mézard, Marc
In recent years statistical physics has proven to be a valuable tool to probe into large dimensional inference problems such as the ones occurring in machine learning. Statistical physics provides analytical tools to study fundamental limitations in their solutions and proposes algorithms to solve individual instances. In these notes, based on the lectures by Marc M\'ezard in 2022 at the summer school in Les Houches, we will present a general framework that can be used in a large variety of problems with weak long-range interactions, including the compressed sensing problem, or the problem of learning in a perceptron. We shall see how these problems can be studied at the replica symmetric level, using developments of the cavity methods, both as a theoretical tool and as an algorithm.
Action and Trajectory Planning for Urban Autonomous Driving with Hierarchical Reinforcement Learning
Lu, Xinyang, Fan, Flint Xiaofeng, Wang, Tianying
Reinforcement Learning (RL) has made promising progress in planning and decision-making for Autonomous Vehicles (AVs) in simple driving scenarios. However, existing RL algorithms for AVs fail to learn critical driving skills in complex urban scenarios. First, urban driving scenarios require AVs to handle multiple driving tasks of which conventional RL algorithms are incapable. Second, the presence of other vehicles in urban scenarios results in a dynamically changing environment, which challenges RL algorithms to plan the action and trajectory of the AV. In this work, we propose an action and trajectory planner using Hierarchical Reinforcement Learning (atHRL) method, which models the agent behavior in a hierarchical model by using the perception of the lidar and birdeye view. The proposed atHRL method learns to make decisions about the agent's future trajectory and computes target waypoints under continuous settings based on a hierarchical DDPG algorithm. The waypoints planned by the atHRL model are then sent to a low-level controller to generate the steering and throttle commands required for the vehicle maneuver. We empirically verify the efficacy of atHRL through extensive experiments in complex urban driving scenarios that compose multiple tasks with the presence of other vehicles in the CARLA simulator. The experimental results suggest a significant performance improvement compared to the state-of-the-art RL methods.
Multi-task Hierarchical Adversarial Inverse Reinforcement Learning
Chen, Jiayu, Tamboli, Dipesh, Lan, Tian, Aggarwal, Vaneet
Multi-task Imitation Learning (MIL) aims to train a policy capable of performing a distribution of tasks based on multi-task expert demonstrations, which is essential for general-purpose robots. Existing MIL algorithms suffer from low data efficiency and poor performance on complex long-horizontal tasks. We develop Multi-task Hierarchical Adversarial Inverse Reinforcement Learning (MH-AIRL) to learn hierarchically-structured multi-task policies, which is more beneficial for compositional tasks with long horizons and has higher expert data efficiency through identifying and transferring reusable basic skills across tasks. To realize this, MH-AIRL effectively synthesizes context-based multi-task learning, AIRL (an IL approach), and hierarchical policy learning. Further, MH-AIRL can be adopted to demonstrations without the task or skill annotations (i.e., state-action pairs only) which are more accessible in practice. Theoretical justifications are provided for each module of MH-AIRL, and evaluations on challenging multi-task settings demonstrate superior performance and transferability of the multi-task policies learned with MH-AIRL as compared to SOTA MIL baselines.