Goto

Collaborating Authors

 Learning Graphical Models


BNNpriors: A library for Bayesian neural network inference with different prior distributions

arXiv.org Machine Learning

Bayesian neural networks have shown great promise in many applications where calibrated uncertainty estimates are crucial and can often also lead to a higher predictive performance. However, it remains challenging to choose a good prior distribution over their weights. While isotropic Gaussian priors are often chosen in practice due to their simplicity, they do not reflect our true prior beliefs well and can lead to suboptimal performance. Our new library, BNNpriors, enables state-of-the-art Markov Chain Monte Carlo inference on Bayesian neural networks with a wide range of predefined priors, including heavy-tailed ones, hierarchical ones, and mixture priors. Moreover, it follows a modular approach that eases the design and implementation of new custom priors. It has facilitated foundational discoveries on the nature of the cold posterior effect in Bayesian neural networks and will hopefully catalyze future research as well as practical applications in this area.


SAT-Based Rigorous Explanations for Decision Lists

arXiv.org Artificial Intelligence

Decision lists (DLs) find a wide range of uses for classification problems in Machine Learning (ML), being implemented in a number of ML frameworks. DLs are often perceived as interpretable. However, building on recent results for decision trees (DTs), we argue that interpretability is an elusive goal for some DLs. As a result, for some uses of DLs, it will be important to compute (rigorous) explanations. Unfortunately, and in clear contrast with the case of DTs, this paper shows that computing explanations for DLs is computationally hard. Motivated by this result, the paper proposes propositional encodings for computing abductive explanations (AXps) and contrastive explanations (CXps) of DLs. Furthermore, the paper investigates the practical efficiency of a MARCO-like approach for enumerating explanations. The experimental results demonstrate that, for DLs used in practical settings, the use of SAT oracles offers a very efficient solution, and that complete enumeration of explanations is most often feasible.


Adapting deep generative approaches for getting synthetic data with realistic marginal distributions

arXiv.org Machine Learning

Synthetic data generation is of great interest in diverse applications, such as for privacy protection. Deep generative models, such as variational autoencoders (VAEs), are a popular approach for creating such synthetic datasets from original data. Despite the success of VAEs, there are limitations when it comes to the bimodal and skewed marginal distributions. These deviate from the unimodal symmetric distributions that are encouraged by the normality assumption typically used for the latent representations in VAEs. While there are extensions that assume other distributions for the latent space, this does not generally increase flexibility for data with many different distributions. Therefore, we propose a novel method, pre-transformation variational autoencoders (PTVAEs), to specifically address bimodal and skewed data, by employing pre-transformations at the level of original variables. Two types of transformations are used to bring the data close to a normal distribution by a separate parameter optimization for each variable in a dataset. We compare the performance of our method with other state-of-the-art methods for synthetic data generation. In addition to the visual comparison, we use a utility measurement for a quantitative evaluation. The results show that the PTVAE approach can outperform others in both bimodal and skewed data generation. Furthermore, the simplicity of the approach makes it usable in combination with other extensions of VAE.


A Monotone Approximate Dynamic Programming Approach for the Stochastic Scheduling, Allocation, and Inventory Replenishment Problem: Applications to Drone and Electric Vehicle Battery Swap Stations

arXiv.org Artificial Intelligence

There is a growing interest in using electric vehicles (EVs) and drones for many applications. However, battery-oriented issues, including range anxiety and battery degradation, impede adoption. Battery swap stations are one alternative to reduce these concerns that allow the swap of depleted for full batteries in minutes. We consider the problem of deriving actions at a battery swap station when explicitly considering the uncertain arrival of swap demand, battery degradation, and replacement. We model the operations at a battery swap station using a finite horizon Markov Decision Process model for the stochastic scheduling, allocation, and inventory replenishment problem (SAIRP), which determines when and how many batteries are charged, discharged, and replaced over time. We present theoretical proofs for the monotonicity of the value function and monotone structure of an optimal policy for special SAIRP cases. Due to the curses of dimensionality, we develop a new monotone approximate dynamic programming (ADP) method, which intelligently initializes a value function approximation using regression. In computational tests, we demonstrate the superior performance of the new regression-based monotone ADP method as compared to exact methods and other monotone ADP methods. Further, with the tests, we deduce policy insights for drone swap stations.


Scaling Ensemble Distribution Distillation to Many Classes with Proxy Targets

arXiv.org Artificial Intelligence

Ensembles of machine learning models yield improved system performance as well as robust and interpretable uncertainty estimates; however, their inference costs may often be prohibitively high. Ensemble Distribution Distillation is an approach that allows a single model to efficiently capture both the predictive performance and uncertainty estimates of an ensemble. For classification, this is achieved by training a Dirichlet distribution over the ensemble members' output distributions via the maximum likelihood criterion. Although theoretically principled, this criterion exhibits poor convergence when applied to large-scale tasks where the number of classes is very high. In our work, we analyze this effect and show that for the Dirichlet log-likelihood criterion classes with low probability induce larger gradients than high-probability classes. This forces the model to focus on the distribution of the ensemble tail-class probabilities. We propose a new training objective which minimizes the reverse KL-divergence to a Proxy-Dirichlet target derived from the ensemble. This loss resolves the gradient issues of Ensemble Distribution Distillation, as we demonstrate both theoretically and empirically on the ImageNet and WMT17 En-De datasets containing 1000 and 40,000 classes, respectively.


Estimating Disentangled Belief about Hidden State and Hidden Task for Meta-RL

arXiv.org Artificial Intelligence

There is considerable interest in designing meta-reinforcement learning (meta-RL) algorithms, which enable autonomous agents to adapt new tasks from small amount of experience. In meta-RL, the specification (such as reward function) of current task is hidden from the agent. In addition, states are hidden within each task owing to sensor noise or limitations in realistic environments. Therefore, the meta-RL agent faces the challenge of specifying both the hidden task and states based on small amount of experience. To address this, we propose estimating disentangled belief about task and states, leveraging an inductive bias that the task and states can be regarded as global and local features of each task. Specifically, we train a hierarchical state-space model (HSSM) parameterized by deep neural networks as an environment model, whose global and local latent variables correspond to task and states, respectively. Because the HSSM does not allow analytical computation of posterior distribution, i.e., belief, we employ amortized inference to approximate it. After the belief is obtained, we can augment observations of a model-free policy with the belief to efficiently train the policy. Moreover, because task and state information are factorized and interpretable, the downstream policy training is facilitated compared with the prior methods that did not consider the hierarchical nature. Empirical validations on a GridWorld environment confirm that the HSSM can separate the hidden task and states information. Then, we compare the meta-RL agent with the HSSM to prior meta-RL methods in MuJoCo environments, and confirm that our agent requires less training data and reaches higher final performance.


A cell type-specific cortico-subcortical brain circuit for investigatory and novelty-seeking behavior

Science

Curiosity is what drives organisms to investigate each other and their environment. It is considered by many to be as intrinsic as hunger and thirst, but the neurobiological mechanisms behind curiosity have remained elusive. In mice, Ahmadlou et al. found that a specific population of genetically identified γ-aminobutyric acid (GABA)—ergic neurons in a brain region called the zona incerta receive excitatory input in the form of novelty and/or arousal information from the prelimbic cortex, and these neurons send inhibitory projections to the periaqueductal gray region (see the Perspective by Farahbakhsh and Siciliano). This circuitry is necessary for the exploration of new objects and conspecifics. Science , this issue p. [eabe9681][1]; see also p. [684][2] ### INTRODUCTION Motivational drives are internal states that can be different even in similar interactions with external stimuli. Curiosity as the motivational drive for novelty-seeking and investigating the surrounding environment is for survival as essential and intrinsic as hunger. Curiosity, hunger, and appetitive aggression drive three different goal-directed behaviors—novelty seeking, food eating, and hunting—but these behaviors are composed of similar actions in animals. This similarity of actions has made it challenging to study novelty seeking and distinguish it from eating and hunting in nonarticulating animals. The brain mechanisms underlying this basic survival drive, curiosity, and novelty-seeking behavior have remained unclear. ### RATIONALE In spite of having well-developed techniques to study mouse brain circuits, there are many controversial and different results in the field of motivational behavior. This has left the functions of motivational brain regions such as the zona incerta (ZI) still uncertain. Not having a transparent, nonreinforced, and easily replicable paradigm is one of the main causes of this uncertainty. Therefore, we chose a simple solution to conduct our research: giving the mouse freedom to choose what it wants—double free-access choice. By examining mice in an experimental battery of object free-access double-choice (FADC) and social interaction tests—using optogenetics, chemogenetics, calcium fiber photometry, multichannel recording electrophysiology, and multicolor mRNA in situ hybridization—we uncovered a cell type–specific cortico-subcortical brain circuit of the curiosity and novelty-seeking behavior. ### RESULTS We analyzed the transitions within action sequences in object FADC and social interaction tests. Frequency and hidden Markov model analyses showed that mice choose different action sequences in interaction with novel objects and in early periods of interaction with novel conspecifics compared with interaction with familiar objects or later periods of interaction with conspecifics, which we categorized as deep and shallow investigation, respectively. This finding helped us to define a measure of depth of investigation that indicates how much a mouse prefers deep over shallow investigation and reflects the mouse’s motivational level to investigate, regardless of total duration of investigation. Optogenetic activation of inhibitory neurons in medial ZI (ZIm), ZImGAD2 neurons, showed a dramatic increase in positive arousal level, depth of investigation, and duration of interaction with conspecifics and novel objects compared with familiar objects, crickets, and food. Optogenetic or chemogenetic deactivation of these neurons decreased depth and duration of investigation. Moreover, we found that ZImGAD2 neurons are more active during deep investigation as compared with during shallow investigation. We found that activation of prelimbic cortex (PL) axons into ZIm increases arousal level, and chemogenetic deactivation of these axons decreases the duration and depth of investigation. Calcium fiber photometry of these axons showed no difference in activity between shallow and deep investigation, suggesting a nonspecific motivation. Optogenetic activation of ZImGAD2 axons into lateral periaqueductal gray (lPAG) increases the arousal level, whereas chemogenetic deactivation of these axons decreases duration and depth of investigation. Calcium fiber photometry of these axons showed high activity during deep investigation and no significant activity during shallow investigation, suggesting a thresholding mechanism. Last, we found a new subpopulation of inhibitory neurons in ZIm expressing tachykinin 1 (TAC1) that monosynaptically receive PL inputs and project to lPAG. Optogenetic activation and deactivation of these neurons, respectively, increased and decreased depth and duration of investigation. ### CONCLUSION Our experiments revealed different action sequences based on the motivational level of novelty seeking. Moreover, we uncovered a new brain circuit underlying curiosity and novelty-seeking behavior, connecting excitatory neurons of PL to lPAG through TAC1+ inhibitory neurons of ZIm. ![Figure][3] Brain mechanism of curiosity. ( A ) How we mapped motivational level to action sequences. ( B ) Experimental battery to distinguish novelty-seeking behavior from food eating and hunting in mice with photoactivation of ZImGAD2 neurons. ( C ) Schematic of calcium activity in PL→ZIm, ZIm, and ZIm→PAG during shallow and deep investigation. ( D ) TAC1+ neurons as a subpopulation of ZImGAD2 neurons receive input from PL and project to PAG. HMM, hidden Markov model. Exploring the physical and social environment is essential for understanding the surrounding world. We do not know how novelty-seeking motivation initiates the complex sequence of actions that make up investigatory behavior. We found in mice that inhibitory neurons in the medial zona incerta (ZIm), a subthalamic brain region, are essential for the decision to investigate an object or a conspecific. These neurons receive excitatory input from the prelimbic cortex to signal the initiation of exploration. This signal is modulated in the ZIm by the level of investigatory motivation. Increased activity in the ZIm instigates deep investigative action by inhibiting the periaqueductal gray region. A subpopulation of inhibitory ZIm neurons expressing tachykinin 1 (TAC1) modulates the investigatory behavior. [1]: /lookup/doi/10.1126/science.abe9681 [2]: /lookup/doi/10.1126/science.abi7270 [3]: pending:yes


Intelligence and Unambitiousness Using Algorithmic Information Theory

arXiv.org Artificial Intelligence

Algorithmic Information Theory has inspired intractable constructions of general intelligence (AGI), and undiscovered tractable approximations are likely feasible. Reinforcement Learning (RL), the dominant paradigm by which an agent might learn to solve arbitrary solvable problems, gives an agent a dangerous incentive: to gain arbitrary "power" in order to intervene in the provision of their own reward. We review the arguments that generally intelligent algorithmic-information-theoretic reinforcement learners such as Hutter's (2005) AIXI would seek arbitrary power, including over us. Then, using an information-theoretic exploration schedule, and a setup inspired by causal influence theory, we present a variant of AIXI which learns to not seek arbitrary power; we call it "unambitious". We show that our agent learns to accrue reward at least as well as a human mentor, while relying on that mentor with diminishing probability. And given a formal assumption that we probe empirically, we show that eventually, the agent's world-model incorporates the following true fact: intervening in the "outside world" will have no effect on reward acquisition; hence, it has no incentive to shape the outside world.


Advances in Machine and Deep Learning for Modeling and Real-time Detection of Multi-Messenger Sources

arXiv.org Artificial Intelligence

This chapter provides a summary of recent developments harnessing the data revolution to realize the science goals of Gravitational Wave Astrophysics. This is an exciting journey that is powered by the renaissance of artificial intelligence, and a new generation of researchers that are willing to embrace disruptive advances in innovative computing and signal processing tools. In this chapter, machine learning refers to a class of algorithms that can learn from data to solve new problems without being explicitly re-programmed. While traditional machine learning algorithms, e.g., random forests, nearest neighbors, etc., have been used successfully in many applications, they are limited in their ability to process raw data, usually requiring time-consuming feature engineering to preprocess data into a suitable representation for each application. On the other hand, deep learning algorithms can learn patterns from unstructured data, finding useful representations and automatically extracting relevant features for each application. The ability of deep learning to deal with poorly defined abstractions and problems has led to major advances in image recognition, speech, computer vision applications, robotics, among others [1]. The following sections describe a few noteworthy applications of modern machine learning for gravitational wave modeling, detection and inference. It is the expectation that by the time this chapter is published, the ongoing developments at the interface of artificial intelligence and extreme-scale computing will have leapt forward, making this chapter a reminiscence of a fast-paced, evolving field of research. The chapter concludes with a summary of recent applications at the interface of deep learning and high performance computing to address computational grand challenges in Gravitational Wave Astrophysics.


Learning Bayesian Networks: A Unification for Discrete and Gaussian Domains

arXiv.org Artificial Intelligence

At last year's conference, we presented approaches for learning Bayesian networks from a combination of prior knowledge and statistical data. These approaches were presented in two papers: one addressing domains containing only discrete variables (Heckerman et al., 1994), and the other addressing domains containing continuous variables related by an unknown multivariate-Gaussian distribution (Geiger and Heckerman, 1994). Unfortunately, these presentations were substantially different, making the parallels between the two methods difficult to appreciate. In this paper, we unify the two approaches. In particular, we abstract our previous assumptions of likelihood equivalence, parameter modularity, and parameter independence such that they are appropriate for discrete and Gaussian domains (as well as other domains). Using these assumptions, we derive a domain-independent Bayesian scoring metric. We then use this general metric in combination with well-known statistical facts about the Dirichlet and normal-Wishart distributions to derive our metrics for discrete and Gaussian domains. In addition, we provide simple proofs that these assumptions are consistent for both domains.