Goto

Collaborating Authors

 Country


Task-Relevant Adversarial Imitation Learning

arXiv.org Artificial Intelligence

We show that a critical problem in adversarial imitation from high-dimensional sensory data is the tendency of discriminator networks to distinguish agent and expert behaviour using task-irrelevant features beyond the control of the agent. We analyze this problem in detail and propose a solution as well as several baselines that outperform standard Generative Adversarial Imitation Learning (GAIL). Our proposed solution, Task-Relevant Adversarial Imitation Learning (TRAIL), uses a constrained optimization objective to overcome task-irrelevant features. Comprehensive experiments show that TRAIL can solve challenging manipulation tasks from pixels by imitating human operators, where other agents such as behaviour cloning (BC), standard GAIL, improved GAIL variants including our newly proposed baselines, and Deterministic Policy Gradients from Demonstrations (DPGfD) fail to find solutions, even when the other agents have access to task reward.


Stabilizing Off-Policy Reinforcement Learning with Conservative Policy Gradients

arXiv.org Artificial Intelligence

In recent years, advances in deep learning have enabled the application of reinforcement learning algorithms in complex domains. However, they lack the theoretical guarantees which are present in the tabular setting and suffer from many stability and reproducibility problems \citep{henderson2018deep}. In this work, we suggest a simple approach for improving stability and providing probabilistic performance guarantees in off-policy actor-critic deep reinforcement learning regimes. Experiments on continuous action spaces, in the MuJoCo control suite, show that our proposed method reduces the variance of the process and improves the overall performance.


TE-ETH: Lower Bounds for QBFs of Bounded Treewidth

arXiv.org Artificial Intelligence

The problem of deciding the validity (QSAT) of quantified Boolean formulas (QBF) is a vivid research area in both theory and practice. In the field of parameterized algorithmics, the well-studied graph measure treewidth turned out to be a successful parameter. A well-known result by Chen in parameterized complexity is that QSAT when parameterized by the treewidth of the primal graph of the input formula together with the quantifier depth of the formula is fixed-parameter tractable. More precisely, the runtime of such an algorithm is polynomial in the formula size and exponential in the treewidth, where the exponential function in the treewidth is a tower, whose height is the quantifier depth. A natural question is whether one can significantly improve these results and decrease the tower while assuming the Exponential Time Hypothesis (ETH). In the last years, there has been a growing interest in the quest of establishing lower bounds under ETH, showing mostly problem-specific lower bounds up to the third level of the polynomial hierarchy. Still, an important question is to settle this as general as possible and to cover the whole polynomial hierarchy. In this work, we show lower bounds based on the ETH for arbitrary QBFs parameterized by treewidth (and quantifier depth). More formally, we establish lower bounds for QSAT and treewidth, namely, that under ETH there cannot be an algorithm that solves QSAT of quantifier depth i in runtime significantly better than i-fold exponential in the treewidth and polynomial in the input size. In doing so, we provide a versatile reduction technique to compress treewidth that encodes the essence of dynamic programming on arbitrary tree decompositions. Further, we describe a general methodology for a more fine-grained analysis of problems parameterized by treewidth that are at higher levels of the polynomial hierarchy.


Clinical Text Generation through Leveraging Medical Concept and Relations

arXiv.org Artificial Intelligence

With a neural sequence generation model, this study aims to develop a method of writing the patient clinical text s given a brief medical history. As a proof - of - a - concept, we have demonstrated that it can be workable t o use medical concept embedding in clinical text generation . Our model was based on the Sequence - to - Sequence architecture and trained with a large set of de - identified clinical text data . T he quantitative result shows that our concept embedding method decr eased the perplexity of the baseline architecture . Also, we discuss the analyzed r esults from a human evaluation performed by medical doctors .


Variational Temporal Abstraction

arXiv.org Artificial Intelligence

We introduce a variational approach to learning and inference of temporally hierarchical structure and representation for sequential data. We propose the Variational Temporal Abstraction (VTA), a hierarchical recurrent state space model that can infer the latent temporal structure and thus perform the stochastic state transition hierarchically. We also propose to apply this model to implement the jumpy-imagination ability in imagination-augmented agent-learning in order to improve the efficiency of the imagination. In experiments, we demonstrate that our proposed method can model 2D and 3D visual sequence datasets with interpretable temporal structure discovery and that its application to jumpy imagination enables more efficient agent-learning in a 3D navigation task.


CWAE-IRL: Formulating a supervised approach to Inverse Reinforcement Learning problem

arXiv.org Artificial Intelligence

Inverse reinforcement learning (IRL) is used to infer the reward function from the actions of an expert running a Markov Decision Process (MDP). A novel approach using variational inference for learning the reward function is proposed in this research. Using this technique, the intractable posterior distribution of the continuous latent variable (the reward function in this case) is analytically approximated to appear to be as close to the prior belief while trying to reconstruct the future state conditioned on the current state and action. The reward function is derived using a well-known deep generative model known as Conditional Variational Auto-encoder (CVAE) with Wasserstein loss function, thus referred to as Conditional Wasserstein Auto-encoder-IRL (CWAE-IRL), which can be analyzed as a combination of the backward and forward inference. This can then form an efficient alternative to the previous approaches to IRL while having no knowledge of the system dynamics of the agent. Experimental results on standard benchmarks such as objectworld and pendulum show that the proposed algorithm can effectively learn the latent reward function in complex, high-dimensional environments.


Acoustic: Artificial intelligence shouldn't be an afterthought - AdNews

#artificialintelligence

Artificial intelligence shouldn't be an afterthought for marketers, says new martech vendor Acoustic. Formerly known as IBM Watson Marketing, Acoustic was sold off by tech giants IBM earlier this year and rebranded with a sole focus on marketers. The independent marketing cloud company has since rebuilt a cloud that it says has a "modern architecture" and brings "humanity" to AI-powered marketing. Jay Henderson, senior vice president of product management at Acoustic, says at the moment a lot of the company's competitors aren't enabling marketers to use AI easily. "I think marketers, generally at the moment, have a little bit of fatigue around the way we've all talked about AI," Henderson says.


Capacity building in artificially intelligent mining systems University of Nevada, Reno

#artificialintelligence

Mining companies from around the world have begun using artificial intelligence in their operations. From safety and maintenance, to exploration and autonomous vehicles, and drills, AI is being used to navigate efficiencies and speed. With this new technology, however, comes an ever-growing need for a workforce who can navigate these new systems. Thanks to a $1.25 million grant from the National Institute for Occupational Safety and Health, an interdisciplinary team at the University of Nevada, Reno, has committed to graduating six doctoral and four master's degree students who will address several challenges related to major safety and health issues in mining operations. "Future mine engineers need to understand emerging technology like AI, drones and big data," Javad Sattarvand, University College of Science assistant professor of mining engineering and the project's principal investigator, said.



Amazon releases Alexa data set to help solve the 'cocktail party problem'

#artificialintelligence

The cocktail party problem, alternatively known as the dinner party problem, is the difficulty automated systems encounter when tasked with isolating audio in noisy, multisource environments. It's widely studied, and a number of academic teams, startups, and corporate giants claim to have solved it with sophisticated machine learning algorithms. But Amazon believes there's room for improvement, and to this end, it's releasing a data set -- the Dinner Party Corpus, or DiPCo -- intended to spur research on the topic. According to Zaid Ahmed, a senior technical program manager in the Alexa Speech group, the corpus was created with the help of Amazon volunteers who simulated a dinner-party scenario in the lab. Over the course of multiple sessions (each involving four participants), the volunteers served themselves food from a buffet table and spoke over music piped into the room.