Reinforcement Learning
Rover Descent: Learning to optimize by learning to navigate on prototypical loss surfaces
Learning to optimize - the idea that we can learn from data algorithms that optimize a numerical criterion - has recently been at the heart of a growing number of research efforts. One of the most challenging issues within this approach is to learn a policy that is able to optimize over classes of functions that are fairly different from the ones that it was trained on. We propose a novel way of framing learning to optimize as a problem of learning a good navigation policy on a partially observable loss surface. To this end, we develop Rover Descent, a solution that allows us to learn a fairly broad optimization policy from training on a small set of prototypical two-dimensional surfaces that encompasses the classically hard cases such as valleys, plateaus, cliffs and saddles and by using strictly zero-order information. We show that, without having access to gradient or curvature information, we achieve state-of-the-art convergence speed on optimization problems not presented at training time such as the Rosenbrock function and other hard cases in two dimensions. We extend our framework to optimize over high dimensional landscapes, while still handling only two-dimensional local landscape information and show good preliminary results.
Predict Responsibly: Increasing Fairness by Learning To Defer
Madras, David, Pitassi, Toniann, Zemel, Richard
In many high-stakes ML applications, there are multiple decision-makers involved, both automated and human. The interaction between these agents often goes unaddressed in algorithmic development. In this work, we explore a simple version of this interaction with a two-stage framework containing an automated model and an external decision-maker. The model can choose to say IDK, and pass the decision downstream, as explored in rejection learning. We extend this concept by proposing learning to defer, which generalizes the rejection learning framework by considering the effect of the other agents in the decision-making process. We propose a learning algorithm which accounts for potential biases held by external decision-makers in a system. Experiments on real-world datasets demonstrate that learning to defer can make a system not only more accurate but also less biased. Even when operated by highly biased users, we show that deferring models can still greatly improve the fairness of the entire system.
What is Artificial General Intelligence? โ Towards Data Science
Artificial Intelligence is a branch of Computer Science ( or Science) which deals with the creation of intelligent systems. Intelligent systems are those systems which posses intelligence just like humans. The science of Artificial intelligence is not new, The term Artificial intelligence has been mentioned in manuscripts of Ancient Greece and Egypt. Greeks believed in god Hephaestus, also known as God of Blacksmiths, according to a Greek mythology Hephaestus made intelligent weapons for all Gods, in their view, the goal of Artificial intelligence is to: be helpful for people to achieve a certain goal, be able to operate automatically and be programmed in advance to react in different ways depending on the situation. Well, The term Artificial Intelligence has become popular in the field of Entertainment, we can see lots of movies based on the concept of Super intelligence.
Reinforcement learning woes, robot doggos, Amazon's homegrown AI chips, and more
Here's a brief roundup of some interesting news from the AI world from the past two weeks, beyond what we've already reported. TL;DR: Deep RL sucks โ A Google engineer has published a long, detailed blog post explaining the current frustrations in deep reinforcement learning, and why it doesn't live up to the hype. Reinforcement learning makes good headlines. Teaching agents to play games like Go well enough to beat human experts like Ke Jie fuels the man versus machine narrative. But a closer look at deep reinforcement learning, a method of machine learning used to train computers to complete a specific task, shows the practice is riddled with problems.
Introduction to Various Reinforcement Learning Algorithms. Part I (Q-Learning, SARSA, DQN, DDPG)
Typically, a RL setup is composed of two components, an agent and an environment. Then environment refers to the object that the agent is acting on (e.g. the game itself in the Atari game), while the agent represents the RL algorithm. The environment starts by sending a state to the agent, which then based on its knowledge to take an action in response to that state. After that, the environment send a pair of next state and reward back to the agent. The agent will update its knowledge with the reward returned by the environment to evaluate its last action.
Inverse Reinforcement Learning Tutorial part I
In this blog post series we will take a closer look at inverse reinforcement learning (IRL) which is the field of learning an agent's objectives, values, or rewards by observing its behavior. For example, we might observe the behavior of a human in some specific task and learn which states of the environment the human is trying to achieve and what the concrete goals might be. This is the first part of this series in which we will get an overview of IRL and look at three basic algorithms to solve the IRL problem. In later parts we will explore more advanced techniques and state of the art methods See section IRL Algorithms . To follow this tutorial, basic knowledge in reinforcement learning (RL) is required.
[D] When is it reasonable to drop the "AI" term? โข r/MachineLearning
I see more Machine Learning / AI articles popping up in my national news sources. The term "Artificial Intelligence" is used a lot, needless to say. Even researchers and unis use the term(to gain hype probably since the word is more sexy than "regression analysis") Even if the article is on something sensible, like: "It might be smart to collect some medical data, to cure new diseases, get better healthcare", the public perception / comments section seems to be heavily influenced by peoples SciFi expectations of the word "AI" (It will turn on us, etc.) My question is: When is it reasonable to use the term AI? I once read that the line is drawn at reinforcement learning, is this reasonable?
Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning
Liang, Qingkai, Que, Fanyu, Modiano, Eytan
Constrained Markov Decision Process (CMDP) is a natural framework for reinforcement learning tasks with safety constraints, where agents learn a policy that maximizes the long-term reward while satisfying the constraints on the long-term cost. A canonical approach for solving CMDPs is the primal-dual method which updates parameters in primal and dual spaces in turn. Existing methods for CMDPs only use on-policy data for dual updates, which results in sample inefficiency and slow convergence. In this paper, we propose a policy search method for CMDPs called Accelerated Primal-Dual Optimization (APDO), which incorporates an off-policy trained dual variable in the dual update procedure while updating the policy in primal space with on-policy likelihood ratio gradient. Experimental results on a simulated robot locomotion task show that APDO achieves better sample efficiency and faster convergence than state-of-the-art approaches for CMDPs.
Improving Mild Cognitive Impairment Prediction via Reinforcement Learning and Dialogue Simulation
Tang, Fengyi, Lin, Kaixiang, Uchendu, Ikechukwu, Dodge, Hiroko H., Zhou, Jiayu
Mild cognitive impairment (MCI) is a prodromal phase in the progression from normal aging to dementia, especially Alzheimers disease. Even though there is mild cognitive decline in MCI patients, they have normal overall cognition and thus is challenging to distinguish from normal aging. Using transcribed data obtained from recorded conversational interactions between participants and trained interviewers, and applying supervised learning models to these data, a recent clinical trial has shown a promising result in differentiating MCI from normal aging. However, the substantial amount of interactions with medical staff can still incur significant medical care expenses in practice. In this paper, we propose a novel reinforcement learning (RL) framework to train an efficient dialogue agent on existing transcripts from clinical trials. Specifically, the agent is trained to sketch disease-specific lexical probability distribution, and thus to converse in a way that maximizes the diagnosis accuracy and minimizes the number of conversation turns. We evaluate the performance of the proposed reinforcement learning framework on the MCI diagnosis from a real clinical trial. The results show that while using only a few turns of conversation, our framework can significantly outperform state-of-the-art supervised learning approaches.
Sim-To-Real Optimization Of Complex Real World Mobile Network with Imperfect Information via Deep Reinforcement Learning from Self-play
Tan, Yongxi, Yang, Jin, Chen, Xin, Song, Qitao, Chen, Yunjun, Ye, Zhangxiang, Su, Zhenqiang
Mobile network that millions of people use every day is one of the most complex systems in real world. Optimization of mobile network to meet exploding customer demand and reduce CAPEX/OPEX poses greater challenges than in prior works. Learning to solve complex problems in real world to benefit everyone and make the world better has long been ultimate goal of AI. However, it still remains an unsolved problem for deep reinforcement learning (DRL), given imperfect information in real world, huge state/action space, lots of data needed for training, associated time/cost, multi-agent interactions, potential negative impact to real world, etc. To bridge this reality gap, we proposed a DRL framework to direct transfer optimal policy learned from multi-tasks in source domain to unseen similar tasks in target domain without any further training in both domains. First, we distilled temporal-spatial relationships between cells and mobile users to scalable 3D image-like tensor to best characterize partially observed mobile network. Second, inspired by AlphaGo, we used a novel self-play mechanism to empower DRL agent to gradually improve its intelligence by competing for best record on multiple tasks. Third, a decentralized DRL method is proposed to coordinate multi-agents to compete and cooperate as a team to maximize global reward and minimize potential negative impact. Using 7693 unseen test tasks over 160 unseen simulated mobile networks and 6 field trials over 4 commercial mobile networks in real world, we demonstrated the capability of our approach to direct transfer the learning from one simulator to another simulator, and from simulation to real world. This is the first time that a DRL agent successfully transfers its learning directly from simulation to very complex real world problems with incomplete and imperfect information, huge state/action space and multi-agent interactions.