Goto

Collaborating Authors

 Education


Option Encoder: A Framework for Discovering a Policy Basis in Reinforcement Learning

arXiv.org Artificial Intelligence

Option discovery and skill acquisition frameworks are integral to the functioning of a Hierarchically organized Reinforcement learning agent. However, such techniques often yield a large number of options or skills, which can potentially be represented succinctly by filtering out any redundant information. Such a reduction can reduce the required computation while also improving the performance on a target task. In order to compress an array of option policies, we attempt to find a policy basis that accurately captures the set of all options. In this work, we propose Option Encoder, an auto-encoder based framework with intelligently constrained weights, that helps discover a collection of basis policies. The policy basis can be used as a proxy for the original set of skills in a suitable hierarchically organized framework. We demonstrate the efficacy of our method on a collection of grid-worlds and on the high-dimensional Fetch-Reach robotic manipulation task by evaluating the obtained policy basis on a set of downstream tasks.


Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning

arXiv.org Artificial Intelligence

We explore fixed-horizon temporal difference (TD) methods, reinforcement learning algorithms for a new kind of value function that predicts the sum of rewards over a $\textit{fixed}$ number of future time steps. To learn the value function for horizon $h$, these algorithms bootstrap from the value function for horizon $h-1$, or some shorter horizon. Because no value function bootstraps from itself, fixed-horizon methods are immune to the stability problems that plague other off-policy TD methods using function approximation (also known as "the deadly triad"). Although fixed-horizon methods require the storage of additional value functions, this gives the agent additional predictive power, while the added complexity can be substantially reduced via parallel updates, shared weights, and $n$-step bootstrapping. We show how to use fixed-horizon value functions to solve reinforcement learning problems competitively with methods such as Q-learning that learn conventional value functions. We also prove convergence of fixed-horizon temporal difference methods with linear and general function approximation. Taken together, our results establish fixed-horizon TD methods as a viable new way of avoiding the stability problems of the deadly triad.


The Best of This Week From the Editors

#artificialintelligence

A secret, pre-CIA field manual for sabotaging enemy organizations during WWII identified two ways of undermining an organization: physical damage (think: pulling out wires and destroying equipment) and human obstruction of processes (think: all the dysfunctional habits companies still struggle with today). When he shares the list of sabotaging tactics with executives today, HBS professor Stefan Thomke finds "that their reaction usually starts with laughter ('I see this in my company.'), Up until last Wednesday, the answer was probably "no." In fact, for years, even the most sophisticated AI systems have been unable to match the language and logic skills students are expected to have mastered before entering high school. But last week (just in time for back to school), the Allen Institute for Artificial Intelligence announced big news about its new system, Aristo -- it passed.


Tests Show That Voice Assistants Still Lack Critical Intelligence

#artificialintelligence

Increasingly, voice assistants from vendors such as Amazon, Apple, Google, Microsoft, and others are starting to find their way into myriad of devices, products, and tools used on a daily basis. While once we might have only interacted with conversational systems on our phones, dedicated desktop appliances, or desktop computers, we can now find conversational interfaces on a wide range of appliances and products from televisions to cars and even toaster ovens. Soon, any device we can interact with will have an audio conversational interface instead of buttons or screens to type or click. The dawn of the conversational computing age is here. However, are these devices intelligent enough to handle the wide range of queries that humans are posing?


Python AI and Machine Learning for Production & Development

#artificialintelligence

When you want to learn a new technology for professional use, there are two mutually exclusive options, either you learn it yourself or you go for instructor based training. Self learning is least expensive but lot of time results in wasting time in finding right contents, setting up the environment, troubleshooting issues and may make you give up in the middle. Instructor based training can be expensive at times and need your time commitment. This course combines the best of both these options. The course is based on one of the most famous books in the field "Python Machine Learning (2nd Ed.)" by Sebastian Raschka and Vahid Mirjalili and provides you video tutorials on how to understand the AI/ML concepts from the books by providing out of box virtual machine with demo examples for each chapter in the book and complete preinstalled setup to execute the code.



Artificial-intelligence voice is used in a theft - The Washington Post

#artificialintelligence

The request was "rather strange," the director noted later in an email, but the voice was so lifelike that he felt he had no choice but to comply. The insurer, whose case was first reported by the Wall Street Journal, provided new details on the theft to The Washington Post on Wednesday, including an email from the employee tricked by what the insurer is referring to internally as "the false Johannes." Now being developed by a wide range of Silicon Valley titans and AI start-ups, such voice-synthesis software can copy the rhythms and intonations of a person's voice and be used to produce convincing speech. Tech giants such as Google and smaller firms such as the "ultrarealistic voice cloning" start-up Lyrebird have helped refine the resulting fakes and made the tools more widely available free for unlimited use. But the synthetic audio and AI-generated videos, known as "deepfakes," have fueled growing anxieties over how the new technologies can erode public trust, empower criminals and make traditional communication -- business deals, family phone calls, presidential campaigns -- that much more vulnerable to computerized manipulation.


R Vs Python Best Programming Language for Machine Learning

#artificialintelligence

Machine learning is one of the hottest skills in the upcoming years. We have advanced it very rapidly in the strong age of AI. A major number of developers are looking towards acquiring machine learning and deep learning skills. And if you are one of these people then you might get confused when it comes to choosing the right programming language for machine learning. Coding languages are becoming more versatile and each one is unique by presenting specific features which you will not find anywhere else. You only need to find the information, analyze it and then sample any one or another language to decide which is the best one or which one fulfills your needs?


How artificial intelligence is creating jobs in India, not just stealing them India News - Times of India

#artificialintelligence

There is a growing demand for data-labelling services that are "localised"- both linguistically and culturally relevant to India From an opportunity point of view, there are about a lakh jobs posted on various portals currently There is a growing demand for data-labelling services that are "localised"- both linguistically and culturally relevant to India NEW DELHI: Five years ago, Hyderabad resident Tulasi Mathi was forced to quit her job as a maths teacher due to health issues and the birth of her two children. But today, the 29-year-old does data labelling and makes up to Rs 15,000 a month. The money isn't much but it's more than she made as a teacher, and enough to pay her kids' school fees and her own expenses. Today, she scans videos and marks and labels objects encountered by self-driving cars. Her output is used to train artificial intelligence algorithms powering such cars. All Mathi knows is that it makes her life easier.