Goto

Collaborating Authors

 Education


Positive-Unlabeled Reward Learning

arXiv.org Machine Learning

Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond to the intended task. These reward delusions can lead to unintended and even dangerous behaviors. On the other hand, adversarial imitation learning frameworks (Ho & Ermon, 2016) tend to suffer the opposite problem, where the discriminator learns to trivially distinguish agent and expert behavior, resulting in reward models that produce low reward signal regardless of the input state. In this paper, we connect these two classes of reward learning methods to positive-unlabeled (PU) learning, and we show that by applying a large-scale PU learning algorithm to the reward learning problem, we can address both the reward under-and overestimation problems simultaneously. Our approach drastically improves both GAIL and supervised reward learning, without any additional assumptions. While Reinforcement Learning (RL) has shown itself to be a powerful tool for automating control and decision making, hand-specifying reward functions requires significant engineering effort, especially in real-world settings. Recent works have made promising progress in learning reward functions directly from human supervision, such as ratings (Cabi et al., 2019) and behavior preferences (Wilson et al., 2012; Ibarz et al., 2018). However, in practice, these supervisions are expensive to curate and thus often only cover a fraction of the state space. As a result, the learned reward functions may have large errors in the unlabeled states, and policy learning algorithms tend to exploit these errors to achieve extremely high pseudo-reward via unintended behaviors (Amodei et al., 2016). Practical solutions often require a human to provide supervision in the policy training loop iteratively (Christiano et al., 2017; Ibarz et al., 2018), resulting in a even more laborious process. On the other hand, works in Inverse Reinforcement Learning (IRL) propose to infer reward functions directly from expert behaviors (Ng et al., 2000; Ziebart et al., 2008), but scaling these methods to high-dimensional state space remains a challenge. Ho & Ermon (2016), and many followup works show that GAIL can learn complex behaviors even in high-dimensional spaces.


Robust Federated Learning with Noisy Communication

arXiv.org Machine Learning

Abstract--Federated learning is a communication-efficient training process that alternates between local training at the edge devices and averaging the updated local model at the central server . Nevertheless, it is impractical to achieve a perfect acquisition of the local models in wireless communication d ue to noise, which also brings serious effects on federated learn ing. T o tackle this challenge, we propose a robust design for federa ted learning to alleviate the effects of noise in this paper . Con sidering noise in the two aforementioned steps, we first formulate the training problem as a parallel optimization for each node un der the expectation-based model and the worst-case model. Due t o the non-convexity of the problem, a regularization for the l oss function approximation method is proposed to make it tracta ble. Regarding the worst-case model, we develop a feasible train ing scheme which utilizes the sampling-based successive conve x approximation algorithm to tackle the unavailable maxima o r minima noise condition and the non-convex issue of the objec tive function. Furthermore, the convergence rates of both new de signs are analyzed from a theoretical point of view. Finally, the improvement of prediction accuracy and the reduction of los s function are demonstrated via simulations for the proposed designs. UTURE wireless computing applications demand higher bandwidth, lower latency and more reliable connections with numerous devices [1].


BERT Goes to Law School: Quantifying the Competitive Advantage of Access to Large Legal Corpora in Contract Understanding

arXiv.org Artificial Intelligence

Fine-tuning language models, such as BERT, on domain specific corpora has proven to be valuable in domains like scientific papers and biomedical text. In this paper, we show that fine-tuning BERT on legal documents similarly provides valuable improvements on NLP tasks in the legal domain. Demonstrating this outcome is significant for analyzing commercial agreements, because obtaining large legal corpora is challenging due to their confidential nature. As such, we show that having access to large legal corpora is a competitive advantage for commercial applications, and academic research on analyzing contracts.


Cognitive Systems Institute Group Speaker Series Thursdays 10:30am US Eastern Time

#artificialintelligence

Join the meetings by pointing your web browser to: https://zoom.us/j/7371462221 Join the CSIG LinkedIn Group to get reminders about talks and discuss them. Replays before Dec 2015: Dial 877.471.6587 or 402.970.2667 and enter the call's Replay ID when prompted for a program ID number. The Replay ID is listed in the Recording column of each date. "Solving Large-Scale Machine Learning Problems in a Distributed Way"


Enterprise Grade Data Labeling - Design Your Ground Truth to Scale in Production Open Data Science Conference

#artificialintelligence

Abstract: Wherever you are in your team's machine learning journey, it's helpful to think about evolving towards large scale production. A key ingredient of this journey is your data labeling and annotation framework. In this talk we focus on how to build your data labeling pipeline to be enterprise grade. We will describe the considerations and insights that go into making your data pipeline a mindful part of your development pipeline. Proactively planning a data process can generate progressively better results during development, but it requires some thought and stakeholder buy-in.


Machine learning in 30 minutes? Are you kidding me?

#artificialintelligence

Shortage of true data-scientists and the increasing demand of AI-ML world, enforced companies like Microsoft, Google, H2O.ai and Data Robots to automate even the machine learning process to make it simpler for people like me or you who has been recruiting novice data scientists and spending money and more importantly time to implement a simple work in the AI-ML world. A few years back when I was in India, everyone coming in the IT field from any stream of education, used to claim themselves as "Engineers". Today's scenario is kind of the same, as many of us are not at all a true data scientist. Just by doing a 12 to 16 months course and learning a few algorithms, one doesn't become "a Scientist". A true scientist is born and grown with a passion for science and that can't be achieved just by doing an online course.


Best Selling Machine Learning Course on the Internet - 2.45 million Enrollments

#artificialintelligence

If you are looking for Machine learning courses or certification, consider looking at the one offered by Stanford University on Coursera (View here). Instructor of the course is Andrew Ng, the biggest names of online teaching space and the co-founder of Coursera. This course is probably the best selling Machine learning course on the internet at the moment! The rating of the course 4.9/5 after 109,078 ratings, and 2.45 million enrollments totally confirm my claim. This Stanford University course, taught is 11 Weeks long.


On EducationDigishock 2.0: Machine Learning for Beginners (No Coding) - CouponED

#artificialintelligence

Learn the basics of machine learning without using code Learn to teach a machine with a camera Use an AI platform to build AI Models and Train the datasets Know about IBM Watson & Wipro Holmes AI technologies Convert a web application/software to an app in less than a minute Digishock 1.0 course from Udemy is a must in order to understand the tools better. No other experience or technical knowledge is necessary. This mind-blowing course takes the huge leap from Digishock 1.0 and is for anyone who want to get introduced with Machine Learning and Deep Learning without learning code. This practical hands-on course involves hands-on exercises with numerous tricks and techniques of analytics, advanced predictive concepts to work on to ensure that all are familiarized with the discipline of machine-learning, deep-learning, big data, analytics etc. The USP of the course is that there is no kind of technical knowledge required whatsoever for students who will participate in this course.


Machine learning Training in Hyderabad

#artificialintelligence

Work towards building a strong knowledge based career foundation in the leading analytics platform of Machine Learning by availing our Analytics Path top-notch Machine Learning Training In Hyderabad. Our experts trainers will be working towards transforming our students into complete career ready professionals. By the time of course completion, our students will become well capable to handling all the real-world complex challenges of the Machine Learning domain. Students will be gaining expertise towards working on the advanced concepts like Support Vector Machines, Naive Bayes Classification, Logistic Regression, Decision Tree Algorithms, K-Means Clustering and more. Machine Learning is the most challenging & innovative platform in the present days analytics domain.


Amazon Alexa Skills and Google Assistant Actions: The Future of Search

#artificialintelligence

In 2018 smart speaker ownership has doubled in the UK. The most popular voice assistant is Amazon Alexa which is part of the Alexa Echo device, followed by Google Assistant which powers Google Home and Google Home Mini, Apple's Home Pod and Sonos One. At the moment people mainly use voice assistants to play music, answer questions, set alarm and reminders. Is there an opportunity there for marketers? Alexa provides a set of built-in capabilities, called skills.