Country
Balanced Off-Policy Evaluation in General Action Spaces
Sondhi, Arjun, Arbour, David, Dimmery, Drew
In many practical applications of contextual bandits, online learning is infeasible and practitioners must rely on off-policy evaluation (OPE) of logged data collected from prior policies. OPE generally consists of a combination of two components: (i) directly estimating a model of the reward given state and action and (ii) importance sampling. While recent work has made significant advances adaptively combining these two components, less attention has been paid to improving the quality of the importance weights themselves. In this work we present balancing off-policy evaluation (BOP-e), an importance sampling procedure that directly optimizes for balance and can be plugged into any OPE estimator that uses importance sampling. BOP-e directly estimates the importance sampling ratio via a classifier which attempts to distinguish state-action pairs from an observed versus a proposed policy. BOP-e can be applied to continuous, mixed, and multi-valued action spaces without modification and is easily scalable to many observations. Further, we show that minimization of regret in the constructed binary classification problem translates directly into minimizing regret in the off-policy evaluation task. Finally, we provide experimental evidence that BOP-e outperforms inverse propensity weighting-based approaches for offline evaluation of policies in the contextual bandit setting under both discrete and continuous action spaces.
Goal-conditioned Imitation Learning
Ding, Yiming, Florensa, Carlos, Phielipp, Mariano, Abbeel, Pieter
Designing rewards for Reinforcement Learning (RL) is challenging because it needs to convey the desired task, be efficient to optimize, and be easy to compute. The latter is particularly problematic when applying RL to robotics, where detecting whether the desired configuration is reached might require considerable supervision and instrumentation. Furthermore, we are often interested in being able to reach a wide range of configurations, hence setting up a different reward every time might be unpractical. Methods like Hindsight Experience Replay (HER) have recently shown promise to learn policies able to reach many goals, without the need of a reward. Unfortunately, without tricks like resetting to points along the trajectory, HER might take a very long time to discover how to reach certain areas of the state-space. In this work we investigate different approaches to incorporate demonstrations to drastically speed up the convergence to a policy able to reach any goal, also surpassing the performance of an agent trained with other Imitation Learning algorithms. Furthermore, our method can be used when only trajectories without expert actions are available, which can leverage kinestetic or third person demonstration. The code is available at https://sites.google.com/view/goalconditioned-il/ .
KCAT: A Knowledge-Constraint Typing Annotation Tool
Lin, Sheng, Zheng, Luye, Chen, Bo, Tang, Siliang, Zhuang, Yueting, Wu, Fei, Chen, Zhigang, Hu, Guoping, Ren, Xiang
Fine-grained Entity Typing is a tough task which suffers from noise samples extracted from distant supervision. Thousands of manually annotated samples can achieve greater performance than millions of samples generated by the previous distant supervision method. Whereas, it's hard for human beings to differentiate and memorize thousands of types, thus making large-scale human labeling hardly possible. In this paper, we introduce a Knowledge-Constraint Typing Annotation Tool (KCAT), which is efficient for fine-grained entity typing annotation. KCAT reduces the size of candidate types to an acceptable range for human beings through entity linking and provides a Multi-step Typing scheme to revise the entity linking result. Moreover, KCAT provides an efficient Annotator Client to accelerate the annotation process and a comprehensive Manager Module to analyse crowdsourcing annotations. Experiment shows that KCAT can significantly improve annotation efficiency, the time consumption increases slowly as the size of type set expands.
Extension of Rough Set Based on Positive Transitive Relation
The application of rough set theory in incomplete information systems is a key problem in practice since missing values almost always occur in knowledge acquisition due to the error of data measuring, the limitation of data collection, or the limitation of data comprehension, etc. An incomplete information system is mainly processed by compressing the indiscernibility relation. The existing rough set extension models based on tolerance or symmetric similarity relations typically discard one relation among the reflexive, symmetric and transitive relations, especially the transitive relation. In order to overcome the limitations of the current rough set extension models, we define a new relation called the positive transitive relation and then propose a novel rough set extension model built upon which. The new model holds the merit of the existing rough set extension models while avoids their limitations of discarding transitivity or symmetry. In comparison to the existing extension models, the proposed model has a better performance in processing the incomplete information systems while substantially reducing the computational complexity, taking into account the relation of tolerance and similarity of positive transitivity, and supplementing the related theories in accordance to the intuitive classification of incomplete information. In summary, the positive transitive relation can improve current theoretical analysis of incomplete information systems and the newly proposed extension model is more suitable for processing incomplete information systems and has a broad application prospect.
Former Defense Secretary Ash Carter On 'The Five-Sided Box'
Even though he started working in government more than 30 years ago, former Defense Secretary Ash Carter said he wouldn't work for President Donald Trump. Carter previously swore off politics but he says his desire not to work for Trump runs deep. "It's not politics, it's personal conduct," Carter said on "CBS This Morning" on Monday. Carter explained that he used to tell officers and soldiers in the military that they "needed to behave themselves" and would have fired an officer for acting as the president has. But that doesn't mean he always agreed with the presidents under whom he did serve.
Uber tests drone food delivery, launches new autonomous SUV
WASHINGTON - Uber is testing restaurant food deliveries by drone. The company's Uber Eats unit began the tests in San Diego with McDonald's and plans to expand to other restaurants later this year. Uber says the service should decrease food delivery times. It works this way: Workers at a restaurant load the meal into a drone and it takes off, tracked and guided by a new aerospace management system. The drone then meets an Uber Eats driver at a drop-off location, and the driver will hand-deliver the meal to the customer.
Forget digital versus physical: The future is programmable
Advances in mixed-reality technologies and machine learning, coupled with the development of reactive and adaptable materials, are changing the way we think about what is real and what is digital. "Programmable reality", a term recently coined by forecasting agency The Future Laboratory, refers to the growing ability for material objects to assume digital attributes. Its advancement is set to transform product and retail development over the next five to 10 years. "Tomorrow's consumer will expect the world around them to become as personalised and responsive as online experiences have been," says Future Laboratory senior foresight writer Rhiannon McGregor. This has significant implications on the ways brands create and speak about their products.
Marketing Benefits of AI and Machine Learning - Chief Marketer
Advances in machine learning and AI are revolutionizing all aspects of business and industry, but most marketers have only scratched the surface of the potential applications of this technology. In the U.S. alone, there are over 300 million potential consumers. Multiply that by the number of possible branded products in a given category, considering all different variants and configurations that are available for purchase. Then, think about how those purchase decisions are affected by previous brand interactions, time of day, weather, type of device, personal preferences, language, sentiment and more. There is no way a person could access all of these variables, integrate them into actionable insights and roll out any activity in real-time.
Alteryx Releases Assisted Modeling Tool
Recognizing the pervasive talent gap that exists between data scientists and data workers in the line of business, Assisted Modeling helps teach data science with a guided walk-through and aims to help all data workers, regardless of technical acumen, advance their skill sets in the process of building machine learning models. "As we continue to deliver innovations for a smarter platform, it is critical to address the human at the center of analytic intelligence. Our approach in building Assisted Modeling is to advance the skills of the data worker, creating next-level citizen data scientists capable of building the machine learning models required to tackle the advanced analytic challenges of the future," said Kramer. "Assisted Modeling provides users the transparency and control needed to build trustworthy machine learning models that drive business outcomes without writing a line of code. I am thrilled to deliver the newest version of our platform today and to invite our customers and partners to be the first to experience Assisted Modeling." As an output of the application, users can access code-free machine learning tools directly within the Alteryx Designer interface. Assisted Modeling allows any data worker to construct machine learning models, understand how and why their models work, and capture modeling decisions, turning raw data into informed business decisions with unprecedented speed and confidence. Alteryx also announced the general availability of the newest version of the Alteryx Platform (2019.2),
ML-powered Automated Insights are here to stay - Sweetspot
It's undeniable that Machine Learning has taken the digital world by storm. But along with its many promises, there has been much hype, making it hard to tell whether particular "AI-powered" labels actually represent true breakthroughs that will have a real impact on our day-to-day life. We certainly believe that Sweetspot's Predictions, Automated Insights, and Smart API Calls, our ML-powered features, have earned this label in their own right. While Predictions have been quietly serving our customers and partners from the very outset (2012!), the last two are only now beginning to unleash their true power. I nevertheless believe they all deserve a proper shout-out.