Goto

Collaborating Authors

 Education


Pre-trained Word Embeddings for Goal-conditional Transfer Learning in Reinforcement Learning

arXiv.org Artificial Intelligence

Reinforcement learning (RL) algorithms typically start tabula rasa, without any prior knowledge of the environment, and without any prior skills. This however often leads to low sample efficiency, requiring a large amount of interaction with the environment. This is especially true in a lifelong learning setting, in which the agent needs to continually extend its capabilities. In this paper, we examine how a pre-trained task-independent language model can make a goal-conditional RL agent more sample efficient. We do this by facilitating transfer learning between different related tasks. We experimentally demonstrate our approach on a set of object navigation tasks.


Learning to Play Sequential Games versus Unknown Opponents

arXiv.org Artificial Intelligence

We consider a repeated sequential game between a learner, who plays first, and an opponent who responds to the chosen action. We seek to design strategies for the learner to successfully interact with the opponent. While most previous approaches consider known opponent models, we focus on the setting in which the opponent's model is unknown. To this end, we use kernel-based regularity assumptions to capture and exploit the structure in the opponent's response. We propose a novel algorithm for the learner when playing against an adversarial sequence of opponents. The algorithm combines ideas from bilevel optimization and online learning to effectively balance between exploration (learning about the opponent's model) and exploitation (selecting highly rewarding actions for the learner). Our results include algorithm's regret guarantees that depend on the regularity of the opponent's response and scale sublinearly with the number of game rounds. Moreover, we specialize our approach to repeated Stackelberg games, and empirically demonstrate its effectiveness in a traffic routing and wildlife conservation task


A Computational Separation between Private Learning and Online Learning

arXiv.org Machine Learning

A recent line of work has shown a qualitative equivalence between differentially private PAC learning and online learning: A concept class is privately learnable if and only if it is online learnable with a finite mistake bound. However, both directions of this equivalence incur significant losses in both sample and computational efficiency. Studying a special case of this connection, Gonen, Hazan, and Moran (NeurIPS 2019) showed that uniform or highly sample-efficient pure-private learners can be time-efficiently compiled into online learners. We show that, assuming the existence of one-way functions, such an efficient conversion is impossible even for general pure-private learners with polynomial sample complexity. This resolves a question of Neel, Roth, and Wu (FOCS 2019).


BERT Learns (and Teaches) Chemistry

arXiv.org Machine Learning

Modern computational organic chemistry is becoming increasingly data-driven. There remain a large number of important unsolved problems in this area such as product prediction given reactants, drug discovery, and metric-optimized molecule synthesis, but efforts to solve these problems using machine learning have also increased in recent years. In this work, we propose the use of attention to study functional groups and other property-impacting molecular substructures from a data-driven perspective, using an transformer-based model (BERT) on datasets of string representations of molecules and analyzing the behavior of its attention heads. We then apply the representations of functional groups and atoms learned by the model to tackle problems of toxicity, solubility, drug-likeness, and synthesis accessibility on smaller datasets using the learned representations as features for graph convolution and attention models on the graph structure of molecules, as well as fine-tuning of BERT. Finally, we propose the use of attention visualization as a helpful tool for chemistry practitioners and students to quickly identify important substructures in various chemical properties.


Transformations between deep neural networks

arXiv.org Machine Learning

We propose to test, and when possible establish, an equivalence between two different artificial neural networks by attempting to construct a data-driven transformation between them, using manifold-learning techniques. In particular, we employ diffusion maps with a Mahalanobis-like metric. If the construction succeeds, the two networks can be thought of as belonging to the same equivalence class. We first discuss transformation functions between only the outputs of the two networks; we then also consider transformations that take into account outputs (activations) of a number of internal neurons from each network. In general, Whitney's theorem dictates the number of measurements from one of the networks required to reconstruct each and every feature of the second network. The construction of the transformation function relies on a consistent, intrinsic representation of the network input space. We illustrate our algorithm by matching neural network pairs trained to learn (a) observations of scalar functions; (b) observations of two-dimensional vector fields; and (c) representations of images of a moving three-dimensional object (a rotating horse). The construction of such equivalence classes across different network instantiations clearly relates to transfer learning. We also expect that it will be valuable in establishing equivalence between different Machine Learning-based models of the same phenomenon observed through different instruments and by different research groups.


The Computational Limits of Deep Learning

arXiv.org Machine Learning

Deep learning's recent history has been one of achievement: from triumphing over humans in the game of Go to world-leading performance in image recognition, voice recognition, translation, and other tasks. But this progress has come with a voracious appetite for computing power. This article reports on the computational demands of Deep Learning applications in five prominent application areas and shows that progress in all five is strongly reliant on increases in computing power. Extrapolating forward this reliance reveals that progress along current lines is rapidly becoming economically, technically, and environmentally unsustainable. Thus, continued progress in these applications will require dramatically more computationally-efficient methods, which will either have to come from changes to deep learning or from moving to other machine learning methods.


Online unsupervised deep unfolding for massive MIMO channel estimation

arXiv.org Machine Learning

Massive MIMO communication systems have a huge potential both in terms of data rate and energy efficiency, although channel estimation becomes challenging for a large number antennas. Using a physical model allows to ease the problem by injecting a priori information based on the physics of propagation. However, such a model rests on simplifying assumptions and requires to know precisely the configuration of the system, which is unrealistic in practice. In this letter, we propose to perform online learning for channel estimation in a massive MIMO context, adding flexibility to physical channel models by unfolding a channel estimation algorithm (matching pursuit) as a neural network. This leads to a computationally efficient neural network structure that can be trained online when initialized with an imperfect model. The method allows a base station to automatically correct its channel estimation algorithm based on incoming data, without the need for a separate offline training phase. It is applied to realistic millimeter wave channels and shows great performance, achieving a channel estimation error almost as low as one would get with a perfectly calibrated system.


Using machine learning to stay connected

#artificialintelligence

Users can create a room in app, invite friends or family and users can sing along in karaoke-style without the vocals. Music students, teachers, or anyone who wants to improve their singing skills can also use it for solo practice. If you're shy about singing karaoke, you can enter a private or solo room and listen to the isolated song vocals. To use MusicBucket for a karaoke social, invite your friends to a karaoke room. They can select a song or upload their own song and start singing, as shown in the following screenshot. Users take turns singing songs with the background music.


LSTM Build your deep learning portfolio: Meditations with

#artificialintelligence

Build your deep learning portfolio: Meditations with LSTM 3.7 (7 ratings) Course Ratings are calculated from individual students' ratings and a variety of other signals, like age of rating and reliability, to ensure that they reflect course quality fairly and accurately. Build your deep learning portfolio: Meditations with LSTM With the help of this course you can In this deep learning tutorial I will teach you how to build an LSTM model, which generates text.. This course was created by David C. It was rated 4.2 out of 5 by approx 13447 ratings. The best Deep Learning courses online & Tutorials to Learn Deep Learning courses for beginners to advanced level. Deep learning courses is an artificial intelligence function that imitates the workings of the human brain in processing data and creating patterns for use in decision making.


Top 8 Free Math Courses For Aspiring Data Scientists

#artificialintelligence

Proficiency in mathematics is essential for aspirants to get started with their data science journey. A strong foundation in mathematics will help beginners to not only learn existing and new machine learning techniques easily but also differentiate themselves from others in the competitive market. Consequently, data science aspirants must ensure that they master algebra, calculus, probability, among others before diving deep into machine learning. Here are top courses on mathematics that aspiring data scientists must take into account while devising their learning strategy. The five-week-long course on Coursera can be the starting point for learners as linear algebra has a wide range of applications in data science practices.