Goto

Collaborating Authors

 Education


MACHINE LEARNING MONDAY – TensorFlow 2.0.0 released @tensorflow #machinelearning #tensorflow

#artificialintelligence

Adafruit's Circuit Playground is jam-packed with LEDs, sensors, buttons, alligator clip pads and more. Build projects with Circuit Playground in a few minutes with the drag-and-drop MakeCode programming site, learn computer science using the CS Discoveries class on code.org, It has a powerful processor, 10 NeoPixels, mini speaker, InfraRed receive and transmit, two buttons, a switch, 14 alligator clip pads, and lots of sensors: capacitive touch, IR proximity, temperature, light, motion and sound. A whole wide world of electronics and coding is waiting for you, and it fits in the palm of your hand. Join 14,000 makers on Adafruit's Discord channels and be part of the community!


An introduction to flexible methods for policy evaluation

arXiv.org Machine Learning

This chapter covers different approaches to policy evaluation for assessing the causal effect of a treatment or intervention on an outcome of interest. As an introduction to causal inference, the discussion starts with the experimental evaluation of a randomized treatment. It then reviews evaluation methods based on selection on observables (assuming a quasi-random treatment given observed covariates), instrumental variables (inducing a quasi-random shift in the treatment), difference-in-differences and changes-in-changes (exploiting changes in outcomes over time), as well as regression discontinuities and kinks (using changes in the treatment assignment at some threshold of a running variable). The chapter discusses methods particularly suited for data with many observations for a flexible (i.e. semi- or nonparametric) modeling of treatment effects, and/or many (i.e. high dimensional) observed covariates by applying machine learning to select and control for covariates in a data-driven way. This is not only useful for tackling confounding by controlling for instance for factors jointly affecting the treatment and the outcome, but also for learning effect heterogeneities across subgroups defined upon observable covariates and optimally targeting those groups for which the treatment is most effective.


Understanding Early Word Learning in Situated Artificial Agents

arXiv.org Artificial Intelligence

Neural network-based systems can now learn to locate the referents of words and phrases in images, answer questions about visual scenes, and execute symbolic instructions as first-person actors in partially-observable worlds. To achieve this so-called grounded language learning, models must overcome challenges that infants face when learning their first words. While it is notable that models with no meaningful prior knowledge overcome these obstacles, researchers currently lack a clear understanding of how they do so, a problem that we attempt to address in this paper. For maximum control and generality, we focus on a simple neural network-based language learning agent, trained via policy-gradient methods, which can interpret single-word instructions in a simulated 3D world. Whilst the goal is not to explicitly model infant word learning, we take inspiration from experimental paradigms in developmental psychology and apply some of these to the artificial agent, exploring the conditions under which established human biases and learning effects emerge. We further propose a novel method for visualising semantic representations in the agent.


The First Google Doodle in 1998 Was a 'Bit of a Joke.' Here's the Story Behind the Design That Started it All

TIME - Tech

When Google co-founders Larry Page and Sergey Brin were headed to Nevada's Burning Man festival in August of 1998, they wanted users and employees to know they wouldn't be at the search engine's helm for a while. The Ph.D. students at Stanford University decided to replace the second'O' in Google's homepage logo with a stick figure resembling the festival's logo. "It was a little bit of a joke," Jessica Yu, the Google Doodle team lead, tells TIME. "It has definitely evolved a lot since then." What began as a joke became Google Doodles that celebrate and honor holidays, people and issues worldwide, now an important venture for the tech giant.


The CS Teacher Shortage

Communications of the ACM

The only exposure Yancarlos Diaz had to computer science during his high school years in New York City was when he used a computer to write essays. When it came time to apply to college, Diaz, who says he was good in math, "blindly signed up" for the computer science program at the Rochester Institute of Technology (RIT), figuring it was a major that would help him easily find a job when he graduated. That decision already is paying off. Now a fourth-year student at RIT, Diaz expects to graduate in 2021 with dual bachelor and master of science degrees in computer science (CS). He then plans to work in the private sector as a software engineer "mainly to pay the loans," he says.


UC Explores Artificial Intelligence in Film Series at Esquire Theatre

#artificialintelligence

A programmer at an internet-search company, Caleb Smith, wins a competition to spend a week at the CEO's estate, tucked away in the mountains. But he learns that he was chosen to be the human half of a Turing Test, a method of determining if a computer is capable of thinking like a real person. Tasked with evaluating Ava, a beautiful cutting-edge robot, they soon find that she is capable of more than they could have conceived. Her - Oct. 14, 7 p.m. Set in a near future in sunny Los Angeles, 2013's Her -- Oscar winner for Best Original Screenplay -- follows Theodore Twombly, a recently-divorced man who writes personal letters for other people for a living. Lonely and mourning his relationship, he begins using an advanced operating system (voiced by Scarlett Johansson) -- think Amazon's Alexa or Google Home -- to which he forms a strong bond.


Gated Linear Networks

arXiv.org Machine Learning

This paper presents a family of backpropagation-free neural architectures, Gated Linear Networks (GLNs),that are well suited to online learning applications where sample efficiency is of paramount importance. The impressive empirical performance of these architectures has long been known within the data compression community, but a theoretically satisfying explanation as to how and why they perform so well has proven difficult. What distinguishes these architectures from other neural systems is the distributed and local nature of their credit assignment mechanism; each neuron directly predicts the target and has its own set of hard-gated weights that are locally adapted via online convex optimization. By providing an interpretation, generalization and subsequent theoretical analysis, we show that sufficiently large GLNs are universal in a strong sense: not only can they model any compactly supported, continuous density function to arbitrary accuracy, but that any choice of no-regret online convex optimization technique will provably converge to the correct solution with enough data. Empirically we show a collection of single-pass learning results on established machine learning benchmarks that are competitive with results obtained with general purpose batch learning techniques.


Hamiltonian Generative Networks

arXiv.org Machine Learning

The Hamiltonian formalism plays a central role in classical and quantum physics. Hamiltonians are the main tool for modelling the continuous time evolution of systems with conserved quantities, and they come equipped with many useful properties, like time reversibility and smooth interpolation in time. These properties are important for many machine learning problems - from sequence prediction to reinforcement learning and density modelling - but are not typically provided out of the box by standard tools such as recurrent neural networks. In this paper, we introduce the Hamiltonian Generative Network (HGN), the first approach capable of consistently learning Hamiltonian dynamics from high-dimensional observations (such as images) without restrictive domain assumptions. Once trained, we can use HGN to sample new trajectories, perform rollouts both forward and backward in time and even speed up or slow down the learned dynamics. We demonstrate how a simple modification of the network architecture turns HGN into a powerful normalising flow model, called Neural Hamiltonian Flow (NHF), that uses Hamiltonian dynamics to model expressive densities. We hope that our work serves as a first practical demonstration of the value that the Hamiltonian formalism can bring to deep learning.


ATOL: Automatic Topologically-Oriented Learning

arXiv.org Machine Learning

There are abundant cases for using Topological Data Analysis (TDA) in a learning context, but robust topological information commonly comes in the form of a set of persistence diagrams, objects that by nature are uneasy to affix to a generic machine learning framework. We introduce a vectorisation method for diagrams that allows to collect information from topological descriptors into a format fit for machine learning tools. Based on a few observations, the method is learned and tailored to discriminate the various important plane regions a diagram is set into. With this tool one can automatically augment any sort of machine learning problem with access to a TDA method, enhance performances, construct features reflecting underlying changes in topological behaviour. The proposed methodology comes with only high level tuning parameters such as the encoding budget for topological features. We provide an open-access, ready-to-use implementation and notebook. We showcase the strengths and versatility of our approach on a number of applications. From emulous and modern graph collections to a highly topological synthetic dynamical orbits data, we prove that the method matches or beats the state-of-the-art in encoding persistence diagrams to solve hard problems. We then apply our method in the context of an industrial, difficult time-series regression problem and show the approach to be relevant.


Distributed SGD Generalizes Well Under Asynchrony

arXiv.org Machine Learning

Jayanth Regatti Gaurav Tendolkar Yi Zhou Abhishek Gupta Yingbin Liang Abstract -- The performance of fully synchronized distributed systems has faced a bottleneck due to the big data trend, under which asynchronous distributed systems are becoming a major popularity due to their powerful scalability. In this paper, we study the generalization performance of stochastic gradient descent (SGD) on a distributed asynchronous system. The system consists of multiple worker machines that compute stochastic gradients which are further sent to and aggregated on a common parameter server to update the variables, and the communication in the system suffers from possible delays. Under the algorithm stability framework, we prove that distributed asynchronous SGD generalizes well given enough data samples in the training optimization. In particular, our results suggest to reduce the learning rate as we allow more asynchrony in the distributed system. Such adaptive learning rate strategy improves the stability of the distributed algorithm and reduces the corresponding generalization error . Then, we confirm our theoretical findings via numerical experiments. I NTRODUCTION Stochastic gradient descent (SGD) and its variants (e.g., Adagrad, Adam, etc) have been very effective in solving many challenging machine learning problems such as training deep neural networks. In practice, the solution found by SGD via solving an empirical risk minimization problem typically has good generalization performance on the test dataset.