Education
Improving the convergence of SGD through adaptive batch sizes
Sievert, Scott, Charles, Zachary
Mini-batch stochastic gradient descent (SGD) approximates the gradient of an objective function with the average gradient of some batch of constant size. While small batch sizes can yield high-variance gradient estimates that prevent the model from learning a good model, large batches may require more data and computational effort. This work presents a method to change the batch size adaptively with model quality. We show that our method requires the same number of model updates as full-batch gradient descent while requiring the same total number of gradient computations as SGD. While this method requires evaluating the objective function, we present a passive approximation that eliminates this constraint and improves computational efficiency. We provide extensive experiments illustrating that our methods require far fewer model updates without increasing the total amount of computation.
Model-Agnostic Meta-Learning using Runge-Kutta Methods
Im, Daniel Jiwoong, Jiang, Yibo, Verma, Nakul
Daniel Jiwoong Im 1, Yibo Jiang 2, and Nakul Verma 3 1 Janelia Research Campus, HHMI, Virgina 2 Harvard University, Massachusetts 3 Columbia University, New York Abstract Meta learning has emerged as an important framework for learning new tasks from just a few examples. The success of any meta-learning model depends on (i) its fast adaptation to new tasks, as well as (ii) having a shared representation across similar tasks. Here we extend the model-agnostic meta-learning (MAML) framework introduced by Finn et al. (2017) to achieve improved performance by analyzing the temporal dynamics of the optimization procedure via the Runge-Kutta method. This method enables us to gain fine-grained control over the optimization and helps us achieve both the adaptation and representation goals across tasks. By leveraging this refined control, we demonstrate that there are multiple principled ways to update MAML and show that the classic MAML optimization is simply a special case of second order Runge-Kutta method that mainly focuses on fast-adaptation. Experiments on benchmark classification, regression and reinforcement learning tasks show that this refined control helps attain improved results. 1 Introduction Building an intelligent system that can learn quickly on a new task with few examples or few experiences is one of the central goals of machine learning. Achieving this goal requires an agent that learns continuously while having the ability to adapt to new tasks with limited data. Meta-learning (Biggs, 1985) has emerged as a compelling framework that strives to attain this challenging goal. There are two main approaches to meta-learning: learning-to-optimize and learning-to-initialize the meta-model (usually encoded as deep network).
A Unified Framework for Tuning Hyperparameters in Clustering Problems
Fan, Xinjie, Yue, Yuguang, Sarkar, Purnamrita, Wang, Y. X. Rachel
Selecting hyperparameters for unsupervised learning problems is difficult in general due to the lack of ground truth for validation. However, this issue is prevalent in machine learning, especially in clustering problems with examples including the Lagrange multipliers of penalty terms in semidefinite programming (SDP) relaxations and the bandwidths used for constructing kernel similarity matrices for Spectral Clustering. Despite this, there are not many provable algorithms for tuning these hyperparameters. In this paper, we provide a unified framework with provable guarantees for the above class of problems. We demonstrate our method on two distinct models. First, we show how to tune the hyperparameters in widely used SDP algorithms for community detection in networks. In this case, our method can also be used for model selection. Second, we show the same framework works for choosing the bandwidth for the kernel similarity matrix in Spectral Clustering for subgaussian mixtures under suitable model specification. In a variety of simulation experiments, we show that our framework outperforms other widely used tuning procedures in a broad range of parameter settings.
Overcoming Forgetting in Federated Learning on Non-IID Data
Shoham, Neta, Avidor, Tomer, Keren, Aviv, Israel, Nadav, Benditkis, Daniel, Mor-Yosef, Liron, Zeitak, Itai
We tackle the problem of Federated Learning in the non i.i.d. case, in which local models drift apart, inhibiting learning. Building on an analogy with Lifelong Learning, we adapt a solution for catastrophic forgetting to Federated Learning. We add a penalty term to the loss function, compelling all local models to converge to a shared optimum. We show that this can be done efficiently for communication (adding no further privacy risks), scaling with the number of nodes in the distributed setting. Our experiments show that this method is superior to competing ones for image recognition on the MNIST dataset.
RTFM: Generalising to Novel Environment Dynamics via Reading
Zhong, Victor, Rocktรคschel, Tim, Grefenstette, Edward
Obtaining policies that can generalise to new environments in reinforcement learning is challenging. In this work, we demonstrate that language understanding via a reading policy learner is a promising vehicle for generalisation to new environments. We propose a grounded policy learning problem, Read to Fight Monsters (RTFM), in which the agent must jointly reason over a language goal, relevant dynamics described in a document, and environment observations. We procedurally generate environment dynamics and corresponding language descriptions of the dynamics, such that agents must read to understand new environment dynamics instead of memorising any particular information. In addition, we propose txt2$\pi$, a model that captures three-way interactions between the goal, document, and observations. On RTFM, txt2$\pi$ generalises to new environments with dynamics not seen during training via reading. Furthermore, our model outperforms baselines such as FiLM and language-conditioned CNNs on RTFM. Through curriculum learning, txt2$\pi$ produces policies that excel on complex RTFM tasks requiring several reasoning and coreference steps.
Making the Mid-career Leap from Urban Design to Deep Learning/Data Science
As an architect with a focus on urban design, Legg Yeung realized the limitations of her impact-driven work given the traditionally creative way of framing solutions. This inspired her to make a leap towards a more data driven career, going back to school at UC Berkeley's School of Information to gain new skills with the vision of bringing more quantitative science and deep learning to the field of architecture and urban design. After working hard at developing new skills, she recently landed a resident position at Microsoft Research AI.
How can quantum computing be useful for Machine Learning - KDnuggets
If you've heard of quantum computing, you might be excited about the possibility of applying it to machine learning applications. I work at Springboard, and we recently launched a machine learning bootcamp that includes a job guarantee. We want to make sure our graduates are exposed to cutting-edge machine learning applications -- so we put together this article as part of our research into the intersection of quantum computing and machine learning. Let's start by examining the difference between quantum computing and classical computing. In classical computing, your data is stored in physical bits and it is binary and mutually exhaustive: a bit is either in a 0 state or in a 1 state and it cannot be both at the same time.
UAE launches 'world's first AI university' with 100% scholarships
After becoming the first nation to appoint a minister for artificial intelligence (AI) in 2017, the United Arab Emirates on Wednesday launched what it calls the'world's first AI university' in Abu Dhabi, with the opening of the Mohammad Bin Zayed University of Artifical Intelligence (MBZUAI). MBZUAI will be the world's first graduate university with a dedicated focus on AI. According to the university's website, MBZUAI is "a graduate-level, research-based academic institution that offers specialized degree programs for local and international students in the field of Artificial Intelligence." Located in Masdar City, Abu Dhabi, the university has begun accepting applications for masters and PhD programmes, with classes scheduled to begin in September 2020. Abu Dhabi's crown prince, Shaikh Mohammad Bin Zayed Al Nahyan, tweeted on the occasion, "Launching of the world's first graduate-level artificial intelligence university in Abu Dhabi echoes the UAE's pioneering spirit, and paves the way towards a new era of innovation and technological advancement that benefits the UAE and a world."
University of Artificial Intelligence launched in Abu Dhabi
Located in Masdar City with the latest state-of-the-art facilities and equipment, the university will offer both masters (two years) and PhD programmes (four years) for local and international graduate students across three main specialised fields โ machine learning, computer vision and natural language processing โ as the UAE looks to equip the next generation of students with the latest expertise in the field of AI.
Artificial Intelligence (AI)
Please note on June 30, 2020, this program will be retiring and no longer available on edX. If you are interested in earning the Professional Certificate you must be complete the program by June 30, 2020, in order to earn the certificate. Artificial Intelligence (AI) will define the next generation of software solutions. Human-like capabilities such as understanding natural language, speech, vision, and making inferences from knowledge will extend software beyond the app. The AI Professional Certificate program takes aspiring AI engineers from a basic introduction of AI to mastery of the skills needed to build deep learning models for AI solutions that exhibit human-like behavior and intelligence.