Deep Learning
Augment your batch: better training with larger batches
Hoffer, Elad, Ben-Nun, Tal, Hubara, Itay, Giladi, Niv, Hoefler, Torsten, Soudry, Daniel
Large-batch SGD is important for scaling training of deep neural networks. However, without fine-tuning hyperparameter schedules, the generalization of the model may be hampered. We propose to use batch augmentation: replicating instances of samples within the same batch with different data augmentations. Batch augmentation acts as a regularizer and an accelerator, increasing both generalization and performance scaling. We analyze the effect of batch augmentation on gradient variance and show that it empirically improves convergence for a wide variety of deep neural networks and datasets. Our results show that batch augmentation reduces the number of necessary SGD updates to achieve the same accuracy as the state-of-the-art. Overall, this simple yet effective method enables faster training and better generalization by allowing more computational resources to be used concurrently.
Fixup Initialization: Residual Learning Without Normalization
Zhang, Hongyi, Dauphin, Yann N., Ma, Tengyu
Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate convergence and improve generalization, though the reason for their effectiveness is still an active research topic. In this work, we challenge the commonly-held beliefs by showing that none of the perceived benefits is unique to normalization. Specifically, we propose fixed-update initialization (Fixup), an initialization motivated by solving the exploding and vanishing gradient problem at the beginning of training via properly rescaling a standard initialization. We find training residual networks with Fixup to be as stable as training with normalization -- even for networks with 10,000 layers. Furthermore, with proper regularization, Fixup enables residual networks without normalization to achieve state-of-the-art performance in image classification and machine translation.
The Oracle of DLphi
Alfke, Dominik, Baines, Weston, Blechschmidt, Jan, Sarmina, Mauricio J. del Razo, Drory, Amnon, Elbrächter, Dennis, Farchmin, Nando, Gambara, Matteo, Glas, Silke, Grohs, Philipp, Hinz, Peter, Kivaranovic, Danijel, Kümmerle, Christian, Kutyniok, Gitta, Lunz, Sebastian, Macdonald, Jan, Malthaner, Ryan, Naisat, Gregory, Neufeld, Ariel, Petersen, Philipp Christian, Reisenhofer, Rafael, Sheng, Jun-Da, Thesing, Laura, Trunschke, Philipp, von Lindheim, Johannes, Weber, David, Weber, Melanie
This paper takes aim at achieving nothing less than the impossible. To be more precise, we seek to predict labels of unknown data from entirely uncorrelated labelled training data. This will be accomplished by an application of an algorithm based on deep learning, as well as, by invoking one of the most fundamental concepts of set theory. Estimating the behaviour of a system in unknown situations is one of the central problems of humanity. Indeed, we are constantly trying to produce predictions for future events to be able to prepare ourselves.
Bayesian Learning of Neural Network Architectures
Dikov, Georgi, van der Smagt, Patrick, Bayer, Justin
In this paper we propose a Bayesian method for estimating architectural parameters of neural networks, namely layer size and network depth. We do this by learning concrete distributions over these parameters. Our results show that regular networks with a learnt structure can generalise better on small datasets, while fully stochastic networks can be more robust to parameter initialisation. The proposed method relies on standard neural variational learning and, unlike randomised architecture search, does not require a retraining of the model, thus keeping the computational overhead at minimum.
Neural Related Work Summarization with a Joint Context-driven Attention Mechanism
Wang, Yongzhen, Liu, Xiaozhong, Gao, Zheng
Conventional solutions to automatic related work summarization rely heavily on human-engineered features. In this paper, we develop a neural data-driven summarizer by leveraging the seq2seq paradigm, in which a joint context-driven attention mechanism is proposed to measure the contextual relevance within full texts and a heterogeneous bibliography graph simultaneously. Our motivation is to maintain the topic coherency between a related work section and its target document, where both the textual and graphic contexts play a big role in characterizing the relationship among scientific publications accurately. Experimental results on a large dataset show that our approach achieves a considerable improvement over a typical seq2seq summarizer and five classical summarization baselines.
On Learning Invariant Representation for Domain Adaptation
Zhao, Han, Combes, Remi Tachet des, Zhang, Kun, Gordon, Geoffrey J.
Due to the ability of deep neural nets to learn rich representations, recent advances in unsupervised domain adaptation have focused on learning domain-invariant features that achieve a small error on the source domain. The hope is that the learnt representation, together with the hypothesis learnt from the source domain, can generalize to the target domain. In this paper, we first construct a simple counterexample showing that, contrary to common belief, the above conditions are not sufficient to guarantee successful domain adaptation. In particular, the counterexample (Fig. 1) exhibits \emph{conditional shift}: the class-conditional distributions of input features change between source and target domains. To give a sufficient condition for domain adaptation, we propose a natural and interpretable generalization upper bound that explicitly takes into account the aforementioned shift. Moreover, we shed new light on the problem by proving an information-theoretic lower bound on the joint error of \emph{any} domain adaptation method that attempts to learn invariant representations. Our result characterizes a fundamental tradeoff between learning invariant representations and achieving small joint error on both domains when the marginal label distributions differ from source to target. Finally, we conduct experiments on real-world datasets that corroborate our theoretical findings. We believe these insights are helpful in guiding the future design of domain adaptation and representation learning algorithms.
Analysis: "The era of deep learning is coming to an end"
Many of the new developments in artificial intelligence that we hear about nowadays are actually just applications of machine learning techniques that have been hammered out for years. And as the research community's attention shifts from deep learning, it remains unclear what will take its place, according to MIT Tech. In the past, older types of artificial intelligence that didn't really take off when they were first developed later resurfaced and taken off with new life. For instance, scientists first developed machine learning decades ago, but it only became commonplace about a decade ago. MIT Tech didn't predict what will come next.
AI Helps Amputees Walk With a Robotic Knee Web Design & Website Hosting Services Creative Digital Agency Mean Web Host
A movie montage for modern artificial intelligence might show a computer playing millions of games of chess or Go against itself to learn how to win. Now, researchers are exploring how the reinforcement learning technique that helped DeepMind's AlphaZero conquer the chess and Go could tackle an even more complex task--training a robotic knee to help amputees walk smoothly. You must log in to article a comment. This site uses Akismet to reduce spam. Learn how your comment data is processed.
Google's StarCraft-playing AI is crushing pro gamers
In December, AlphaStar played as a Protoss and won five games against Dario Wünsch, a German player who goes by the gamer handle TLO and who also played as a Protoss (although it is not the group in which he specializes). A week later, the AI won five games again, this time against a tougher Protoss competitor: Grzegorz Komincz, a professional gamer from Poland who goes by the name MaNa. DeepMind announced the victories Thursday during a live stream on YouTube and Twitch. The researchers used a sort of tournament-style approach to train AlphaStar. First, they spent three days training a neural network -- a machine-learning algorithm modeled after the way neurons work in a brain -- on replays of human players' StarCraft II games. This neural network was used to create a number of computer-based competitors that played many, many rounds of the game against each other, learning from their experiences, over the course of two weeks.
Artificial Intelligence Now Simplifies Evolution
The sequencing of ancient Neanderthal and Denisovan fossils supported introgression events into anatomically modern humans (AMH) (out of Africa). However, recent studies also support the presence of gene flow from AMH into Neanderthals, thus suggesting a complex hominin evolution. All modern humans are genetically related to each other at a time depth of up to 300 thousand years ago and share a common African root. The migratory routes used by AMH after the African diaspora and aspects of the interbreeding between AMH and presently extinct hominins living at the time in Eurasia (here referred as Eurasian Extinct Hominins, EEH) are still under debate. Recently, for the first time, deep learning has been successfully used to explain human history, paving the way for this technology to be applied in other questions in medicine, genomics, and evolution.