Deep Learning
Robust Generalization of Quadratic Neural Networks via Function Identification
Xu, Kan, Bastani, Hamsa, Bastani, Osbert
A key challenge facing deep learning is that neural networks are often not robust to shifts in the underlying data distribution. We study this problem from the perspective of the statistical concept of parameter identification. Generalization bounds from learning theory often assume that the test distribution is close to the training distribution. In contrast, if we can identify the "true" parameters, then the model generalizes to arbitrary distribution shifts. However, neural networks are typically overparameterized, making parameter identification impossible. We show that for quadratic neural networks, we can identify the function represented by the model even though we cannot identify its parameters. Thus, we can obtain robust generalization bounds even in the overparameterized setting. We leverage this result to obtain new bounds for contextual bandits and transfer learning with quadratic neural networks. Overall, our results suggest that we can improve robustness of neural networks by designing models that can represent the true data generating process. In practice, the true data generating process is often very complex; thus, we study how our framework might connect to neural module networks, which are designed to break down complex tasks into compositions of simpler ones. We prove robust generalization bounds when individual neural modules are identifiable.
Entropic Issues in Likelihood-Based OOD Detection
Caterini, Anthony L., Loaiza-Ganem, Gabriel
Deep generative models trained by maximum likelihood remain very popular methods for reasoning about data probabilistically. However, it has been observed that they can assign higher likelihoods to out-of-distribution (OOD) data than in-distribution data, thus calling into question the meaning of these likelihood values. In this work we provide a novel perspective on this phenomenon, decomposing the average likelihood into a KL divergence term and an entropy term. We argue that the latter can explain the curious OOD behaviour mentioned above, suppressing likelihood values on datasets with higher entropy. Although our idea is simple, we have not seen it explored yet in the literature. This analysis provides further explanation for the success of OOD detection methods based on likelihood ratios, as the problematic entropy term cancels out in expectation. Finally, we discuss how this observation relates to recent success in OOD detection with manifold-supported models, for which the above decomposition does not hold.
Introducing Symmetries to Black Box Meta Reinforcement Learning
Kirsch, Louis, Flennerhag, Sebastian, van Hasselt, Hado, Friesen, Abram, Oh, Junhyuk, Chen, Yutian
Meta reinforcement learning (RL) attempts to discover new RL algorithms automatically from environment interaction. In so-called black-box approaches, the policy and the learning algorithm are jointly represented by a single neural network. These methods are very flexible, but they tend to underperform in terms of generalisation to new, unseen environments. In this paper, we explore the role of symmetries in meta-generalisation. We show that a recent successful meta RL approach that meta-learns an objective for backpropagation-based learning exhibits certain symmetries (specifically the reuse of the learning rule, and invariance to input and output permutations) that are not present in typical black-box meta RL systems. We hypothesise that these symmetries can play an important role in meta-generalisation. Building off recent work in black-box supervised meta learning, we develop a black-box meta RL system that exhibits these same symmetries. We show through careful experimentation that incorporating these symmetries can lead to algorithms with a greater ability to generalise to unseen action & observation spaces, tasks, and environments.
DeepMind tells Google it has no idea how to make AI less toxic
Did you know Neural is taking the stage this fall? Together with an amazing line-up of experts, we will explore the future of AI during TNW Conference 2021. Reducing the massive power consumption it takes to train deep learning models. These are among the loftiest outstanding problems in artificial intelligence. Whoever has the talent and budget to solve them will be handsomely rewarded with gobs and gobs of money.
The Imperative for Sustainable AI Systems
AI systems are compute-intensive: the AI lifecycle often requires long-running training jobs, hyperparameter searches, inference jobs, and other costly computations. They also require massive amounts of data that might be moved over the wire, and require specialized hardware to operate effectively, especially large-scale AI systems. All of these activities require electricity -- which has a carbon cost. There are also carbon emissions in ancillary needs like hardware and datacenter cooling [1]. Thus, AI systems have a massive carbon footprint[2]. This carbon footprint also has consequences in terms of social justice as we will explore in this article.
How to use Conv2d layers as fully connected layers.
In Deep learning, a convolutional neural network (CNN) is a class of deep NN, that are typically used to recognize patterns present in images but they are also used for spatial data analysis, computer vision, natural language processing, signal processing, and various other purposes. The focus of this article is to demonstrate how do we use Convolutional layers as Fully connected layers or How to convert fully connected layers to Conv layers. I have tried to find relevant material on the internet, but only could find one research paper on this subject, so we will use that to build up on the theoretical part, then we will dive deep into the coding part. What is a Conv layer in Deep networks?. It's nothing but a convolution operation that is done on an image/input. Such an operation requires a kernel (matrix) with values to move on the image (which itself is converted into a matrix) with a set value of stride or steps performing an elementwise multiplication.
Falsehoods more likely with large language models
The Transform Technology Summits start October 13th with Low-Code/No Code: Enabling Enterprise Agility. The use of AI language models to generate text for business applications is gaining steam. Large companies are deploying their own systems, while others are leveraging models like OpenAI's GPT-3 via APIs. According to OpenAI, GPT-3 is now being used in over 300 apps by thousands of developers, producing an average of more than 4.5 billion novel words per day. But while recent language models are impressively fluent, they have a tendency to write falsehoods ranging from factual inaccuracies to potentially harmful disinformation.
Improved algorithms may be more important for AI performance than faster hardware
The Transform Technology Summits start October 13th with Low-Code/No Code: Enabling Enterprise Agility. When it comes to AI, algorithmic innovations are substantially more important than hardware -- at least where the problems involve billions to trillions of data points. That's the conclusion of a team of scientists at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL), who conducted what they claim is the first study on how fast algorithms are improving across a broad range of examples. Algorithms tell software how to make sense of text, visual, and audio data so that they can, in turn, draw inferences from it. For example, OpenAI's GPT-3 was trained on webpages, ebooks, and other documents to learn how to write papers in a humanlike way.
A review of deep learning methods for MRI reconstruction
Following the success of deep learning in a wide range of applications, neural network-based machine-learning techniques have received significant interest for accelerating magnetic resonance imaging (MRI) acquisition and reconstruction strategies. A number of ideas inspired by deep learning techniques for computer vision and image processing have been successfully applied to nonlinear image reconstruction in the spirit of compressed sensing for accelerated MRI. Given the rapidly growing nature of the field, it is imperative to consolidate and summarize the large number of deep learning methods that have been reported in the literature, to obtain a better understanding of the field in general. This article provides an overview of the recent developments in neural-network based approaches that have been proposed specifically for improving parallel imaging. A general background and introduction to parallel MRI is also given from a classical view of k-space based reconstruction methods.
Best Laptops for Deep Learning, Machine Learning, and Data Science for 2021
Machine learners, deep learning practitioners, and data scientists are continually looking for the edge on their performance-oriented devices. That's why we looked at over 2,000 laptops to bring you what we consider the best laptops for your projects on machine learning, deep learning, and data science. We will continuously update this resource with powerful and more performant laptops for every budget as technology continues to evolve to bring you the best suggestions for your machine learning, data science, and deep learning projects and adventures. Our mailbox is full of emails from AI enthusiasts asking us for the best laptops for AI projects. That's why we decided to make this list.