Deep Learning
Local Competition and Uncertainty for Adversarial Robustness in Deep Learning
Alexos, Antonios, Panousis, Konstantinos P., Chatzis, Sotirios
This work attempts to address adversarial robustness of deep networks by means of novel learning arguments. Specifically, inspired from results in neuroscience, we propose a local competition principle as a means of adversarially-robust deep learning. We argue that novel local winner-takes-all (LWTA) nonlinearities, combined with posterior sampling schemes, can greatly improve the adversarial robustness of traditional deep networks against difficult adversarial attack schemes. We combine these LWTA arguments with tools from the field of Bayesian non-parametrics, specifically the stick-breaking construction of the Indian Buffet Process, to flexibly account for the inherent uncertainty in data-driven modeling. As we experimentally show, the new proposed model achieves high robustness to adversarial perturbations on MNIST and CIFAR10 datasets. Our model achieves state-of-the-art results in powerful white-box attacks, while at the same time retaining its benign accuracy to a high degree. Equally importantly, our approach achieves this result while requiring far less trainable model parameters than the existing state-of-the-art.
Riemannian Continuous Normalizing Flows
Mathieu, Emile, Nickel, Maximilian
Normalizing flows have shown great promise for modelling flexible probability distributions in a computationally tractable way. However, whilst data is often naturally described on Riemannian manifolds such as spheres, torii, and hyperbolic spaces, most normalizing flows implicitly assume a flat geometry, making them either misspecified or ill-suited in these situations. To overcome this problem, we introduce Riemannian continuous normalizing flows, a model which admits the parametrization of flexible probability measures on smooth manifolds by defining flows as the solution to ordinary differential equations. We show that this approach can lead to substantial improvements on both synthetic and real-world data when compared to standard flows or previously introduced projected flows.
Shapeshifter Networks: Cross-layer Parameter Sharing for Scalable and Effective Deep Learning
Plummer, Bryan A., Dryden, Nikoli, Frost, Julius, Hoefler, Torsten, Saenko, Kate
We present Shapeshifter Networks (SSNs), a flexible neural network framework that improves performance and reduces memory requirements on a diverse set of scenarios over standard neural networks. Our approach is based on the observation that many neural networks are severely overparameterized, resulting in significant waste in computational resources as well as being susceptible to overfitting. SSNs address this by learning where and how to share parameters between layers in a neural network while avoiding degenerate solutions that result in underfitting. Specifically, we automatically construct parameter groups that identify where parameter sharing is most beneficial. Then, we map each group's weights to construct layers with learned combinations of candidates from a shared parameter pool. SSNs can share parameters across layers even when they have different sizes, perform different operations, and/or operate on features from different modalities. We evaluate our approach on a diverse set of tasks, including image classification, bidirectional image-sentence retrieval, and phrase grounding, creating high performing models even when using as little as 1% of the parameters. We also apply SSNs to knowledge distillation, where we obtain state-of-the-art results when combined with traditional distillation methods.
SXL: Spatially explicit learning of geographic processes with auxiliary tasks
Klemmer, Konstantin, Neill, Daniel B.
From earth system sciences to climate modeling and ecology, many of the greatest empirical modeling challenges are geographic in nature. As these processes are characterized by spatial dynamics, we can exploit their autoregressive nature to inform learning algorithms. We introduce SXL, a method for learning with geospatial data using explicitly spatial auxiliary tasks. We embed the local Moran's I, a well-established measure of local spatial autocorrelation, into the training process, "nudging" the model to learn the direction and magnitude of local autoregressive effects in parallel with the primary task. Further, we propose an expansion of Moran's I to multiple resolutions to capture effects at different spatial granularities and over varying distance scales. We show the superiority of this method for training deep neural networks using experiments with real-world geospatial data in both generative and predictive modeling tasks. Our approach can be used with arbitrary network architectures and, in our experiments, consistently improves their performance. We also outperform appropriate, domain-specific interpolation benchmarks. Our work highlights how integrating the geographic information sciences and spatial statistics into machine learning models can address the specific challenges of spatial data.
An Investigation of the Weight Space for Version Control of Neural Networks
Schรผrholt, Konstantin, Borth, Damian
Deployed Deep Neural Networks (DNNs) are often trained further to improve in performance. This complicates the tracking of DNN model versions and the synchronization between already deployed models and upstream updates of the same architecture. Software Version Control cannot be applied straight-forwardly to DNNs due to the different nature of software and DNN models. In this paper we investigate if the weight space of DNN models contains a structure, which can be used for the identification of individual DNN models. Our results show that DNN models evolve on unique, smooth trajectories in weight space which we can exploit as feature for DNN version control.
Sparse Bottleneck Networks for Exploratory Analysis and Visualization of Neural Patch-seq Data
Bernaerts, Yves, Berens, Philipp, Kobak, Dmitry
In recent years, increasingly large datasets with two different sets of features measured for each sample have become prevalent in many areas of biology. For example, a recently developed method called Patch-seq provides single-cell RNA sequencing data together with electrophysiological measurements of the same neurons. However, the efficient and interpretable analysis of such paired data has remained a challenge. As a tool for exploration and visualization of Patch-seq data, we introduce neural networks with a two-dimensional bottleneck, trained to predict electrophysiological measurements from gene expression. To make the model biologically interpretable and perform gene selection, we enforce sparsity by using a group lasso penalty, followed by pruning of the input units and subsequent fine-tuning. We applied this method to a recent dataset with $>$1000 neurons from mouse motor cortex and found that the resulting bottleneck model had the same predictive performance as a full-rank linear model with much higher latent dimensionality. Exploring the two-dimensional latent space in terms of neural types showed that the nonlinear bottleneck approach led to much better visualizations and higher biological interpretability.
DREAM: Deep Regret minimization with Advantage baselines and Model-free learning
Steinberger, Eric, Lerer, Adam, Brown, Noam
We introduce DREAM, a deep reinforcement learning algorithm that finds optimal strategies in imperfect-information games with multiple agents. Formally, DREAM converges to a Nash Equilibrium in two-player zero-sum games and to an extensive-form coarse correlated equilibrium in all other games. Our primary innovation is an effective algorithm that, in contrast to other regret-based deep learning algorithms, does not require access to a perfect simulator of the game to achieve good performance. We show that DREAM empirically achieves state-of-the-art performance among model-free algorithms in popular benchmark games, and is even competitive with algorithms that do use a perfect simulator.
Overcoming Classifier Imbalance for Long-tail Object Detection with Balanced Group Softmax
Li, Yu, Wang, Tao, Kang, Bingyi, Tang, Sheng, Wang, Chunfeng, Li, Jintao, Feng, Jiashi
Solving long-tail large vocabulary object detection with deep learning based models is a challenging and demanding task, which is however under-explored.In this work, we provide the first systematic analysis on the underperformance of state-of-the-art models in front of long-tail distribution. We find existing detection methods are unable to model few-shot classes when the dataset is extremely skewed, which can result in classifier imbalance in terms of parameter magnitude. Directly adapting long-tail classification models to detection frameworks can not solve this problem due to the intrinsic difference between detection and classification.In this work, we propose a novel balanced group softmax (BAGS) module for balancing the classifiers within the detection frameworks through group-wise training. It implicitly modulates the training process for the head and tail classes and ensures they are both sufficiently trained, without requiring any extra sampling for the instances from the tail classes.Extensive experiments on the very recent long-tail large vocabulary object recognition benchmark LVIS show that our proposed BAGS significantly improves the performance of detectors with various backbones and frameworks on both object detection and instance segmentation. It beats all state-of-the-art methods transferred from long-tail image classification and establishes new state-of-the-art.Code is available at https://github.com/FishYuLi/BalancedGroupSoftmax.
Neural Architecture Optimization with Graph VAE
Li, Jian, Liu, Yong, Liu, Jiankun, Wang, Weiping
Due to their high computational efficiency on a continuous space, gradient optimization methods have shown great potential in the neural architecture search (NAS) domain. The mapping of network representation from the discrete space to a latent space is the key to discovering novel architectures, however, existing gradient-based methods fail to fully characterize the networks. In this paper, we propose an efficient NAS approach to optimize network architectures in a continuous space, where the latent space is built upon variational autoencoder (VAE) and graph neural networks (GNN). The framework jointly learns four components: the encoder, the performance predictor, the complexity predictor and the decoder in an end-to-end manner. The encoder and the decoder belong to a graph VAE, mapping architectures between continuous representations and network architectures. The predictors are two regression models, fitting the performance and computational complexity, respectively. Those predictors ensure the discovered architectures characterize both excellent performance and high computational efficiency. Extensive experiments demonstrate our framework not only generates appropriate continuous representations but also discovers powerful neural architectures.
GAT-GMM: Generative Adversarial Training for Gaussian Mixture Models
Farnia, Farzan, Wang, William, Das, Subhro, Jadbabaie, Ali
Learning the distribution of observed data is a basic task in unsupervised learning which has been studied for decades. The recently-introduced concept of Generative Adversarial Networks (GANs) [1] has demonstrated great success in various distribution learning tasks. Unlike the traditional maximum-likelihood-based approaches, GANs learn the distribution of observed data through a zero-sum game between two machine players, a generator G mimicking the true distribution of data and a discriminator D distinguishing the generator's produced samples from real data points. This zero-sum game is typically formulated through a minimax optimization problem where G and D optimize a minimax objective quantifying how dissimilar G's generated samples and real training samples are. In GAN minimax optimization problems, the generator and discriminator functions are commonly chosen as two deep neural networks (DNNs). Leveraging the expressive power of DNNs, GANs have achieved state-of-the-art performance in learning complex distributions of image data [2, 3, 4]. This success, however, is achieved at the cost of their notoriously difficult training procedure which has introduced several challenges to the machine learning community. Addressing these challenges requires a deeper theoretical understanding of GANs, including their approximation, generalization, and optimization properties. Specifically, GANs have been frequently observed to fail in learning multi-modal distributions [5].