Goto

Collaborating Authors

 Deep Learning


An Investigation of the Weight Space for Version Control of Neural Networks

arXiv.org Machine Learning

Deployed Deep Neural Networks (DNNs) are often trained further to improve in performance. This complicates the tracking of DNN model versions and the synchronization between already deployed models and upstream updates of the same architecture. Software Version Control cannot be applied straight-forwardly to DNNs due to the different nature of software and DNN models. In this paper we investigate if the weight space of DNN models contains a structure, which can be used for the identification of individual DNN models. Our results show that DNN models evolve on unique, smooth trajectories in weight space which we can exploit as feature for DNN version control.


Sparse Bottleneck Networks for Exploratory Analysis and Visualization of Neural Patch-seq Data

arXiv.org Machine Learning

In recent years, increasingly large datasets with two different sets of features measured for each sample have become prevalent in many areas of biology. For example, a recently developed method called Patch-seq provides single-cell RNA sequencing data together with electrophysiological measurements of the same neurons. However, the efficient and interpretable analysis of such paired data has remained a challenge. As a tool for exploration and visualization of Patch-seq data, we introduce neural networks with a two-dimensional bottleneck, trained to predict electrophysiological measurements from gene expression. To make the model biologically interpretable and perform gene selection, we enforce sparsity by using a group lasso penalty, followed by pruning of the input units and subsequent fine-tuning. We applied this method to a recent dataset with $>$1000 neurons from mouse motor cortex and found that the resulting bottleneck model had the same predictive performance as a full-rank linear model with much higher latent dimensionality. Exploring the two-dimensional latent space in terms of neural types showed that the nonlinear bottleneck approach led to much better visualizations and higher biological interpretability.


DREAM: Deep Regret minimization with Advantage baselines and Model-free learning

arXiv.org Machine Learning

We introduce DREAM, a deep reinforcement learning algorithm that finds optimal strategies in imperfect-information games with multiple agents. Formally, DREAM converges to a Nash Equilibrium in two-player zero-sum games and to an extensive-form coarse correlated equilibrium in all other games. Our primary innovation is an effective algorithm that, in contrast to other regret-based deep learning algorithms, does not require access to a perfect simulator of the game to achieve good performance. We show that DREAM empirically achieves state-of-the-art performance among model-free algorithms in popular benchmark games, and is even competitive with algorithms that do use a perfect simulator.


Overcoming Classifier Imbalance for Long-tail Object Detection with Balanced Group Softmax

arXiv.org Machine Learning

Solving long-tail large vocabulary object detection with deep learning based models is a challenging and demanding task, which is however under-explored.In this work, we provide the first systematic analysis on the underperformance of state-of-the-art models in front of long-tail distribution. We find existing detection methods are unable to model few-shot classes when the dataset is extremely skewed, which can result in classifier imbalance in terms of parameter magnitude. Directly adapting long-tail classification models to detection frameworks can not solve this problem due to the intrinsic difference between detection and classification.In this work, we propose a novel balanced group softmax (BAGS) module for balancing the classifiers within the detection frameworks through group-wise training. It implicitly modulates the training process for the head and tail classes and ensures they are both sufficiently trained, without requiring any extra sampling for the instances from the tail classes.Extensive experiments on the very recent long-tail large vocabulary object recognition benchmark LVIS show that our proposed BAGS significantly improves the performance of detectors with various backbones and frameworks on both object detection and instance segmentation. It beats all state-of-the-art methods transferred from long-tail image classification and establishes new state-of-the-art.Code is available at https://github.com/FishYuLi/BalancedGroupSoftmax.


Neural Architecture Optimization with Graph VAE

arXiv.org Machine Learning

Due to their high computational efficiency on a continuous space, gradient optimization methods have shown great potential in the neural architecture search (NAS) domain. The mapping of network representation from the discrete space to a latent space is the key to discovering novel architectures, however, existing gradient-based methods fail to fully characterize the networks. In this paper, we propose an efficient NAS approach to optimize network architectures in a continuous space, where the latent space is built upon variational autoencoder (VAE) and graph neural networks (GNN). The framework jointly learns four components: the encoder, the performance predictor, the complexity predictor and the decoder in an end-to-end manner. The encoder and the decoder belong to a graph VAE, mapping architectures between continuous representations and network architectures. The predictors are two regression models, fitting the performance and computational complexity, respectively. Those predictors ensure the discovered architectures characterize both excellent performance and high computational efficiency. Extensive experiments demonstrate our framework not only generates appropriate continuous representations but also discovers powerful neural architectures.


GAT-GMM: Generative Adversarial Training for Gaussian Mixture Models

arXiv.org Machine Learning

Learning the distribution of observed data is a basic task in unsupervised learning which has been studied for decades. The recently-introduced concept of Generative Adversarial Networks (GANs) [1] has demonstrated great success in various distribution learning tasks. Unlike the traditional maximum-likelihood-based approaches, GANs learn the distribution of observed data through a zero-sum game between two machine players, a generator G mimicking the true distribution of data and a discriminator D distinguishing the generator's produced samples from real data points. This zero-sum game is typically formulated through a minimax optimization problem where G and D optimize a minimax objective quantifying how dissimilar G's generated samples and real training samples are. In GAN minimax optimization problems, the generator and discriminator functions are commonly chosen as two deep neural networks (DNNs). Leveraging the expressive power of DNNs, GANs have achieved state-of-the-art performance in learning complex distributions of image data [2, 3, 4]. This success, however, is achieved at the cost of their notoriously difficult training procedure which has introduced several challenges to the machine learning community. Addressing these challenges requires a deeper theoretical understanding of GANs, including their approximation, generalization, and optimization properties. Specifically, GANs have been frequently observed to fail in learning multi-modal distributions [5].


Robust compressed sensing of generative models

arXiv.org Machine Learning

The goal of compressed sensing is to estimate a high dimensional vector from an underdetermined system of noisy linear equations. In analogy to classical compressed sensing, here we assume a generative model as a prior, that is, we assume the vector is represented by a deep generative model $G: \mathbb{R}^k \rightarrow \mathbb{R}^n$. Classical recovery approaches such as empirical risk minimization (ERM) are guaranteed to succeed when the measurement matrix is sub-Gaussian. However, when the measurement matrix and measurements are heavy-tailed or have outliers, recovery may fail dramatically. In this paper we propose an algorithm inspired by the Median-of-Means (MOM). Our algorithm guarantees recovery for heavy-tailed data, even in the presence of outliers. Theoretically, our results show our novel MOM-based algorithm enjoys the same sample complexity guarantees as ERM under sub-Gaussian assumptions. Our experiments validate both aspects of our claims: other algorithms are indeed fragile and fail under heavy-tailed and/or corrupted data, while our approach exhibits the predicted robustness.


A new measure for overfitting and its implications for backdooring of deep learning

arXiv.org Machine Learning

Overfitting describes the phenomenon that a machine learning model fits the given data instead of learning the underlying distribution. Existing approaches are computationally expensive, require large amounts of labeled data, consider overfitting global phenomenon, and often compute a single measurement. Instead, we propose a local measurement around a small number of unlabeled test points to obtain features of overfitting. Our extensive evaluation shows that the measure can reflect the model's different fit of training and test data, identify changes of the fit during training, and even suggest different fit among classes. We further apply our method to verify if backdoors rely on overfitting, a common claim in security of deep learning. Instead, we find that backdoors rely on underfitting. Our findings also provide evidence that even unbackdoored neural networks contain patterns similar to backdoors that are reliably classified as one class.


CERT: Contrastive Self-supervised Learning for Language Understanding

arXiv.org Machine Learning

Pretrained language models such as BERT, GPT have shown great effectiveness in language understanding. The auxiliary predictive tasks in existing pretraining approaches are mostly defined on tokens, thus may not be able to capture sentence-level semantics very well. To address this issue, we propose CERT: Contrastive self-supervised Encoder Representations from Transformers, which pretrains language representation models using contrastive self-supervised learning at the sentence level. CERT creates augmentations of original sentences using back-translation. Then it finetunes a pretrained language encoder (e.g., BERT) by predicting whether two augmented sentences originate from the same sentence. CERT is simple to use and can be flexibly plugged into any pretraining-finetuning NLP pipeline. We evaluate CERT on 11 natural language understanding tasks in the GLUE benchmark where CERT outperforms BERT on 7 tasks, achieves the same performance as BERT on 2 tasks, and performs worse than BERT on 2 tasks. On the averaged score of the 11 tasks, CERT outperforms BERT. The data and code are available at https://github.com/UCSD-AI4H/CERT


Crop Disease Detection Using Machine Learning and Computer Vision - KDnuggets

#artificialintelligence

International Conference on Learning Representations (ICLR) and Consultative Group on International Agricultural Research (CGIAR) jointly conducted a challenge where over 800 data scientists globally competed to detect diseases in crops based on close shot pictures. The objective of this challenge is to build a machine learning algorithm to correctly classify if a plant is healthy, has stem rust, or has leaf rust. Wheat rust is a devastating plant disease affecting many crops, reducing yields and affecting the livelihoods of farmers and decreasing food security across Africa. The disease is difficult to monitor at a large scale, making it difficult to control and eradicate. An accurate image recognition model that can detect wheat rust from any image will enable a crowd-sourced approach to monitor crops. The imagery data came from a variety of sources.