Goto

Collaborating Authors

 Deep Learning


Data Visualization

#artificialintelligence

When designing and evaluating a new algorithm, one of the first steps is exploratory data analysis (EDA). The point is to find the most efficient learning approach for a given problem. For the human researcher to understand what's working and what's not, the model's results are often displayed graphically. Since these datasets cover many variables and are so "high dimensional," several new data visualization techniques have been developed specifically for deep learning systems.


Is AI an agent of big tech hegemony or multi-disciplinary research and innovation?

#artificialintelligence

A recent New York Times article fretting about the soaring costs of developing and training leading-edge deep learning models and my admittedly provocative Tweet questioning the premise and motives of the article's sources led to the type of online banter that indicates a nuanced question ill-suited for pithy Twitter responses. Fears of AI creating a chasm between haves and have-nots are common, however the topic of AI-fueled inequality typically centers on its economic effects, namely that the growing substitution of manual labor with algorithmic automation serves to further polarize income distributions as the knowledge class controlling and using the algorithms get richer while the working class being displaced by machines suffers. Many new technologies -- those we call'automation technologies' -- do not increase laborรญs productivity, but are explicitly aimed at replacing it by substituting cheaper capital (machines) in a range of tasks performed by humans. As a result, automation technologies always reduce the laborรญs share in value added (because they increase productivity by more than wages and employment). They may also reduce overall labor demand because they displace workers from the tasks they were previously performing.


The Economist's essay contest featured an AI submission. Here's what the judges thought.

#artificialintelligence

Earlier this summer, the Economist announced a competition for young people. They asked contestants to answer this question: "What fundamental economic and political change, if any, is needed for an effective response to climate change?" More than 2,400 people responded, from over 110 countries. And the Economist slipped one essay into the stack of submissions that their judges would review: an essay written by an artificial intelligence. The AI in question was GPT-2, a language-generating system developed by San Francisco AI lab OpenAI and announced this spring.


Operational Calibration: Debugging Confidence Errors for DNNs in the Field

arXiv.org Machine Learning

Trained DNN models are increasingly adopted as integral parts of software systems. However, they are often over-confident, especially in practical operation domains where slight divergence from their training data almost always exists. To minimize the loss due to inaccurate confidence, operational calibration, i.e., calibrating the confidence function of a DNN classifier against its operation domain, becomes a necessary debugging step in the engineering of the whole system. Operational calibration is difficult considering the limited budget of labeling operation data and the weak interpretability of DNN models. We propose a Bayesian approach to operational calibration that gradually corrects the confidence given by the model under calibration with a small number of labeled operational data deliberately selected from a larger set of unlabeled operational data. Exploiting the locality of the learned representation of the DNN model and modeling the calibration as Gaussian Process Regression, the approach achieves impressive efficacy and efficiency. Comprehensive experiments with various practical data sets and DNN models show that it significantly outperformed alternative methods, and in some difficult tasks it eliminated about 71% to 97% high-confidence errors with only about 10% of the minimal amount of labeled operation data needed for practical learning techniques to barely work.


Neural Multisensory Scene Inference

arXiv.org Machine Learning

For embodied agents to infer representations of the underlying 3D physical world they inhabit, they should efficiently combine multisensory cues from numerous trials, e.g., by looking at and touching objects. Despite its importance, multisensory 3D scene representation learning has received less attention compared to the unimodal setting. In this paper, we propose the Generative Multisensory Network (GMN) for learning latent representations of 3D scenes which are partially observable through multiple sensory modalities. We also introduce a novel method, called the Amortized Product-of-Experts, to improve the computational efficiency and the robustness to unseen combinations of modalities at test time. Experimental results demonstrate that the proposed model can efficiently infer robust modality-invariant 3D-scene representations from arbitrary combinations of modalities and perform accurate cross-modal generation. To perform this exploration, we also develop the Multisensory Embodied 3D-Scene Environment (MESE).


Minimum "Norm" Neural Networks are Splines

arXiv.org Machine Learning

We develop a general framework based on splines to understand the interpolation properties of overparameterized neural networks. We prove that minimum "norm" two-layer neural networks (with appropriately chosen activation functions) that interpolate scattered data are minimal knot splines. Our results follow from understanding key relationships between notions of neural network "norms", linear operators, and continuous-domain linear inverse problems.


Characterizing Membership Privacy in Stochastic Gradient Langevin Dynamics

arXiv.org Machine Learning

Bayesian deep learning is recently regarded as an intrinsic way to characterize the weight uncertainty of deep neural networks~(DNNs). Stochastic Gradient Langevin Dynamics~(SGLD) is an effective method to enable Bayesian deep learning on large-scale datasets. Previous theoretical studies have shown various appealing properties of SGLD, ranging from the convergence properties to the generalization bounds. In this paper, we study the properties of SGLD from a novel perspective of membership privacy protection (i.e., preventing the membership attack). The membership attack, which aims to determine whether a specific sample is used for training a given DNN model, has emerged as a common threat against deep learning algorithms. To this end, we build a theoretical framework to analyze the information leakage (w.r.t. the training dataset) of a model trained using SGLD. Based on this framework, we demonstrate that SGLD can prevent the information leakage of the training dataset to a certain extent. Moreover, our theoretical analysis can be naturally extended to other types of Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods. Empirical results on different datasets and models verify our theoretical findings and suggest that the SGLD algorithm can not only reduce the information leakage but also improve the generalization ability of the DNN models in real-world applications.


Pay Attention: Leveraging Sequence Models to Predict the Useful Life of Batteries

arXiv.org Machine Learning

We use data on 124 batteries released by Stanford University to first try to solve the binary classification problem of determining if a battery is "good" or "bad" given only the first 5 cycles of data (i.e., will it last longer than a certain threshold of cycles), as well as the prediction problem of determining the exact number of cycles a battery will last given the first 100 cycles of data. We approach the problem from a purely data-driven standpoint, hoping to use deep learning to learn the patterns in the sequences of data that the Stanford team engineered by hand. For both problems, we used a similar deep network design, that included an optional 1-D convolution, LSTMs, an optional Attention layer, followed by fully connected layers to produce our output. For the classification task, we were able to achieve very competitive results, with validation accuracies above 90%, and a test accuracy of 95%, compared to the 97.5% test accuracy of the current leading model. For the prediction task, we were also able to achieve competitive results, with a test MAPE error of 12.5% as compared with a 9.1% MAPE error achieved by the current leading model (Severson et al. 2019).


Stein Bridging: Enabling Mutual Reinforcement between Explicit and Implicit Generative Models

arXiv.org Machine Learning

Deep generative models are generally categorized into explicit models and implicit models. The former defines an explicit density form, whose normalizing constant is often unknown; while the latter, including generative adversarial networks (GANs), generates samples without explicitly defining a density function. In spite of substantial recent advances demonstrating the power of the two classes of generative models in many applications, both of them, when used alone, suffer from respective limitations and drawbacks. To mitigate these issues, we propose Stein Bridging, a novel joint training framework that connects an explicit density estimator and an implicit sample generator with Stein discrepancy. We show that the Stein Bridge induces new regularization schemes for both explicit and implicit models. Convergence analysis and extensive experiments demonstrate that the Stein Bridging i) improves the stability and sample quality of the GAN training, and ii) facilitates the density estimator to seek more modes in data and alleviate the mode-collapse issue. Additionally, we discuss several applications of Stein Bridging and useful tricks in practical implementation used in our experiments.


Making sense of sensory input

arXiv.org Artificial Intelligence

This paper attempts to answer a central question in unsupervised learning: what does it mean to "make sense" of a sensory sequence? In our formalization, making sense involves constructing a symbolic causal theory that explains the sensory sequence and satisfies a set of unity conditions. This model was inspired by Kant's discussion of the synthetic unity of apperception in the Critique of Pure Reason. On our account, making sense of sensory input is a type of program synthesis, but it is unsupervised program synthesis. Our second contribution is a computer implementation, the Apperception Engine, that was designed to satisfy the above requirements. Our system is able to produce interpretable human-readable causal theories from very small amounts of data, because of the strong inductive bias provided by the Kantian unity constraints. A causal theory produced by our system is able to predict future sensor readings, as well as retrodict earlier readings, and "impute" (fill in the blanks of) missing sensory readings, in any combination. We tested the engine in a diverse variety of domains, including cellular automata, rhythms and simple nursery tunes, multi-modal binding problems, occlusion tasks, and sequence induction IQ tests. In each domain, we test our engine's ability to predict future sensor values, retrodict earlier sensor values, and impute missing sensory data. The Apperception Engine performs well in all these domains, significantly out-performing neural net baselines. We note in particular that in the sequence induction IQ tasks, our system achieved human-level performance. This is notable because our system is not a bespoke system designed specifically to solve IQ tasks, but a general purpose apperception system that was designed to make sense of any sensory sequence.