Deep Learning
Connecting First and Second Order Recurrent Networks with Deterministic Finite Automata
Wang, Qinglong, Zhang, Kaixuan, Liu, Xue, Giles, C. Lee
We propose an approach that connects recurrent networks with different orders of hidden interaction with regular grammars of different levels of complexity. We argue that the correspondence between recurrent networks and formal computational models gives understanding to the analysis of the complicated behaviors of recurrent networks. We introduce an entropy value that categorizes all regular grammars into three classes with different levels of complexity, and show that several existing recurrent networks match grammars from either all or partial classes. As such, the differences between regular grammars reveal the different properties of these models. We also provide a unification of all investigated recurrent networks. Our evaluation shows that the unified recurrent network has improved performance in learning grammars, and demonstrates comparable performance on a real-world dataset with more complicated models.
Robust Design of Deep Neural Networks against Adversarial Attacks based on Lyapunov Theory
Rahnama, Arash, Nguyen, Andre T., Raff, Edward
Deep neural networks (DNNs) are vulnerable to subtle adversarial perturbations applied to the input. These adversarial perturbations, though imperceptible, can easily mislead the DNN. In this work, we take a control theoretic approach to the problem of robustness in DNNs. We treat each individual layer of the DNN as a nonlinear dynamical system and use Lyapunov theory to prove stability and robustness locally. We then proceed to prove stability and robustness globally for the entire DNN. We develop empirically tight bounds on the response of the output layer, or any hidden layer, to adversarial perturbations added to the input, or the input of hidden layers. Recent works have proposed spectral norm regularization as a solution for improving robustness against l2 adversarial attacks. Our results give new insights into how spectral norm regularization can mitigate the adversarial effects. Finally, we evaluate the power of our approach on a variety of data sets and network architectures and against some of the well-known adversarial attacks.
Model-Augmented Nearest-Neighbor Estimation of Conditional Mutual Information for Feature Selection
Yang, Alan, Ghassami, AmirEmad, Raginsky, Maxim, Kiyavash, Negar, Rosenbaum, Elyse
Markov blanket feature selection, while theoretically optimal, generally is challenging to implement. This is due to the shortcomings of existing approaches to conditional independence (CI) testing, which tend to struggle either with the curse of dimensionality or computational complexity. We propose a novel two-step approach which facilitates Markov blanket feature selection in high dimensions. First, neural networks are used to map features to low-dimensional representations. In the second step, CI testing is performed by applying the k-NN conditional mutual information estimator to the learned feature maps. The mappings are designed to ensure that mapped samples both preserve information and share similar information about the target variable if and only if they are close in Euclidean distance. We show that these properties boost the performance of the k-NN estimator in the second step. The performance of the proposed method is evaluated on synthetic, as well as real data pertaining to datacenter hard disk drive failures.
Adaptive Probabilistic Vehicle Trajectory Prediction Through Physically Feasible Bayesian Recurrent Neural Network
Tang, Chen, Chen, Jianyu, Tomizuka, Masayoshi
Probabilistic vehicle trajectory prediction is essential for robust safety of autonomous driving. Current methods for long-term trajectory prediction cannot guarantee the physical feasibility of predicted distribution. Moreover, their models cannot adapt to the driving policy of the predicted target human driver. In this work, we propose to overcome these two shortcomings by a Bayesian recurrent neural network model consisting of Bayesian-neural-network-based policy model and known physical model of the scenario. Bayesian neural network can ensemble complicated output distribution, enabling rich family of trajectory distribution. The embedded physical model ensures feasibility of the distribution. Moreover, the adopted gradient-based training method allows direct optimization for better performance in long prediction horizon. Furthermore, a particle-filter-based parameter adaptation algorithm is designed to adapt the policy Bayesian neural network to the predicted target online. Effectiveness of the proposed methods is verified with a toy example with multi-modal stochastic feedback gain and naturalistic car following data.
Evaluating Combinatorial Generalization in Variational Autoencoders
Bozkurt, Alican, Esmaeili, Babak, Brooks, Dana H., Dy, Jennifer G., van de Meent, Jan-Willem
We evaluate the ability of variational autoencoders to generalize to unseen examples in domains with a large combinatorial space of feature values. Our experiments systematically evaluate the effect of network width, depth, regularization, and the typical distance between the training and test examples. Increasing network capacity benefits generalization in easy problems, where test-set examples are similar to training examples. In more difficult problems, increasing capacity deteriorates generalization when optimizing the standard VAE objective, but once again improves generalization when we decrease the KL regularization. Our results establish that interplay between model capacity and KL regularization is not clear cut; we need to take the typical distance between train and test examples into account when evaluating generalization.
Neural Contextual Bandits with Upper Confidence Bound-Based Exploration
Zhou, Dongruo, Li, Lihong, Gu, Quanquan
We study the stochastic contextual bandit problem, where the reward is generated from an unknown bounded function with additive noise. We propose the NeuralUCB algorithm, which leverages the representation power of deep neural networks and uses a neural network-based random feature mapping to construct an upper confidence bound (UCB) of reward for efficient exploration. We prove that, under mild assumptions, NeuralUCB achieves $\tilde O(\sqrt{T})$ regret, where $T$ is the number of rounds. To the best of our knowledge, our algorithm is the first neural network-based contextual bandit algorithm with near-optimal regret guarantee. Preliminary experiment results on synthetic data corroborate our theory, and shed light on potential applications of our algorithm to real-world problems.
Structural Pruning in Deep Neural Networks: A Small-World Approach
Krishnan, Gokul, Du, Xiaocong, Cao, Yu
--Deep Neural Networks (DNNs) are usually over-parameterized, causing excessive memory and interconnection cost on the hardware platform. Existing pruning approaches remove secondary parameters at the end of training to reduce the model size; but without exploiting the intrinsic network property, they still require the full interconnection to prepare the network. Inspired by the observation that brain networks follow the Small-World model, we propose a novel structural pruning scheme, which includes (1) hierarchically trimming the network into a Small-World model before training, (2) training the network for a given dataset, and (3) optimizing the network for accuracy. The new scheme effectively reduces both the model size and the interconnection needed before training, achieving a locally clustered and globally sparse model. We demonstrate our approach on LeNet-5 for MNIST and VGG-16 for CIF AR-10, decreasing the number of parameters to 2.3% and 9.02% of the baseline model, respectively. Recent developments in Deep Neural Networks (DNNs) have made them an integral part of modern day data processing which enable applications such as image recognition [1], object detection [2], speech recognition [3] and other applications.
Fault Detection and Identification using Bayesian Recurrent Neural Networks
Sun, Weike, Paiva, Antonio R. C., Xu, Peng, Sundaram, Anantha, Braatz, Richard D.
In processing and manufacturing industries, there has been a large push to produce higher quality products and ensure maximum efficiency of processes. This requires approaches to effectively detect and resolve disturbances to ensure optimal operations. While the control system can compensate for many types of disturbances, there are changes to the process which it still cannot handle adequately. It is therefore important to further develop monitoring systems to effectively detect and identify those faults such that they can be quickly resolved by operators. In this paper, a novel probabilistic fault detection and identification method is proposed which adopts a newly developed deep learning approach using Bayesian recurrent neural networks (BRNNs) with variational dropout. The BRNN model is general and can model complex nonlinear dynamics. Moreover, compared to traditional statistic-based data-driven fault detection and identification methods, the proposed BRNN-based method yields uncertainty estimates which allow for simultaneous fault detection of chemical processes, direct fault identification, and fault propagation analysis. The outstanding performance of this method is demonstrated and contrasted to (dynamic) principal component analysis, which are widely applied in the industry, in the benchmark Tennessee Eastman process (TEP) and a real chemical manufacturing dataset.
Modeling EEG data distribution with a Wasserstein Generative Adversarial Network to predict RSVP Events
Panwar, Sharaj, Rad, Paul, Jung, Tzyy-Ping, Huang, Yufei
Electroencephalography (EEG) data are difficult to obtain due to complex experimental setups and reduced comfort with prolonged wearing. This poses challenges to train powerful deep learning model with the limited EEG data. B eing able to generate EEG data computationally could address this limitation . We propose a novel Wasserstein Generative Adversarial Network with gradient penalty ( W GAN - GP) to synthesize EEG data. We further extend ed this network to a class - conditioned variant that also includes a classification branch to perform event - related classification. We trained the proposed networks to generate one and 64 - channel data resembling EEG signals routinely seen in a rapid serial visual presentation (RSVP) experiment and demonstrate d the validity of the generated samples . We also tested intra - subject cross - session classification performance for classifying the RSVP target events and show ed that class - conditioned W GAN - GP can achieve improved event - classification performance over EEGNet . LECTROENCEPHAL OGRAPHY (EEG) i s an attractive neuroimaging tool for measuring brain activities due to its portability, noninvasiveness and its ability to capture spatiotemporal dynamics of human brains . However, obtaining high - quality EEG data could be labor - intensive an d costly. The scarcity of high - quality EEG data poses significant challenges in the era of deep learning (DL) to train high - performing deep models to predict cognitive events and understand associated brain dynamics and mechanisms. It is thus of great interest in developing cost - effective approaches to augment the limited EEG samples so that the superb ability of DL in learning data representation can be fully exploited for EEG - based cognitive event classification.
Stronger Convergence Results for Deep Residual Networks: Network Width Scales Linearly with Training Data Size
Deep neural networks have gained remarkable success over a l arge variety of applications, including computer vision [ 1 ], natural language processing [ 2 ], speech recognition [ 3 ] and Go games [ 4 ]. But the reason why deep networks perform well over various tasks is still not exactly understood. The optimization performance of deep networks is one of the subj ects which requires an involved theoretical study, given that gradient descent can achieve zero training loss even for random labels [ 5 ], and the loss of deep networks is highly non-convex. There are different lines of works investigating the optimization of deep networks from different perspec tives. For example, a large number of works consider the optimization landscape correspondin g to different activation functions [ 6 - 11 ], whereas some others [ 12 - 15 ] ensure global convergence by imposing some restrictions o n the input distribution. In the recent years, there has been considerably many papers providing convergence guarantees for over-parameterized two-layer and deep networks. It is s hown in [ 16 ] that gradient descent can find the near-global minima of a single hidden layer network i n polynomial time with respect to the accuracy and sample size.