Goto

Collaborating Authors

 Deep Learning


How transferable are features in deep neural networks?

Neural Information Processing Systems

Many deep neural networks trained on natural images exhibit a curious phenomenon in common: on the first layer they learn features similar to Gabor filters and color blobs. Such first-layer features appear not to be specific to a particular dataset or task, but general in that they are applicable to many datasets and tasks. Features must eventually transition from general to specific by the last layer of the network, but this transition has not been studied extensively. In this paper we experimentally quantify the generality versus specificity of neurons in each layer of a deep convolutional neural network and report a few surprising results. Transferability is negatively affected by two distinct issues: (1) the specialization of higher layer neurons to their original task at the expense of performance on the target task, which was expected, and (2) optimization difficulties related to splitting networks between co-adapted neurons, which was not expected. In an example network trained on ImageNet, we demonstrate that either of these two issues may dominate, depending on whether features are transferred from the bottom, middle, or top of the network. We also document that the transferability of features decreases as the distance between the base task and target task increases, but that transferring features even from distant tasks can be better than using random features. A final surprising result is that initializing a network with transferred features from almost any number of layers can produce a boost to generalization that lingers even after fine-tuning to the target dataset.







Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal Attention

Neural Information Processing Systems

Although recent attempts to apply GNN to the FC network have shown promising results, there is still a common limitation that they usually do not incorporate the dynamic characteristics of the FC network which fluctuates over time.


The Synthesis of XNOR Recurrent Neural Networks with Stochastic Logic

Neural Information Processing Systems

These limitations make LSTM models difficult to deploy on specialized hardware requiring real-time processes with inferior hardware resources and power budget.


1f6591cc41be737e9ba4cc487ac8082d-Paper-Conference.pdf

Neural Information Processing Systems

Although OOD detection methods have advanced by a great deal, they are still susceptible to adversarial examples, which is a violation of their purpose. To mitigate this issue, several defenses have recently been proposed. Nevertheless, these efforts remained ineffective, as their evaluations are based on either small perturbation sizes, or weak attacks.


On Uniform Convergence and Low-Norm Interpolation Learning

Neural Information Processing Systems

In the past several years, it has become empirically clear that - contrary to traditional intuition - it is possible for models which exactly interpolate noisy training data to reliably generalize well on practical problems, especially in deep learning [7, 25, 34].