Goto

Collaborating Authors

 Deep Learning


Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect Teacher

Neural Information Processing Systems

On the other hand, recent finding on neural tangent kernel enables us to approximate a wide neural network with a linear model of the network's random features. In this paper, we theoretically analyze the knowledge distillation of a wide neural network. First we provide a transfer risk bound for the linearized model of the network. Then we propose a metric of the task's training difficulty, called data inefficiency.



A Preliminaries on Transformers

Neural Information Processing Systems

Perplexity is a widely used metric for evaluating the performance of autoregressive language models. This metric encapsulates how well the model can predict a word.



Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack Transferability

Neural Information Processing Systems

We focus on the blackbox transfer-based adversarial threat model for DNN image classifiers. In the standard case, blackbox means the attacker does not have access to the gradients of the target model and makes no assumptions about its architecture.