A Method for Computing Class-wise Universal Adversarial Perturbations
Gupta, Tejus, Sinha, Abhishek, Kumari, Nupur, Singh, Mayank, Krishnamurthy, Balaji
We present an algorithm for computing class-specific universal adversarial perturbations for deep neural networks. Such perturbations can induce mis-classification in a large fraction of images of a specific class. Unlike previous methods that use iterative optimization for computing a universal perturbation, the proposed method employs a perturbation that is a linear function of weights of the neural network and hence can be computed much faster. The method does not require any training data and has no hyper-parameters. We also study the characteristics of the decision boundaries learned by standard and adversarially trained models to understand the universal adversarial perturbations. The vulnerability of state-of-the-art neural networks to adversarial perturbations was first studied in (Szegedy et al., 2014).
Dec-1-2019