Adversarial Defense via Data Dependent Activation Function and Total Variation Minimization

Wang, Bao, Lin, Alex T., Shi, Zuoqiang, Zhu, Wei, Yin, Penghang, Bertozzi, Andrea L., Osher, Stanley J.

arXiv.org Machine Learning 

The adversarial vulnerability [27] of deep neural nets (DNNs) threatens their applicability in security critical tasks, e.g., autonomous cars [1], robotics [9], DNN-based malware detection systems [21, 8]. Since the pioneering work by Szegedy et al. [27], many advanced adversarial attack schemes have been devised to generate imperceptible perturbations to sufficiently fool the DNNs [7, 20, 6, 30, 12, 3]. And not only are adversarial attacks successful in white-box attacks, i.e. when the adversary has access to the DNN parameters, but attacks are also successful in black-box attacks, i.e. it has no access to the parameters. Black-box attacks are successful because one can perturb an image so it misclassifies on one DNN, and the same perturbed image also has a significant chance to be misclassified by another DNN; this is known as transferability of adversarial examples [23]. Due to the transferability of adversarial examples, it is very easy to attack neural nets in a black-box fashion [15, 5]. In fact, there exist universal perturbations that can imperceptibly perturb any image and cause misclassification for any given network [17]. There is much recent research on designing advanced adversarial attacks and defending against adversarial perturbation.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found