Adversarial Examples Are Not Bugs, They Are Features

Ilyas, Andrew, Santurkar, Shibani, Tsipras, Dimitris, Engstrom, Logan, Tran, Brandon, Madry, Aleksander

arXiv.org Machine Learning 

The pervasive brittleness of deep neural networks [Sze 14; Eng 19; HD19; Ath 18] has attracted significant attention in recent years. Particularly worrisome is the phenomenon of adversarial examples [Big 13; Sze 14], imperceptibly perturbed natural inputs that induce erroneous predictions in state-of-the-art classifiers. Previous work has proposed a variety of explanations for this phenomenon, ranging from theoretical models [Sch 18; BPR18] to arguments based on concentration of measure in high-dimensions [Gil 18; MDM18; Sha 19a]. These theories, however, are often unable to fully capture behaviors we observe in practice (we discuss this further in Section 5). More broadly, previous work in the field tends to view adversarial examples as aberrations arising either from the high dimensional nature of the input space or statistical fluctuations in the training data [Sze 14; GSS15; Gil 18]. From this point of view, it is natural to treat adversarial robustness as a goal that can be disentangled and pursued independently from maximizing accuracy [Mad 18; SHS19; Sug 19], either through improved standard regularization methods [TG16] or pre/post-processing of network inputs/outputs [Ues 18; CW17a; He 17].

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found