Learning Universal Adversarial Perturbations with Generative Models
Abstract--Neural networks are known to be vulnerable to adversarial examples, inputs that have been intentionally perturbed to remain visually similar to the source input, but cause a misclassification. It was recently shown that given a dataset and classifier, there exists so called universal adversarial perturbations, a single perturbation that causes a misclassification when applied to any input. In this work, we introduce universal adversarial networks, a generative network that is capable of fooling a target classifier when it's generated output is added to a clean sample from a dataset. We show that this technique improves on known universal adversarial attacks. Machine Learning models are increasingly relied upon for safety and business critical tasks such as in medicine [22], [30], [41], robotics and automotive [28], [32], [40], security [2], [17], [38] and financial [13], [18], [36] applications. Recent research shows that machine learning models trained on entirely uncorrupted data, are still vulnerable to adversarial examples [7], [12], [23], [24], [35], [37]: samples that have been maliciously altered so as to be misclassified by a target model while appearing unaltered to the human eye.
Jan-5-2018
- Genre:
- Research Report > New Finding (0.66)
- Industry:
- Information Technology > Security & Privacy (1.00)
- Health & Medicine > Therapeutic Area (0.94)
- Technology: