Adversarial Robustness by Design through Analog Computing and Synthetic Gradients

Cappelli, Alessandro, Ohana, Ruben, Launay, Julien, Meunier, Laurent, Poli, Iacopo, Krzakala, Florent

arXiv.org Artificial Intelligence 

Neural networks are sensitive to small, imperceptible to humans, perturbations of their inputs that can cause state-of-the-art classifiers to completely fail [1]. As deep learning models are deployed in realworld applications, guaranteeing their robustness to malicious actors becomes increasingly important: for instance, an adversarial image could evade automated content filtering on social networks [2]. Adversarial attacks can be carried out in different frameworks: in the white-box setting, the attacker has full access to the model, while black-box attacks only rely on queries. It is also possible to craft an attack on a different model and transfer it to the model targeted [3]. There is no universal defense, and state-ofthe-art techniques often come with a large computational cost, as well as reduced natural accuracy [4]. Some of these defenses rely on obfuscated gradients: the model is designed so that the gradients are unsuitable for attacks, for instance by using non-differentiable layers. However, attackers can choose to alter the network structure, using Backward Pass Differentiable Approximation (BPDA) [5], replacing obfuscating layers with well-behaved approximations. Furthermore, approaches relying on obfuscation do not generally provide robustness against transfer and black-box attacks.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found