Adversarial Robustness by Design through Analog Computing and Synthetic Gradients
Cappelli, Alessandro, Ohana, Ruben, Launay, Julien, Meunier, Laurent, Poli, Iacopo, Krzakala, Florent
–arXiv.org Artificial Intelligence
Neural networks are sensitive to small, imperceptible to humans, perturbations of their inputs that can cause state-of-the-art classifiers to completely fail [1]. As deep learning models are deployed in realworld applications, guaranteeing their robustness to malicious actors becomes increasingly important: for instance, an adversarial image could evade automated content filtering on social networks [2]. Adversarial attacks can be carried out in different frameworks: in the white-box setting, the attacker has full access to the model, while black-box attacks only rely on queries. It is also possible to craft an attack on a different model and transfer it to the model targeted [3]. There is no universal defense, and state-ofthe-art techniques often come with a large computational cost, as well as reduced natural accuracy [4]. Some of these defenses rely on obfuscated gradients: the model is designed so that the gradients are unsuitable for attacks, for instance by using non-differentiable layers. However, attackers can choose to alter the network structure, using Backward Pass Differentiable Approximation (BPDA) [5], replacing obfuscating layers with well-behaved approximations. Furthermore, approaches relying on obfuscation do not generally provide robustness against transfer and black-box attacks.
arXiv.org Artificial Intelligence
Jan-6-2021
- Country:
- North America
- United States > California
- Los Angeles County > Long Beach (0.04)
- Canada > Ontario
- Toronto (0.04)
- United States > California
- Europe
- Switzerland (0.04)
- France > Île-de-France
- North America
- Genre:
- Research Report (0.82)
- Industry:
- Information Technology > Security & Privacy (0.50)
- Government > Military (0.36)
- Technology: