Learn2Weight: Parameter Adaptation against Similar-domain Adversarial Attacks
–arXiv.org Artificial Intelligence
Recent work in black-box adversarial attacks for NLP systems has attracted much attention. Prior black-box attacks assume that attackers can observe output labels from target models based on selected inputs. In this work, inspired by adversarial transferability, we propose a new type of black-box NLP adversarial attack that an attacker can choose a similar domain and transfer the adversarial examples to the target domain and cause poor performance in target model. Based on domain adaptation theory, we then propose a defensive strategy, called Learn2Weight, which trains to predict the weight adjustments for a target model in order to defend against an attack of similar-domain adversarial examples. Using Amazon multi-domain sentiment classification datasets, we empirically show that Learn2Weight is effective against the attack compared to standard black-box defense methods such as adversarial training and defensive distillation. This work contributes to the growing literature on machine learning safety.
arXiv.org Artificial Intelligence
Sep-20-2022
- Country:
- North America > United States
- Utah > Salt Lake County
- Salt Lake City (0.04)
- New York > New York County
- New York City (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Massachusetts > Middlesex County
- Cambridge (0.04)
- California > Los Angeles County
- Long Beach (0.04)
- Utah > Salt Lake County
- Europe
- United Kingdom > England
- Oxfordshire > Oxford (0.04)
- Italy > Tuscany
- Florence (0.04)
- France > Hauts-de-France
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- United Kingdom > England
- Asia > China
- Hong Kong (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Information Technology > Security & Privacy (1.00)
- Health & Medicine (1.00)
- Banking & Finance (0.93)
- Government > Military (0.92)
- Technology: