Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
Li, Yang, Wang, Chenyu, Wang, Tingrui, Wang, Yongwei, Li, Haonan, Liu, Zhunga, Pan, Quan
–arXiv.org Artificial Intelligence
We identify and address the challenge of attacking fundamentally different model architectures in black-box settings. JAD is the first generative attack framework to bridge CNN and Transformer vulnerabilities, enabling effective transferable attacks across architectures. We introduce a novel attention distillation technique that fuses saliency maps from a CNN and a ViT into the training of a latent diffusion generator. This guides the generator to align perturbations with critical regions common to both architectures, greatly enhancing attack generalization. By focusing on shared weak spots and leveraging a diffusion model, JAD achieves higher success rates on both CNN and ViT targets than existing methods (e.g., CDMA) under stringent query limits. Our approach demonstrates state-of-the-art transferability across heterogeneous models while often requiring only a single query at attack time, marking a substantial step toward practical and architecture-agnostic black-box adversarial attacks. These results underline the potential of combining insights from disparate network architectures within a generative attack paradigm. By jointly distilling CNN and Transformer perspectives into adversarial example generation, JAD opens a new avenue for highly transferable and query-efficient black-box attacks, moving closer to real-world applicability.
arXiv.org Artificial Intelligence
Sep-24-2025
- Genre:
- Research Report (0.83)
- Industry:
- Transportation > Air (1.00)
- Information Technology > Security & Privacy (1.00)
- Technology: