Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models

Xu, Zane, Sun, Jason

arXiv.org Artificial Intelligence 

A central challenge in this domain is the inherent trade-off between enhancing adversarial robustness and preserving the model's zero-shot generalization capabilities. We analyze two primary defense paradigms that have emerged from recent research. The first, Adversarial Fine-Tuning (AFT), involves modifying model parameters. We trace its evolution from early methods focused on protecting vision-language alignment (TeCoA) and mitigating overfitting through pre-trained model guidance (PMG-AFT, TGA-ZSR), to more advanced strategies that proactively re-engineer the embedding space's geometry to build intrinsic robustness (LAA T, TIMA). The second paradigm, Training-Free and Test-Time Defenses, avoids parameter modification altogether. We examine its progression from simple, heuristic-based input manipulations (AOM, TTC) to theoretically grounded purification methods in the model's latent space (CLIPure). By analyzing these works, we distill the field's core problems--including the robustness-generalization dilemma, geometric vulnerabilities, and computational costs--and identify key insights. We conclude by outlining promising future directions, such as the development of hybrid defense models and the pursuit of large-scale adversarial pre-training to create natively robust foundation models.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found