Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models
–arXiv.org Artificial Intelligence
A central challenge in this domain is the inherent trade-off between enhancing adversarial robustness and preserving the model's zero-shot generalization capabilities. We analyze two primary defense paradigms that have emerged from recent research. The first, Adversarial Fine-Tuning (AFT), involves modifying model parameters. We trace its evolution from early methods focused on protecting vision-language alignment (TeCoA) and mitigating overfitting through pre-trained model guidance (PMG-AFT, TGA-ZSR), to more advanced strategies that proactively re-engineer the embedding space's geometry to build intrinsic robustness (LAA T, TIMA). The second paradigm, Training-Free and Test-Time Defenses, avoids parameter modification altogether. We examine its progression from simple, heuristic-based input manipulations (AOM, TTC) to theoretically grounded purification methods in the model's latent space (CLIPure). By analyzing these works, we distill the field's core problems--including the robustness-generalization dilemma, geometric vulnerabilities, and computational costs--and identify key insights. We conclude by outlining promising future directions, such as the development of hybrid defense models and the pursuit of large-scale adversarial pre-training to create natively robust foundation models.
arXiv.org Artificial Intelligence
Aug-8-2025
- Genre:
- Research Report (0.83)
- Industry:
- Information Technology (0.46)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Machine Learning (1.00)
- Natural Language > Large Language Model (0.65)
- Information Technology > Artificial Intelligence