Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections

Maloyan, Narek, Namiot, Dmitry

arXiv.org Artificial Intelligence 

Abstract--Large Language Models (LLMs) are increasingly used as automated judges for evaluating text quality, code c or-rectness, and argument strength. However, these LLM-as-a-judge systems are vulnerable to adversarial attacks that can mani pulate their assessments. This paper investigates the vulnerabil ity of LLM-as-a-judge systems to prompt injection attacks, drawi ng insights from both academic literature and practical solutio ns from the "LLMs: Y ou Can't Please Them All" Kaggle competition. We present a comprehensive framework for developing and evaluating adversarial attacks against LLM judges, distinguis hing between content-author attacks and system-prompt attacks . Through rigorous statistical a nalysis (n=50 prompts per condition, bootstrap confidence interval s), we demonstrate that sophisticated attacks can achieve succes s rates of up to 73.8% against popular LLM judges, with Contextual Misdirection being the most effective method against Gemma models at 67.7%. We find that smaller models like Gemma-3-4B-Instruct are more vulnerable (65.9% average success rat e) than their larger counterparts, and that attacks show high transferability (50.5-62.6%) We compare our approach with recent work including Universal-Prompt-Injection [1] and AdvPrompter [2], demonstrating b oth complementary insights and novel contributions. Our findin gs highlight critical vulnerabilities in current LLM-as-a-j udge systems and provide recommendations for developing more robus t evaluation frameworks, including using multi-model commi ttees with diverse architectures and preferring comparative ass essment over absolute scoring methods. T o ensure reproducibility, we release our code, evaluation harness, and processed datase ts.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found