Benchmarking and Improving LLM Robustness for Personalized Generation
Okite, Chimaobi, Deng, Naihao, Bodipati, Kiran, Hou, Huaidian, Chai, Joyce, Mihalcea, Rada
–arXiv.org Artificial Intelligence
Recent years have witnessed a growing interest in personalizing the responses of large language models (LLMs). While existing evaluations primarily focus on whether a response aligns with a user's preferences, we argue that factuality is an equally important yet often overlooked dimension. In the context of personalization, we define a model as robust if its responses are both factually accurate and align with the user preferences. To assess this, we introduce PERG, a scalable framework for evaluating robustness in LLMs, along with a new dataset, PERGData. We evaluate fourteen models from five different model families using different prompting methods. Our findings show that current LLMs struggle with robust personalization: even the strongest models (GPT-4.1, LLaMA3-70B) fail to maintain correctness in 5% of previously successful cases without personalization, while smaller models (e.g., 7B-scale) can fail more than 20% of the time. Further analysis reveals that robustness is significantly affected by the nature of the query and the type of user preference. To mitigate these failures, we propose Pref-Aligner, a two-stage approach that improves robustness by an average of 25% across models. Our work highlights critical gaps in current evaluation practices and introduces tools and metrics to support more reliable, user-aligned LLM deployments.
arXiv.org Artificial Intelligence
Sep-25-2025
- Country:
- Asia > Thailand
- Europe
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Middle East > Malta
- Eastern Region > Northern Harbour District > St. Julian's (0.04)
- Ireland > Leinster
- North America
- Canada > Ontario
- Toronto (0.04)
- Dominican Republic (0.04)
- United States
- Florida > Miami-Dade County
- Miami (0.04)
- Michigan (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Florida > Miami-Dade County
- Canada > Ontario
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Education
- Curriculum > Subject-Specific Education (1.00)
- Educational Setting (0.68)
- Government (0.93)
- Health & Medicine > Therapeutic Area
- Cardiology/Vascular Diseases (0.67)
- Education
- Technology: