Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
Xu, Wenda, Zhu, Guanglei, Zhao, Xuandong, Pan, Liangming, Li, Lei, Wang, William Yang
–arXiv.org Artificial Intelligence
Recent studies show that large language models (LLMs) improve their performance through self-feedback on certain tasks while degrade on others. We discovered that such a contrary is due to LLM's bias in evaluating their own output. In this paper, we formally define LLM's self-bias - the tendency to favor its own generation - using two statistics. We analyze six LLMs (GPT-4, GPT-3.5, Gemini, LLaMA2, Mixtral and DeepSeek) on translation, constrained text generation, and mathematical reasoning tasks. We find that self-bias is prevalent in all examined LLMs across multiple languages and tasks. Our analysis reveals that while the self-refine pipeline improves the fluency and understandability of model outputs, it further amplifies self-bias. To mitigate Figure 1: How LLM's self-feedback inflates scores compared such biases, we discover that larger model to human assessment. Bias is the mean difference size and external feedback with accurate assessment between LLM and human scores, while skewness can significantly reduce bias in the (Dskew) measures the asymmetry of their distribution self-refine pipeline, leading to actual performance around zero. Non-biased estimation will have Dskew=0.
arXiv.org Artificial Intelligence
Jun-18-2024
- Country:
- North America > United States
- Europe > Austria
- Vienna (0.04)
- Asia
- Singapore (0.05)
- Middle East > UAE
- Abu Dhabi Emirate > Abu Dhabi (0.04)
- Genre:
- Research Report > New Finding (1.00)
- Technology: