Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Xu, Wenda, Zhu, Guanglei, Zhao, Xuandong, Pan, Liangming, Li, Lei, Wang, William Yang

arXiv.org Artificial Intelligence 

Recent studies show that large language models (LLMs) improve their performance through self-feedback on certain tasks while degrade on others. We discovered that such a contrary is due to LLM's bias in evaluating their own output. In this paper, we formally define LLM's self-bias - the tendency to favor its own generation - using two statistics. We analyze six LLMs (GPT-4, GPT-3.5, Gemini, LLaMA2, Mixtral and DeepSeek) on translation, constrained text generation, and mathematical reasoning tasks. We find that self-bias is prevalent in all examined LLMs across multiple languages and tasks. Our analysis reveals that while the self-refine pipeline improves the fluency and understandability of model outputs, it further amplifies self-bias. To mitigate Figure 1: How LLM's self-feedback inflates scores compared such biases, we discover that larger model to human assessment. Bias is the mean difference size and external feedback with accurate assessment between LLM and human scores, while skewness can significantly reduce bias in the (Dskew) measures the asymmetry of their distribution self-refine pipeline, leading to actual performance around zero. Non-biased estimation will have Dskew=0.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found