Promoting Equality in Large Language Models: Identifying and Mitigating the Implicit Bias based on Bayesian Theory
Deng, Yongxin, Qiu, Xihe, Tan, Xiaoyu, Pan, Jing, Jue, Chen, Fang, Zhijun, Xu, Yinghui, Chu, Wei, Qi, Yuan
–arXiv.org Artificial Intelligence
Large language models (LLMs) are trained on extensive text corpora, which inevitably include biased information. Although techniques such as Affective Alignment can mitigate some negative impacts of these biases, existing prompt-based attack methods can still extract these biases from the model's weights. Moreover, these biases frequently appear subtly when LLMs are prompted to perform identical tasks across different demographic groups, thereby camouflaging their presence. To address this issue, we have formally defined the implicit bias problem and developed an innovative framework for bias removal based on Bayesian theory, Bayesian-Theory based Bias Removal (BTBR). BTBR employs likelihood ratio screening to pinpoint data entries within publicly accessible biased datasets that represent biases inadvertently incorporated during the LLM training phase. It then automatically constructs relevant knowledge triples and expunges bias information from LLMs using model editing techniques. Through extensive experimentation, we have confirmed the presence of the implicit bias problem in LLMs and demonstrated the effectiveness of our BTBR approach.
arXiv.org Artificial Intelligence
Aug-20-2024
- Country:
- North America
- Dominican Republic (0.04)
- United States > Louisiana
- Orleans Parish > New Orleans (0.04)
- Europe
- Austria (0.04)
- Italy > Tuscany
- Florence (0.04)
- Croatia > Dubrovnik-Neretva County
- Dubrovnik (0.04)
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Asia
- Middle East
- UAE > Abu Dhabi Emirate
- Abu Dhabi (0.04)
- Israel > Tel Aviv District
- Tel Aviv (0.04)
- UAE > Abu Dhabi Emirate
- China
- Middle East
- North America
- Genre:
- Research Report (1.00)
- Industry:
- Information Technology (0.46)
- Technology: