Shifting Perspectives: Steering Vector Ensembles for Robust Bias Mitigation in LLMs
Siddique, Zara, Khalid, Irtaza, Turner, Liam D., Espinosa-Anke, Luis
–arXiv.org Artificial Intelligence
We present a novel approach to bias mitigation in large language models (LLMs) by applying steering vectors to modify model activations in forward passes. We employ Bayesian optimization to systematically identify effective contrastive pair datasets across nine bias axes. When optimized on the BBQ dataset, our individually tuned steering vectors achieve average improvements of 12.2%, 4.7%, and 3.2% over the baseline for Mistral, Llama, and Qwen, respectively. Building on these promising results, we introduce Steering Vector Ensembles (SVE), a method that averages multiple individually optimized steering vectors, each targeting a specific bias axis such as age, race, or gender. By leveraging their collective strength, SVE outperforms individual steering vectors in both bias reduction and maintaining model performance. The work presents the first systematic investigation of steering vectors for bias mitigation, and we demonstrate that SVE is a powerful and computationally efficient strategy for reducing bias in LLMs, with broader implications for enhancing AI safety.
arXiv.org Artificial Intelligence
Mar-7-2025
- Country:
- North America
- United States
- New York > New York County
- New York City (0.04)
- Florida > Miami-Dade County
- Miami (0.04)
- New York > New York County
- Mexico > Mexico City
- Mexico City (0.04)
- Canada > Ontario
- Toronto (0.04)
- United States
- Europe
- United Kingdom (0.04)
- Italy (0.04)
- Middle East > Malta
- Eastern Region > Northern Harbour District > St. Julian's (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Asia
- North America
- Genre:
- Research Report > New Finding (1.00)
- Technology: