AI Bias Mitigation Efforts May Reduce Model Confidence Rather Than Correcting Underlying Preferences
Research indicates that current methods for mitigating bias in large language models (LLMs), particularly activation steering, may not be effectively correcting the models' underlying biases. Instead, these methods appear to reduce the models' confidence, leading to an increased tendency to abstain from answering or to provide more generalized, less specific responses. This effect is observed across various LLMs, including Llama-3.1-8B, Falcon3-7B, Qwen3.5-9B, and Ministral-3-8B. When models are given an option to abstain, steering towards debiasing directions significantly increases the probability of choosing the abstain option. In scenarios where abstention is not an option, the steering leads to an equilibration of answer probabilities, increasing entropy and decreasing overall confidence. This suggests that the debiasing directions identified in latent space are highly aligned with model confidence rather than solely targeting and correcting biased preferences.