Papers
arxiv:2609.32720

BiasReducer: Adaptive Bias Mitigation for Reward Models

Published on Sep 26
· Submitted by
Shuang Liu
on Oct 1
Authors:
,
,
,
,

Abstract

Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.

Community

Paper submitter

Reward models often favor superficial attributes like length, confident tone, or sycophancy, so LLMs trained on them learn to score higher without being more correct.

We introduce BiasReducer, a lightweight framework that edits only the linear reward head, with no retraining of the reward model. For each new dataset, it ranks attributes by their influence on reward scores, selects the relevant ones, and applies the matching edits, instead of one fixed correction for one pre-specified bias.

Across five public reward models, BiasReducer-M improves pairwise preference accuracy on three bias-focused benchmarks by 8.3 / 18.0 / 6.9 percentage points on average, outperforming two fine-tuning baselines. It beats the original reward model in every model–benchmark combination.

📄 Paper: https://arxiv.org/abs/2609.32720
💻 Code: https://github.com/olivialiu121/BiasReducer-RM

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32720 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32720 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32720 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.