Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers
Abstract
In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.
Community
Rerankers, reward models and multi-document QA systems can score several candidates together in one LLM prompt. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, what a reader answers, or which response is preferred.
This paper shows that equal ranking quality does not imply equal decisions. On passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when the candidates are reordered. No prompt-time change tested resolves this.
The paper introduces order-consistency SFT (OC-SFT), which reduces this order dependence in the weights by penalizing disagreement between a candidate's scores across orderings. It:
• holds ranking quality
• leads every decision-stability measure among trained scorers, across passage reranking, response ranking and multi-document QA
• is more stable than order-averaged distillation on 12 base models across three families
• with a single permutation, retains sets that overlap more than ten averaged permutations of an off-the-shelf scorer
💡 Takeaway: a comparison of such scorers should report what a threshold retains and a reader answers, not ranking quality alone.
Get this paper in your agent:
hf papers read 2608.26762 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper