Papers
arxiv:2608.26762

Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers

Published on Sep 26
· Submitted by
Markus Frohmann
on Oct 5
Authors:
,
,
,

Abstract

In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.

Community

Paper submitter

Rerankers, reward models and multi-document QA systems can score several candidates together in one LLM prompt. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, what a reader answers, or which response is preferred.
This paper shows that equal ranking quality does not imply equal decisions. On passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when the candidates are reordered. No prompt-time change tested resolves this.

The paper introduces order-consistency SFT (OC-SFT), which reduces this order dependence in the weights by penalizing disagreement between a candidate's scores across orderings. It:
• holds ranking quality
• leads every decision-stability measure among trained scorers, across passage reranking, response ranking and multi-document QA
• is more stable than order-averaged distillation on 12 base models across three families
• with a single permutation, retains sets that overlap more than ten averaged permutations of an off-the-shelf scorer

💡 Takeaway: a comparison of such scorers should report what a threshold retains and a reader answers, not ranking quality alone.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.26762
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.26762 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.26762 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.26762 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.