Papers
arxiv:2609.30467

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

Published on Sep 24
Β· Submitted by
HeyuanHuang
on Oct 5
Authors:
,
,
,
,
,
,
,

Abstract

Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5 for the full reproducibility of our results.

Community

Your fact-checker scores well on benchmarks. But is it right for the right reasons?

πŸŽ‰ Oral Spotlight @ NeurIPS 2026 Generative AI for Healthcare (GenAI4Health) Workshop (Top 4%)

TL;DR: Retrieve-then-verify factuality evaluators can look strong on closed-ended biomedical benchmarks, but on clinician-annotated open-ended medical answers (MedExpert), even a SOTA pipeline (Qwen3 retriever + GPT-5.4 verifier) reaches only 0.06 F1. We open up that black box with two interpretable, human-validated taxonomies of where and why the pipeline fails.

Key findings

  • πŸ“‰ Aggregate P/R/F1 hide a sharp open- vs. closed-ended gap: F1 score drops from 0.78 (close-ended benchmark)β†’ 0.06 on MedExpert (truly open-ended benchmark).
  • πŸ”Ž Retrieval β‰  evidence: across 4 retrievers, only 18.5–42% of retrieved passages directly affirm or contradict the claim; many are merely on-topic, off-population, or from low-evidence-grade sources.
  • 🧠 Verifier failures decompose into 6 reasoning steps (evidence selection β†’ interpretation β†’ clinical inference β†’ evidence grounding β†’ confidence calibration β†’ label consistency).
  • βš™οΈ In our tested settings, more reasoning effort burns far more tokens without improving recall; larger models, medical fine-tuning, and web corpora don't remove the dominant failure patterns.
  • πŸ§ͺ Breadth: 6 verifier LLMs Γ— 4 retrievers Γ— 3 corpora; taxonomies are LLM-assisted, human-guided, applied at scale (800 claim–evidence pairs, 576 reasoning traces) and checked against clinician adjudication.

Why it matters: If you build medical RAG or LLM-as-judge factuality evaluators, a single F1 won't tell you whether failures come from retrieval or reasoning. Our codebooks give you a diagnostic checklist. High F1 β‰  correct verification: find what's really going on inside your medical fact-checker.

πŸ“„ Paper: https://arxiv.org/abs/2609.30467 Β· πŸ’» Code/data: https://anonymous.4open.science/r/Medical_RAG_eval-4AB5

πŸ‘ Upvote if useful, and tell us which failure mode you've hit in your own pipeline!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.30467
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.30467 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.30467 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.30467 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.