Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
Abstract
Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5 for the full reproducibility of our results.
Community
Your fact-checker scores well on benchmarks. But is it right for the right reasons?
π Oral Spotlight @ NeurIPS 2026 Generative AI for Healthcare (GenAI4Health) Workshop (Top 4%)
TL;DR: Retrieve-then-verify factuality evaluators can look strong on closed-ended biomedical benchmarks, but on clinician-annotated open-ended medical answers (MedExpert), even a SOTA pipeline (Qwen3 retriever + GPT-5.4 verifier) reaches only 0.06 F1. We open up that black box with two interpretable, human-validated taxonomies of where and why the pipeline fails.
Key findings
- π Aggregate P/R/F1 hide a sharp open- vs. closed-ended gap: F1 score drops from 0.78 (close-ended benchmark)β 0.06 on MedExpert (truly open-ended benchmark).
- π Retrieval β evidence: across 4 retrievers, only 18.5β42% of retrieved passages directly affirm or contradict the claim; many are merely on-topic, off-population, or from low-evidence-grade sources.
- π§ Verifier failures decompose into 6 reasoning steps (evidence selection β interpretation β clinical inference β evidence grounding β confidence calibration β label consistency).
- βοΈ In our tested settings, more reasoning effort burns far more tokens without improving recall; larger models, medical fine-tuning, and web corpora don't remove the dominant failure patterns.
- π§ͺ Breadth: 6 verifier LLMs Γ 4 retrievers Γ 3 corpora; taxonomies are LLM-assisted, human-guided, applied at scale (800 claimβevidence pairs, 576 reasoning traces) and checked against clinician adjudication.
Why it matters: If you build medical RAG or LLM-as-judge factuality evaluators, a single F1 won't tell you whether failures come from retrieval or reasoning. Our codebooks give you a diagnostic checklist. High F1 β correct verification: find what's really going on inside your medical fact-checker.
π Paper: https://arxiv.org/abs/2609.30467 Β· π» Code/data: https://anonymous.4open.science/r/Medical_RAG_eval-4AB5
π Upvote if useful, and tell us which failure mode you've hit in your own pipeline!
Get this paper in your agent:
hf papers read 2609.30467 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper