Attributed Decision Model — evidence model

Demo: Ask your contract, where every answer is shown with the quotes it was made from · Companion model: decider

Query-conditioned verbatim evidence extractor over documents of up to 8,192 tokens per window (longer documents are read in overlapping windows). Given a question or claim and a document, it returns character spans of the document that bear on it. It is the first half of a two-model pipeline whose decisions are faithful by construction: the companion decider answers from these quotes only.

Fine-tuned from KRLabsOrg/verbatim-rag-modern-bert-v2 (Apache-2.0; itself built on Alibaba-NLP/gte-reranker-modernbert-base) on documents with gold evidence across contracts, Wikipedia, multi-hop QA, science, medicine and tables, with v2's own training data replayed so its original domains are kept. Same interface as v2:

from transformers import AutoModel
model = AutoModel.from_pretrained("<this repo>", trust_remote_code=True)
out = model.process(question="Is the receiving party allowed to share information with employees?",
                    context=contract_text, threshold=0.5)
for s in out["spans"]:
    print(s["score"], s["text"])

Evaluation

ev-s1 is the earlier stage-1 checkpoint. End-to-end rows use the calibrated stage-1b decider at threshold 0.5.

Benchmark v2 ev-s1 this model
ContractNLI test, evidence F1 vs gold units 0.371 0.786 0.791
ContractNLI test, end to end with the decider, macro-F1 0.663 0.828 0.826
Medical consultations (validation), accuracy / top-span tIoU 0.839 / 0.298 0.951 / 0.489 0.954 / 0.371
Evidence Inference test (never trained on), word F1 0.305 0.282 0.281
Evidence Inference test, end to end, macro-F1 0.581 0.596 0.589
ACL test (v2's domain), word F1 @0.2 / @0.5 0.470 / 0.402 0.426 / 0.292 0.466 / 0.399

Gold recall within the decider's budget (share of gold evidence characters the decider actually reads), threshold 0.3: ContractNLI 0.82, Evidence Inference 0.62, medical transcripts 0.90.

Training data

Tier A sources with character-level gold evidence: ContractNLI, CUAD, BoolQ, MultiRC, Movie Rationales, SciFact, FEVER, HotpotQA, Natural Questions, IIRC, MuSiQue, QASPER, MASH-QA, 2WikiMultiHopQA, FEVEROUS (tables rendered as rows), plus replay of KRLabsOrg/verbatim-spans (ACL, RAGBench, Squeez). Sources that mark only some of the evidence are trained with a lower weight on unmarked tokens.

Licence

CC BY-NC 4.0: research and non-commercial use only. The base model is Apache-2.0, but the training data includes non-commercial sources: MultiRC (CogComp Research and Academic Use License) and Movie Rationales (no licence).

  • Permissive: ContractNLI, CUAD, SciFact, QASPER, MuSiQue (CC BY 4.0); 2WikiMultiHopQA, MASH-QA, KRLabsOrg/verbatim-spans (Apache-2.0).
  • Share-alike: BoolQ, Natural Questions, FEVER, FEVEROUS (CC BY-SA 3.0); HotpotQA (CC BY-SA 4.0).
  • Licence not stated: IIRC (Hugging Face mirror).
  • The RAGBench part of verbatim-spans was labelled by GPT-4o.

Source-by-source details: LICENSE_AUDIT.md in the project repository.

Limitations

  • Spans are broader than v2's: on a medical-transcript validation set, the overlap between the top span and the annotated interval drops from 0.49 (ev-s1) to 0.37. Use ev-s1 if tight spans matter more than recall.
  • Recall is the main gap: on Evidence Inference (biomedical papers) about a third of items get no gold evidence among the quotes at threshold 0.5.
  • English only.
Downloads last month
8
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ricardocs99/attributed-decision-evidence

Collection including ricardocs99/attributed-decision-evidence