Attributed Decision Model — evidence model
Demo: Ask your contract, where every answer is shown with the quotes it was made from · Companion model: decider
Query-conditioned verbatim evidence extractor over documents of up to 8,192 tokens per window (longer documents are read in overlapping windows). Given a question or claim and a document, it returns character spans of the document that bear on it. It is the first half of a two-model pipeline whose decisions are faithful by construction: the companion decider answers from these quotes only.
Fine-tuned from KRLabsOrg/verbatim-rag-modern-bert-v2
(Apache-2.0; itself built on Alibaba-NLP/gte-reranker-modernbert-base) on documents with gold evidence across
contracts, Wikipedia, multi-hop QA, science, medicine and tables, with v2's own training data replayed so its original
domains are kept. Same interface as v2:
from transformers import AutoModel
model = AutoModel.from_pretrained("<this repo>", trust_remote_code=True)
out = model.process(question="Is the receiving party allowed to share information with employees?",
context=contract_text, threshold=0.5)
for s in out["spans"]:
print(s["score"], s["text"])
Evaluation
ev-s1 is the earlier stage-1 checkpoint. End-to-end rows use the calibrated stage-1b decider at threshold 0.5.
| Benchmark | v2 | ev-s1 | this model |
|---|---|---|---|
| ContractNLI test, evidence F1 vs gold units | 0.371 | 0.786 | 0.791 |
| ContractNLI test, end to end with the decider, macro-F1 | 0.663 | 0.828 | 0.826 |
| Medical consultations (validation), accuracy / top-span tIoU | 0.839 / 0.298 | 0.951 / 0.489 | 0.954 / 0.371 |
| Evidence Inference test (never trained on), word F1 | 0.305 | 0.282 | 0.281 |
| Evidence Inference test, end to end, macro-F1 | 0.581 | 0.596 | 0.589 |
| ACL test (v2's domain), word F1 @0.2 / @0.5 | 0.470 / 0.402 | 0.426 / 0.292 | 0.466 / 0.399 |
Gold recall within the decider's budget (share of gold evidence characters the decider actually reads), threshold 0.3: ContractNLI 0.82, Evidence Inference 0.62, medical transcripts 0.90.
Training data
Tier A sources with character-level gold evidence: ContractNLI, CUAD, BoolQ, MultiRC, Movie Rationales, SciFact, FEVER, HotpotQA, Natural Questions, IIRC, MuSiQue, QASPER, MASH-QA, 2WikiMultiHopQA, FEVEROUS (tables rendered as rows), plus replay of KRLabsOrg/verbatim-spans (ACL, RAGBench, Squeez). Sources that mark only some of the evidence are trained with a lower weight on unmarked tokens.
Licence
CC BY-NC 4.0: research and non-commercial use only. The base model is Apache-2.0, but the training data includes non-commercial sources: MultiRC (CogComp Research and Academic Use License) and Movie Rationales (no licence).
- Permissive: ContractNLI, CUAD, SciFact, QASPER, MuSiQue (CC BY 4.0); 2WikiMultiHopQA, MASH-QA, KRLabsOrg/verbatim-spans (Apache-2.0).
- Share-alike: BoolQ, Natural Questions, FEVER, FEVEROUS (CC BY-SA 3.0); HotpotQA (CC BY-SA 4.0).
- Licence not stated: IIRC (Hugging Face mirror).
- The RAGBench part of verbatim-spans was labelled by GPT-4o.
Source-by-source details: LICENSE_AUDIT.md in the project repository.
Limitations
- Spans are broader than v2's: on a medical-transcript validation set, the overlap between the top span and the annotated interval drops from 0.49 (ev-s1) to 0.37. Use ev-s1 if tight spans matter more than recall.
- Recall is the main gap: on Evidence Inference (biomedical papers) about a third of items get no gold evidence among the quotes at threshold 0.5.
- English only.
- Downloads last month
- 8
Model tree for ricardocs99/attributed-decision-evidence
Base model
answerdotai/ModernBERT-base