Instructions to use subhamoybasu/laya-evidence-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use subhamoybasu/laya-evidence-v2 with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya Evidence
Laya Evidence finds the sentences (or list items) of a document that support or contradict
each of one or more claims. It adds a fourth request type, evidence, to
Laya, next to choice, score and noul, and packs the
claims into one encoder input per document window.
It is a specialist. Its strongest evaluated setting is the contract questions it was trained on (ContractNLI's 17 hypothesis types), asked of contracts it has not seen. On claim types and kinds of document unlike its training data it is much weaker (see Limitations). It is not a general-purpose fact checker.
Summary
For each claim the model returns a status (evidence_found or insufficient_evidence) and the
segments it keeps. Each segment comes with its character offsets in the document, its exact text, a
label (supports or contradicts), an evidence score and a polarity confidence.
Each encoder row is [CLS] head [SEP] [MASK] claim 1 … [MASK] claim n [SEP] document window [SEP].
The row runs through Laya's encoder; Laya's head layers are not used on this path. The rows are centred, a trained evidence type vector marking the request type is added to every token vector, and the rows are layer-normalised and passed through 2 adapter layers. An evidence head then scores every (claim, segment) pair as none, supports or
contradicts, from the claim's [MASK] vector and the segment's token vectors. A segment is kept when
its evidence score, P(supports) + P(contradicts), is at least the threshold τ.
This repository holds only the parameters Laya Evidence trains: the evidence type vector, a LayerNorm, 2 adapter layers and the evidence head, plus LoRA adapters on Laya's encoder (lora_adapter/). The Laya checkpoint
is downloaded from the Hub at a pinned revision when the model is loaded (see Base model).
| Developed by | Subhamoy Basu (subhamoy.basu@berkeley.edu) |
| Code | https://github.com/subhamoy-basu/laya-evidence |
| License | CC BY-SA 4.0 (see LICENSE and NOTICE): trained on FEVER (CC BY-SA 3.0); the code is Apache-2.0 |
| Base checkpoint | convaiinnovations/laya (root checkpoint) at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 |
| Configuration | C: Laya's encoder gets trained LoRA adapters on the evidence path; its base weights and head layers stay frozen |
| Encoder budget | 1,024 tokens per row |
| Claims per encoder pass | up to 4, within 448 claim tokens |
| Default threshold τ | 0.88, the value that maximises evidence F1 on the dev data of the training datasets (almost all of it ContractNLI's) |
| Segmenter | pysbd 0.3.4 (or the caller's own segments) |
| Training run | v2_joint_C_cuad_aug_fever_s13, seed 13, code at commit cb765a989d4ca93787afd6f12de7802279ac9501 |
| Weights | evidence.safetensors, sha256 ffe1bf2a25f7e241331dbc1c79f99b319d2b682b5d2e50d52c26c33ea7782fc0; LoRA adapter lora_adapter/adapter_model.safetensors, sha256 97d614501bd7b253bc3c5e8a50952914920aa582f79f0bd5bd8309b4982d693a |
Intended use
- Finding the evidence for recurring contract questions like the ones it was evaluated on, with a person reviewing the output, after you have checked it on documents like yours.
- Adding evidence extraction to an existing Laya workflow, as an experimental request type.
- Research on scoring several claims per encoder pass, calibration, parameter-efficient adaptation, and where such a model stops working.
Scientific abstracts are experimental: results are reported, but on abstracts the model has not trained on, a zero-shot LLM (Claude Haiku 4.5) has the higher evidence F1 (0.628 against 0.521). Other kinds of document and claim have not been evaluated on a benchmark.
Out-of-scope uses
- Deciding whether a claim is true. The model judges a claim only against the document it is given, and it gives no document-level verdict.
- Reading an empty result as a negative.
insufficient_evidencemeans that no segment passed the threshold. It does not mean that the claim is false or that the document has no relevant evidence: evidence can be missed. - Reading one returned segment as the whole evidence. A segment can be evidence only together with other segments.
- General-purpose verification. Checking arbitrary generated answers or summaries, or claim types and domains unlike the training data, without your own evaluation on representative data.
- Decisions without review. Legal, medical, scientific or other consequential decisions. The output is an unreviewed model prediction, not a legal review of a contract.
- Trusting the scores on new domains. The thresholds and probabilities were fitted on contract and abstract dev data.
- Languages other than English. The model was trained and evaluated on English text only.
How to use
Requires Python 3.11 or later.
pip install "laya-evidence[lora] @ git+https://github.com/subhamoy-basu/laya-evidence"
from laya_evidence import EvidenceExtractor
ex = EvidenceExtractor.from_pretrained("subhamoybasu/laya-evidence-v2")
doc = ("3. The Receiving Party may disclose Confidential Information to its employees and advisers who "
"need to know it for the Purpose. 4. The Receiving Party shall not reverse engineer, decompile or "
"disassemble any samples or software provided under this Agreement.")
claims = ['Receiving Party may share some Confidential Information with some of its employees.', 'Receiving Party may independently develop information similar to Confidential Information.']
for r in ex.find(doc, claims):
print(r.claim_id, r.status, [(s.label, s.text) for s in r.spans])
from_pretrained also downloads the frozen Laya checkpoint at its pinned revision. find accepts a list
of claims (ids c1, c2, …) or a dict of id → claim, and the keyword arguments threshold (default
τ), segments (your own sorted (start, end) character spans instead of the built-in segmenter),
max_claims_per_pass, top_k and debug. find_batch takes lists of documents and claims. For every
span, document[span.start_char:span.end_char] == span.text (end_char is exclusive).
Each kept span also has evidence_probability: its evidence score calibrated on the dev data, the
probability that the segment is gold evidence. The decisions stay on the raw evidence_score, and the
calibration never reorders segments. calibration= picks the map: "abstract", "contract",
"default" (by default, the one named by threshold, else "default"), or False to leave the field
out.
LayaEvidence exposes the same model as a fourth Laya request type, next to Laya's own:
from laya_evidence import LayaEvidence
agent = LayaEvidence.from_pretrained("subhamoybasu/laya-evidence-v2")
doc = ("3. The Receiving Party may disclose Confidential Information to its employees and advisers who "
"need to know it for the Purpose. 4. The Receiving Party shall not reverse engineer, decompile or "
"disassemble any samples or software provided under this Agreement.")
out = agent.predict(doc, {
"shares": {"type": "noul", "instructions": "Does the agreement allow sharing with employees?"},
"grounding": {"type": "evidence", "claims": ['Receiving Party may share some Confidential Information with some of its employees.']},
})
print(out["answers"]["shares"])
print(out["answers"]["grounding"]["claims"]["c1"])
Interpreting the output
evidence_found: at least one segment passed the evidence-score threshold. It is the model's prediction, not a verification of the claim.insufficient_evidence: no segment passed the threshold. Relevant evidence may have been missed.supports/contradicts: the predicted direction of a returned segment. It can be wrong, and some evidence means something only together with other segments.evidence_score: the raw score the threshold applies to, P(supports) + P(contradicts). It is not a probability.evidence_probability: the evidence score calibrated on dev data, an estimate of the probability that the segment is gold evidence. It is not the probability that the claim is true or that the label is right.polarity_confidence: the share of the evidence score that the returned label holds. It is not calibrated.- Offsets:
document[start_char:end_char] == textalways holds, so a span can be located exactly. That does not make it relevant, complete or correctly labelled. - The result depends on the threshold, on the segmentation and on the other claims packed into the same encoder pass.
Training data
The weights were trained on ContractNLI, SciFact, CUAD and FEVER. Documents are split into segments, and every (claim, segment) pair gets one label: none, supports or contradicts.
| Dataset | Use | Train | Dev (model selection, τ) | Test (final evaluation only) |
|---|---|---|---|---|
| ContractNLI | training and evaluation | 423 docs / 7,191 claims | 61 docs / 1,037 claims | 123 docs / 2,091 claims |
| SciFact | training and evaluation | 491 docs / 788 claims | 74 docs / 131 claims | 283 docs / 339 claims |
| CUAD | training only (no dev or test split here) | 510 docs / 15,810 claims | — | — |
| FEVER | training only (no dev or test split here) | 75,000 docs / 75,000 claims | — | — |
Claims are counted once per document they are asked of, so every ContractNLI contract contributes its
17 hypotheses. More detail on sources, conversion and splits is in
docs/data.md.
- ContractNLI (Koreeda and Manning, 2021) asks the same 17 hypotheses of every non-disclosure agreement. The official splits are used, and each official span is one segment. Evidence spans take the (contract, hypothesis) label: Entailment → supports, Contradiction → contradicts. Every other span, and every span of a NotMentioned pair, is none.
- SciFact (Wadden et al., 2020) pairs scientific claims with research abstracts. Each abstract sentence is one segment (the title is not used). Rationale sentences take the rationale's label, and all other sentences are none. The official dev claims are the test split, and an internal dev split (about 15% of the official train claims, split by connected groups of claims and abstracts) is held out from training.
- CUAD (Hendrycks et al., 2021) annotates 510 commercial contracts with 41 clause categories. 31 of the categories became claims (for example "The contract includes a cap on liability for the breach of a party's obligation"), written by Qwen3-32B-AWQ from the category definitions and reviewed by the coding assistant. A segment overlapping an annotated clause of the category supports its claim; a contract without such a clause does not mention it. Used for training only.
- FEVER (Thorne et al., 2018) checks claims against Wikipedia introductions; the data comes from MultiVerS's pretraining files, which mark the evidence sentences. 25,000 claims of each label (supported, refuted, not enough information) were sampled. Each sentence is one segment, and the evidence sentences support or contradict the claim. Used for training only.
- Reworded hypotheses. In training, each ContractNLI hypothesis is replaced, with probability 0.5 per contract and epoch, by one of its paraphrases or its negation, whose labels are reversed. Qwen3-32B-AWQ wrote them; they were filtered with an NLI model and reviewed by the coding assistant, and the evaluation-only ones also by the author. Two paraphrases and one negation per hypothesis are kept for evaluation only.
Attribution
Training data attribution.
ContractNLI: "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Yuta Koreeda and Christopher D. Manning; licensor Hitachi America, Ltd.; https://stanfordnlp.github.io/contract-nli/; licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: document-level labels and evidence-span indices were converted to one label per span (supports, contradicts or none).
SciFact: "Fact or Fiction: Verifying Scientific Claims", David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan and Hannaneh Hajishirzi; https://github.com/allenai/scifact; claims and evidence annotations licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: rationale annotations were converted to one label per abstract sentence. Contains information from the SciFact corpus (https://github.com/allenai/scifact), whose abstracts are drawn from S2ORC (https://github.com/allenai/s2orc), which is made available under the ODC Attribution License (https://opendatacommons.org/licenses/by/1-0/).
CUAD: "CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review", Dan Hendrycks, Collin Burns, Anya Chen and Spencer Ball; licensor The Atticus Project; https://huggingface.co/datasets/theatticusproject/cuad; licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: 31 of its 41 clause categories were turned into claims, written by the generator model Qwen3-32B-AWQ from the category questions and reviewed by the coding assistant (two rewritten), and each segment overlapping an annotated clause of a category was labelled as supporting its claim.
FEVER: "FEVER: a Large-scale Dataset for Fact Extraction and VERification", James Thorne, Andreas Vlachos, Christos Christodoulopoulos and Arpit Mittal; https://fever.ai/; the annotations incorporate material from Wikipedia and are made available under the license terms of the Wikipedia articles or, where those are unavailable, CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/), and are provided without warranties; taken from the MultiVerS pretraining data (https://github.com/dwadden/multivers). Modified: 75,000 claims with their Wikipedia introductions were sampled (25,000 per label), each sentence became one segment, and evidence sentences were labelled supports or contradicts. These weights were trained on FEVER and are therefore released under CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/).
Base model: Laya by Convai Innovations (https://huggingface.co/convaiinnovations/laya), Apache-2.0, loaded at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Laya's weights are not included in this repository. This model adds trained layers on top of Laya's frozen encoder; Laya itself is not modified. Laya Evidence is an independent project, not affiliated with or endorsed by Convai Innovations. LoRA adapters on Laya's encoder (lora_adapter/) were also trained; they modify the encoder only on the evidence path.
Evaluation
All numbers are on the test splits: ContractNLI's official test set, and SciFact's official dev claims, which this project uses as its test split. They were opened only after the shipped run and its settings were frozen in configs/final_v2.lock.json; every access is logged in results/v2/final/TEST_ACCESS.log. Threshold metrics use each model's own τ (τ = 0.88 here). The first release was evaluated on the same test splits and its results set this release's goals, so these are not results on an untouched test set (docs/evaluation.md in the code repository). mAP / P@R80 are the official ContractNLI evidence-identification metrics, over the (contract, hypothesis) pairs whose gold label is not NotMentioned; ev-F1@τ is segment-level evidence F1, and ev+pol F1 also requires the right label. The first table is on the clean test subset, the test documents that do not also appear in the training or dev data; the second is on the full official test splits, which published results use.
| Model | CNLI mAP (macro) | CNLI P@R80 (macro) | CNLI mAP (micro) | CNLI P@R80 (micro) | CNLI ev-F1@τ | CNLI ev+pol F1 | CNLI false-alarm | SciFact ev-F1@τ | SciFact ev+pol F1 | Latency p50 GPU / CPU (s/doc) |
|---|---|---|---|---|---|---|---|---|---|---|
| Ours, mean ± std over 3 seeds | 0.847 ± 0.001 | 0.779 ± 0.007 | 0.830 ± 0.002 | 0.760 ± 0.005 | 0.747 ± 0.007 | 0.715 ± 0.003 | 0.148 ± 0.026 | 0.521 ± 0.046 | 0.492 ± 0.030 | — |
| Ours, shipped run (v2_joint_C_cuad_aug_fever_s13) | 0.848 [0.827, 0.876] | 0.775 [0.721, 0.823] | 0.832 | 0.759 | 0.746 [0.722, 0.770] | 0.715 [0.692, 0.739] | 0.176 [0.151, 0.201] | 0.560 [0.489, 0.622] | 0.520 [0.442, 0.594] | 0.257 / 8.64 |
| BL-1C (one claim per pass) | 0.850 | 0.783 | 0.837 | 0.757 | 0.752 | 0.722 | 0.167 | 0.559 | 0.518 | 0.708 / 17.4 |
| BL-MB (plain ModernBERT) | 0.696 | 0.519 | 0.689 | 0.511 | 0.616 | 0.589 | 0.346 | 0.208 | 0.135 | — |
| BL-NLI (zero-shot DeBERTa NLI) | 0.261 | 0.050 | 0.119 | 0.036 | 0.181 | 0.165 | 0.462 | 0.361 | 0.353 | 1.22 / 27.1 |
| Claude Haiku 4.5 (zero-shot LLM) | — | — | — | — | 0.600 | 0.565 | 0.199 | 0.628 | 0.613 | — |
- Clean test subset: the official test splits without the test documents that also appear in our training or dev data (
data/splits/test_clean.json,scripts/split_overlap.py): 109 of 123 ContractNLI test contracts are kept (dropped: word 8-gram Jaccard similarity >= 0.5 with a train or dev contract) and 102 of 283 SciFact test abstracts (dropped: abstract also in our SciFact train split or dev_int). The subset was defined for the first release (v1), after its test results, and is used unchanged: it was fixed before this release's test split was opened, and none of this release's added training data (CUAD, FEVER) overlaps a test document (results/v2/overlap_added_data.json).scripts/evaluate_clean.pycomputes these numbers from the saved test scores. Models, τ and settings are those of the official-split table. - Rows as in the official table: mean ± sample std over the seed runs v2_joint_C_cuad_aug_fever_s13, v2_joint_C_cuad_aug_fever_s14, v2_joint_C_cuad_aug_fever_s15, each at its own τ; the shipped run's row has its 95% document-level bootstrap interval on the clean subset and its latency; the baselines are single runs. BL-NLI keeps its dev-tuned τ = 0.93.
- The published Span NLI BERT results are for the full test split, so they are compared in the official-split table only. The experiments E1–E5 (
results/v2/final/summary.md) compare settings of the same model on the same documents and use the full split; E3 also has clean rows. - The rows answer different questions. BL-1C is the shipped weights at one claim per pass. BL-MB is the first release's frozen-encoder recipe on plain ModernBERT: it is not a matched control for this release's LoRA and added data (the matched comparison, plain ModernBERT with LoRA, is on dev only:
docs/training.md§3.10). BL-NLI is an off-the-shelf segment-level NLI model. BL-LLM is a prompted hosted LLM. The published Span NLI BERT numbers (official-split table) are copied from the paper and were not rerun here. - Zero-shot NLI: an existing natural language inference model (DeBERTa-v3-large, fine-tuned for NLI on MNLI, FEVER-NLI, ANLI, LingNLI and WANLI), used as it is: this project did not fine-tune it on ContractNLI or SciFact. It reads one document segment and one claim at a time and says whether the segment entails or contradicts the claim. It is a reference point, not a model trained on the same data as Laya Evidence.
- Zero-shot LLM (claude-haiku-4-5-20251001, one frozen prompt, temperature 0): one API call per document; its picks are 0/1, so it has no mAP and its decisions are at τ = 0.5. API latency over the 50 E5 contracts, one call each, measured from a laptop over the internet, so it includes network and service time (not comparable with the GPU/CPU column): p50 2.97 s, p95 4.55 s.
- Zero-shot means that this project did not fine-tune the baseline on these datasets and gave it no worked examples. It does not mean that the public benchmarks were absent from the baseline's own training: for the LLM that is unknown. The clean subset removes overlap with this project's training and dev data only.
- Released model minus the zero-shot LLM on the same documents (paired document bootstrap, 95% interval): ContractNLI evidence F1 +0.146 [+0.120, +0.171], ContractNLI false alarms -0.022 [-0.059, +0.009], SciFact evidence F1 -0.068 [-0.136, +0.007]. The intervals reflect the choice of test documents for these two sets of predictions, not retraining or repeated API calls.
Official test splits (overlapping the training data):
| Model | CNLI mAP (macro) | CNLI P@R80 (macro) | CNLI mAP (micro) | CNLI P@R80 (micro) | CNLI ev-F1@τ | CNLI ev+pol F1 | CNLI false-alarm | SciFact ev-F1@τ | SciFact ev+pol F1 | Latency p50 GPU / CPU (s/doc) |
|---|---|---|---|---|---|---|---|---|---|---|
| Ours, mean ± std over 3 seeds | 0.855 ± 0.005 | 0.788 ± 0.005 | 0.838 ± 0.005 | 0.783 ± 0.002 | 0.758 ± 0.005 | 0.729 ± 0.002 | 0.147 ± 0.025 | 0.673 ± 0.024 | 0.601 ± 0.013 | — |
| Ours, shipped run (v2_joint_C_cuad_aug_fever_s13) | 0.855 [0.837, 0.880] | 0.784 [0.744, 0.832] | 0.840 | 0.783 | 0.757 [0.736, 0.777] | 0.730 [0.706, 0.751] | 0.173 [0.151, 0.195] | 0.701 [0.659, 0.742] | 0.591 [0.538, 0.642] | 0.257 / 8.64 |
| BL-1C (one claim per pass) | 0.860 | 0.797 | 0.846 | 0.788 | 0.762 | 0.735 | 0.167 | 0.700 | 0.592 | 0.708 / 17.4 |
| BL-MB (plain ModernBERT) | 0.705 | 0.558 | 0.702 | 0.532 | 0.622 | 0.595 | 0.357 | 0.560 | 0.211 | — |
| BL-NLI (zero-shot DeBERTa NLI) | 0.253 | 0.049 | 0.118 | 0.035 | 0.178 | 0.163 | 0.466 | 0.434 | 0.422 | 1.22 / 27.1 |
| Claude Haiku 4.5 (zero-shot LLM) | — | — | — | — | 0.597 | 0.565 | 0.210 | 0.653 | 0.624 | — |
| BL-PUB: Span NLI BERT (BERT-base), published | — | — | 0.885 ± 0.025 | 0.663 ± 0.093 | — | — | — | — | — | — |
| BL-PUB: Span NLI BERT (BERT-large), published | — | — | 0.922 ± 0.006 | 0.793 ± 0.018 | — | — | — | — | — | — |
Ours: mean ± sample std (ddof=1) over the seed runs v2_joint_C_cuad_aug_fever_s13, v2_joint_C_cuad_aug_fever_s14, v2_joint_C_cuad_aug_fever_s15 of the shipped configuration, each at its own tau_default (0.88, 0.85, 0.86); max_claims_per_pass = 4. The shipped run v2_joint_C_cuad_aug_fever_s13 alone has its own row, with its 95% document-level bootstrap interval [lo, hi] (evaluate.py --bootstrap) where present and its latency; the baselines below are single runs too.
Macro = macro_label_micro_doc (our headline metric); micro = micro_label_micro_doc (the averaging of the ContractNLI paper, see BL-PUB); model selection uses macro alone. ev-F1@τ and ev+pol F1 are over (document, claim, segment) pairs; false-alarm is over (document, claim) with no gold evidence (docs/evaluation.md).
BL-1C: the shipped weights (v2_joint_C_cuad_aug_fever_s13) at max_claims_per_pass = 1, τ = 0.88.
BL-MB: ablation_modernbert_s13 at τ = 0.68; the same head over a frozen plain ModernBERT-large, trained on ContractNLI and SciFact without LoRA or the added data (the first release's recipe), carried unchanged from
configs/final.lock.jsonand not re-evaluated.BL-NLI: MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli@b3546ea6b0346eb6f8d5d68b13c7dc6d0376b3d7; τ = 0.93 maximises evidence F1 on the dev union of both sources (dev F1 0.204).
BL-PUB: Koreeda and Manning, ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts, Findings of EMNLP 2021 (arXiv:2110.01799), Table 4 (Main results), rows 'Ours (BERTbase)' and 'Ours (BERTlarge)'; the same numbers appear in Table 3 (backbone 'None' rows) (test split). The paper averages as micro_label_micro_doc.span.map, so only those columns are filled; ± is the paper's spread: Mean over the 3 of 10 hyperparameter runs with the best development macro-average NLI accuracy; '_std' is the standard deviation over those three runs (the paper's subscripts).
BL-LLM: claude-haiku-4-5-20251001 through the Anthropic API, temperature 0.0, no extended thinking, structured JSON output; prompt
configs/baselines/llm_prompt.txt(sha256 78322db9d1a8…, 2 dev revisions). One call per document with all of its claims; the model names segments as supports or contradicts, so its scores are 0/1, decisions are at τ = 0.5 and it has no mAP, P@R80 or AP. Per contract 3757 input and 437 output tokens ($0.0059 at list price as of 2026-10-04); per abstract 1097 and 39 ($0.0013). 0 of 406 calls failed and count as no evidence. API latency over the 50 E5 contracts, one call each, measured from a laptop over the internet, so it includes network and service time (not comparable with the GPU/CPU column): p50 2.97 s, p95 4.55 s.BL-MAJ (per-hypothesis majority polarity in train; polarity accuracy only): 0.886 over 2404 gold-evidence segments of ContractNLI test (0.886 restricted to doc_label ≠ none pairs, the same segments). Ours: polarity accuracy TP′/TP 0.962 ± 0.004 at each run's τ_default, over the gold-evidence segments each run also predicts, so the denominators differ. On those same segments BL-MAJ's label is right for 0.896 ± 0.004 (further analyses).
Latency: p50 seconds per document, from E5 (
results/v2/final/summary.md).E5
results/v2/final/e5_latency_test_api.jsondeparts from the E5 protocol (n_docs=50, warmup_docs=5, repeats=3): warmup_docs=1, repeats=1.Latency hardware: device=api, platform=macOS-14.4.1-arm64-arm-64bit
Latency hardware: cpu_count=40, cpu_model=unknown, device=cpu, machine=x86_64, platform=Linux-4.19.0-gvisor-x86_64-with-glibc2.36, sched_affinity_cpus=40, torch_num_interop_threads=40, torch_num_threads=40
Latency hardware: bf16_supported=True, cpu_count=17, cpu_model=unknown, cuda=13.0, device=cuda, gpu=NVIDIA L40S, gpu_total_mem_bytes=47665709056, machine=x86_64, platform=Linux-4.19.0-gvisor-x86_64-with-glibc2.36, sched_affinity_cpus=17, torch_num_interop_threads=9, torch_num_threads=1
E1: claims per encoder pass
Claims of each ContractNLI test contract are shuffled (seed 13) and packed n per pass. The shipped model packs up to 4 claims per pass, chosen on dev as the largest n whose mAP is within 0.01 of one claim per pass.
| n | mAP (macro) | P@R80 (macro) | ev-F1@τ | Forwards/doc | Tokens/doc |
|---|---|---|---|---|---|
| 1 | 0.860 | 0.797 | 0.763 | 50.5 | 42613 |
| 2 | 0.857 | 0.794 | 0.763 | 27.0 | 22968 |
| 4 | 0.856 | 0.788 | 0.757 | 15.3 | 13218 |
| 8 | 0.855 | 0.789 | 0.752 | 9.6 | 8396 |
| 17 | 0.841 | 0.763 | 0.738 | 4.0 | 3729 |
E2: interference between claims
Each ContractNLI test (contract, claim) is scored alone and with k random companion claims from the same contract (seed 13). The flip rates are the fractions of the claim's segments whose keep decision (ev ≥ τ), or whose label among segments kept both times, differs from scoring the claim alone.
| k | mean |Δev| | Decision flip rate | Polarity flip rate |
|---|---|---|---|
| 1 | 0.002 | 0.001 | 0.006 |
| 4 | 0.005 | 0.002 | 0.009 |
| 8 | 0.006 | 0.003 | 0.011 |
| 16 | 0.007 | 0.003 | 0.014 |
Almost every segment scores far below τ either way, so the decision flip rate over all segments is small. The same flips counted over the segments that matter:
| k | Segments kept alone or with companions: flipped | Gold evidence segments: flipped |
|---|---|---|
| 1 | 0.078 | 0.027 |
| 4 | 0.137 | 0.054 |
| 8 | 0.179 | 0.076 |
| 16 | 0.198 | 0.080 |
E3: unseen hypotheses
A separate model (the shipped recipe) was trained with 4 of the 17 ContractNLI hypotheses removed from its training and dev data; the held-out and the seen hypotheses are scored separately on the test split.
| Model | Hypotheses | mAP (macro) | P@R80 (macro) | mAP (micro) | P@R80 (micro) |
|---|---|---|---|---|---|
| E3 run v2_holdout_C_cuad_aug_fever_s13 (held-out hypotheses not trained on) | held-out | 0.485 | 0.331 | 0.526 | 0.281 |
| E3 run v2_holdout_C_cuad_aug_fever_s13 (held-out hypotheses not trained on) | seen | 0.863 | 0.829 | 0.848 | 0.809 |
| E3 run v2_holdout_C_cuad_aug_fever_s13, clean test subset | held-out | 0.469 | 0.302 | 0.514 | 0.265 |
| E3 run v2_holdout_C_cuad_aug_fever_s13, clean test subset | seen | 0.856 | 0.823 | 0.841 | 0.789 |
Full results: results/v2/final/summary.md.
Limitations
- Segment level only. The model returns whole segments (sentences or list items), never shorter
spans. Its output depends on the segmentation; pass
segments=to use your own. - Calibration.
evidence_scoreis a raw score, and it is overconfident.evidence_probabilityis calibrated on the dev data (Platt scaling, one map per preset). What is calibrated is the evidence event, whether the segment is gold evidence: not the supports/contradicts label, and not the truth of the claim. On the clean test subset, among segments with a raw score of at least 0.1, the expected calibration error falls from 0.238 to 0.019 on contracts and falls from 0.205 to 0.069 on abstracts (post hoc, mean over 3 seeds). That cut-off of 0.1 defines the analysis; it is not the decision threshold. The maps were fitted on contract and abstract dev data and are not guaranteed on other domains, claim types, segmentations or rates of evidence.polarity_confidenceis not calibrated. The default τ = 0.88 maximises evidence F1 on the dev data; choose your own τ for the precision and recall you need. - Spans inherit document-level labels. The segment labels are converted from coarser annotations:
every ContractNLI evidence span takes the label of its (contract, hypothesis) pair, and every sentence
of a SciFact rationale takes the rationale's label. A segment can therefore be labelled
contradictsalthough it contradicts the claim only together with other segments. - 17 fixed ContractNLI hypotheses. ContractNLI asks the same 17 hypotheses of every contract, so claims unlike them may be handled worse. In E3, a model trained without 4 of the hypotheses reached a test mAP (macro) of 0.485 on them, against 0.863 on the hypotheses it was trained on. By evidence F1 on the clean test subset it reached 0.311 on them, where a zero-shot LLM (Claude Haiku 4.5) reached 0.622.
- Domain-specific. Trained only on non-disclosure agreements (ContractNLI), scientific research
abstracts (SciFact), commercial contracts (CUAD's clause categories) and claims about Wikipedia
introductions (FEVER), English only. It works best on documents like its training data. On SciFact it
does well mainly on abstracts it was trained on: test evidence F1 is 0.765 on the 157 test abstracts
that are also in its training data, but 0.521 on the 102 that are in neither its training nor its dev
data (mean over 3 seeds), against a zero-shot NLI baseline's 0.361 there. A zero-shot LLM (Claude Haiku
4.5) reaches 0.628 there. On a fixed probe (
scripts/probe_out_of_domain.pyin the code repository; short documents written by the coding assistant) it keeps the evidence segment of 14 of the 18 claims that have one, at τ = 0.88: in-domain clauses as a bare three-clause snippet, 3 of 3; in-domain: the same three clauses inside a full NDA frame, 3 of 3; in-domain agreement (fictional): claim phrasing and polarity, 5 of 5; near-domain (biomedical abstract written for this probe), 2 of 3; out-of-domain (a two-sentence drug description), 0 of 2; out-of-domain (news-style paragraph), 1 of 2. Check it on your own data before relying on it. - Polarity errors remain. A contradicting clause can be returned as
supports. Among the gold evidence segments it returns (its true positives), its label is right for 0.958 on ContractNLI. On SciFact abstracts it has not seen it is right for 0.946 (always answering supports: 0.697, on about 59 segments per seed); on the official split, where most abstracts were seen in training, it is right for 0.893, above always answering supports (0.683). These rates are conditional: they leave out the evidence it misses and the segments it returns wrongly. The results table's ev+pol F1 counts both. - Claim interference. Claims packed into one encoder pass attend to each other, so a claim's scores
can change with the other claims in the request. In E2, with k = 16 companion claims, the keep decisions of a claim's segments flipped at a rate of 0.003 (mean |Δev| 0.007) compared with scoring the claim alone; among the segments kept alone or with the companions, 0.198 flipped. Pass
max_claims_per_pass=1to score every claim on its own, at the cost of more encoder passes.
Base model
Laya Evidence runs on Laya by Convai Innovations
(Apache-2.0): convaiinnovations/laya (root checkpoint) at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851, loaded with the laya==0.3.21 package. Laya's English
checkpoints are fine-tuned from ModernBERT-large
(Apache-2.0).
Laya's weights are not included in this repository. from_pretrained downloads them from the Hub
at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 and attaches this repository's weights with a strict state-dict load, so
later changes to the base repository do not change this model. This is a configuration C model: lora_adapter/ holds LoRA adapters (peft format) for Laya's encoder. They are applied only to evidence requests; choice, score and noul requests run with the adapters disabled, so they get Laya's own answers. The adapters are switched off and on again around those requests on the one shared encoder, so a model object must serve one call at a time: do not call it from several threads at once.
Laya's own capabilities are unchanged. On Laya's published benchmark (49 suites, 17,416 choice, score and noul questions, rebuilt with Laya's own code), this model answers exactly as vanilla Laya (accuracy):
| Benchmark | Type | Laya, published | Laya (vanilla, our run) | Laya inside Laya Evidence |
|---|---|---|---|---|
| typed-decisions, all | all three | 0.3615 | 0.3615 | 0.3615 |
| typed-decisions | choice | 0.2883 | 0.2883 | 0.2883 |
| typed-decisions | score | 0.3225 | 0.3225 | 0.3225 |
| typed-decisions | noul | 0.4867 | 0.4867 | 0.4867 |
| SST-5 | score | 0.3717 | 0.3717 | 0.3717 |
| BoolQ | noul | 0.8300 | 0.8300 | 0.8300 |
| prompt-injections | noul | 0.6983 | 0.6983 | 0.6983 |
| AG News | choice | 0.9467 | 0.9467 | 0.9467 |
| DAIR Emotion | choice | 0.5733 | 0.5733 | 0.5733 |
| MASSIVE intent, 14 languages | choice | 0.3405 | 0.3402 | 0.3402 |
| MASSIVE scenario, 14 languages | choice | 0.3036 | 0.3036 | 0.3036 |
| XNLI, 15 languages | choice | 0.5436 | 0.5438 | 0.5438 |
Laya's published numbers: typed-decisions from Laya's CPU run, the rest from its T4 run. Our runs: fp32 on an NVIDIA L40S GPU; 40 of the 49 suites match the published T4 accuracy exactly. Comparing every full answer as predict returns it (probabilities, confidence and action probability, which Laya rounds to 4 decimals) with vanilla Laya's: 17,416 of 17,416 answers identical in fp32; 17,416 of 17,416 in bf16; 17,416 of 17,416 with an evidence request before every call; 2,000 of 2,000 with an evidence question in the same call.
Details: docs/evaluation.md.
Citation
If you use this model, please cite Laya Evidence, and also ContractNLI, SciFact (and its S2ORC corpus), CUAD, FEVER, Laya and ModernBERT:
@software{basu-2026-laya-evidence,
title = "{Laya Evidence}",
author = "Basu, Subhamoy",
year = "2026",
version = "2.1.1",
url = "https://github.com/subhamoy-basu/laya-evidence",
note = "Model weights: \url{https://huggingface.co/subhamoybasu/laya-evidence-v2}. Apache License 2.0; model weights: CC BY-SA 4.0"
}
@inproceedings{koreeda-manning-2021-contractnli-dataset,
title = "{C}ontract{NLI}: A Dataset for Document-level Natural Language Inference for Contracts",
author = "Koreeda, Yuta and
Manning, Christopher",
editor = "Moens, Marie-Francine and
Huang, Xuanjing and
Specia, Lucia and
Yih, Scott Wen-tau",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
month = nov,
year = "2021",
address = "Punta Cana, Dominican Republic",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2021.findings-emnlp.164/",
doi = "10.18653/v1/2021.findings-emnlp.164",
pages = "1907--1919"
}
@inproceedings{wadden-etal-2020-fact,
title = "Fact or Fiction: Verifying Scientific Claims",
author = "Wadden, David and
Lin, Shanchuan and
Lo, Kyle and
Wang, Lucy Lu and
van Zuylen, Madeleine and
Cohan, Arman and
Hajishirzi, Hannaneh",
editor = "Webber, Bonnie and
Cohn, Trevor and
He, Yulan and
Liu, Yang",
booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)",
month = nov,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.emnlp-main.609/",
doi = "10.18653/v1/2020.emnlp-main.609",
pages = "7534--7550"
}
@inproceedings{lo-etal-2020-s2orc,
title = "{S}2{ORC}: The Semantic Scholar Open Research Corpus",
author = "Lo, Kyle and
Wang, Lucy Lu and
Neumann, Mark and
Kinney, Rodney and
Weld, Daniel",
editor = "Jurafsky, Dan and
Chai, Joyce and
Schluter, Natalie and
Tetreault, Joel",
booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.acl-main.447/",
doi = "10.18653/v1/2020.acl-main.447",
pages = "4969--4983"
}
@misc{convai-2026-laya,
title = "{Laya}",
author = "{Convai Innovations}",
year = "2026",
howpublished = "\url{https://github.com/NandhaKishorM/laya}",
note = "Python package laya 0.3.21; model weights at \url{https://huggingface.co/convaiinnovations/laya}. Apache License 2.0"
}
@misc{warner-etal-2024-modernbert,
title = "Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference",
author = {Warner, Benjamin and
Chaffin, Antoine and
Clavi{\'e}, Benjamin and
Weller, Orion and
Hallstr{\"o}m, Oskar and
Taghadouini, Said and
Gallagher, Alexis and
Biswas, Raja and
Ladhak, Faisal and
Aarsen, Tom and
Cooper, Nathan and
Adams, Griffin and
Howard, Jeremy and
Poli, Iacopo},
year = "2024",
eprint = "2412.13663",
archivePrefix = "arXiv",
primaryClass = "cs.CL",
url = "https://arxiv.org/abs/2412.13663",
note = "Also published at ACL 2025, doi:10.18653/v1/2025.acl-long.127, with a different author list"
}
@inproceedings{hendrycks-etal-2021-cuad,
title = "{CUAD}: An Expert-Annotated {NLP} Dataset for Legal Contract Review",
author = "Hendrycks, Dan and
Burns, Collin and
Chen, Anya and
Ball, Spencer",
editor = "Vanschoren, Joaquin and
Yeung, Serena",
booktitle = "Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks",
volume = "1",
year = "2021",
eprint = "2103.06268",
archivePrefix = "arXiv",
primaryClass = "cs.CL",
url = "https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/6ea9ab1baa0efb9e19094440c317e21b-Abstract-round1.html"
}
@inproceedings{thorne-etal-2018-fever,
title = "{FEVER}: a Large-scale Dataset for Fact Extraction and {VER}ification",
author = "Thorne, James and
Vlachos, Andreas and
Christodoulopoulos, Christos and
Mittal, Arpit",
editor = "Walker, Marilyn and
Ji, Heng and
Stent, Amanda",
booktitle = "Proceedings of the 2018 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)",
month = jun,
year = "2018",
address = "New Orleans, Louisiana",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/N18-1074/",
doi = "10.18653/v1/N18-1074",
pages = "809--819"
}
Model tree for subhamoybasu/laya-evidence-v2
Base model
convaiinnovations/laya