Laya Evidence

Laya Evidence finds the sentences (or list items) of a document that support or contradict each of one or more claims. It adds a fourth request type, evidence, to Laya, next to choice, score and noul, and packs the claims into one encoder input per document window.

It is a specialist. Its strongest evaluated setting is the contract questions it was trained on (ContractNLI's 17 hypothesis types), asked of contracts it has not seen. On claim types and kinds of document unlike its training data it is much weaker (see Limitations). It is not a general-purpose fact checker.

Summary

For each claim the model returns a status (evidence_found or insufficient_evidence) and the segments it keeps. Each segment comes with its character offsets in the document, its exact text, a label (supports or contradicts), an evidence score and a polarity confidence.

Each encoder row is [CLS] head [SEP] [MASK] claim 1 … [MASK] claim n [SEP] document window [SEP]. The row runs through Laya's encoder; Laya's head layers are not used on this path. The rows are centred, a trained evidence type vector marking the request type is added to every token vector, and the rows are layer-normalised and passed through 2 adapter layers. An evidence head then scores every (claim, segment) pair as none, supports or contradicts, from the claim's [MASK] vector and the segment's token vectors. A segment is kept when its evidence score, P(supports) + P(contradicts), is at least the threshold τ.

This repository holds only the parameters Laya Evidence trains: the evidence type vector, a LayerNorm, 2 adapter layers and the evidence head, plus LoRA adapters on Laya's encoder (lora_adapter/). The Laya checkpoint is downloaded from the Hub at a pinned revision when the model is loaded (see Base model).

Developed by Subhamoy Basu (subhamoy.basu@berkeley.edu)
Code https://github.com/subhamoy-basu/laya-evidence
License CC BY-SA 4.0 (see LICENSE and NOTICE): trained on FEVER (CC BY-SA 3.0); the code is Apache-2.0
Base checkpoint convaiinnovations/laya (root checkpoint) at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
Configuration C: Laya's encoder gets trained LoRA adapters on the evidence path; its base weights and head layers stay frozen
Encoder budget 1,024 tokens per row
Claims per encoder pass up to 4, within 448 claim tokens
Default threshold τ 0.88, the value that maximises evidence F1 on the dev data of the training datasets (almost all of it ContractNLI's)
Segmenter pysbd 0.3.4 (or the caller's own segments)
Training run v2_joint_C_cuad_aug_fever_s13, seed 13, code at commit cb765a989d4ca93787afd6f12de7802279ac9501
Weights evidence.safetensors, sha256 ffe1bf2a25f7e241331dbc1c79f99b319d2b682b5d2e50d52c26c33ea7782fc0; LoRA adapter lora_adapter/adapter_model.safetensors, sha256 97d614501bd7b253bc3c5e8a50952914920aa582f79f0bd5bd8309b4982d693a

Intended use

  • Finding the evidence for recurring contract questions like the ones it was evaluated on, with a person reviewing the output, after you have checked it on documents like yours.
  • Adding evidence extraction to an existing Laya workflow, as an experimental request type.
  • Research on scoring several claims per encoder pass, calibration, parameter-efficient adaptation, and where such a model stops working.

Scientific abstracts are experimental: results are reported, but on abstracts the model has not trained on, a zero-shot LLM (Claude Haiku 4.5) has the higher evidence F1 (0.628 against 0.521). Other kinds of document and claim have not been evaluated on a benchmark.

Out-of-scope uses

  • Deciding whether a claim is true. The model judges a claim only against the document it is given, and it gives no document-level verdict.
  • Reading an empty result as a negative. insufficient_evidence means that no segment passed the threshold. It does not mean that the claim is false or that the document has no relevant evidence: evidence can be missed.
  • Reading one returned segment as the whole evidence. A segment can be evidence only together with other segments.
  • General-purpose verification. Checking arbitrary generated answers or summaries, or claim types and domains unlike the training data, without your own evaluation on representative data.
  • Decisions without review. Legal, medical, scientific or other consequential decisions. The output is an unreviewed model prediction, not a legal review of a contract.
  • Trusting the scores on new domains. The thresholds and probabilities were fitted on contract and abstract dev data.
  • Languages other than English. The model was trained and evaluated on English text only.

How to use

Requires Python 3.11 or later.

pip install "laya-evidence[lora] @ git+https://github.com/subhamoy-basu/laya-evidence"
from laya_evidence import EvidenceExtractor
ex = EvidenceExtractor.from_pretrained("subhamoybasu/laya-evidence-v2")
doc = ("3. The Receiving Party may disclose Confidential Information to its employees and advisers who "
       "need to know it for the Purpose. 4. The Receiving Party shall not reverse engineer, decompile or "
       "disassemble any samples or software provided under this Agreement.")
claims = ['Receiving Party may share some Confidential Information with some of its employees.', 'Receiving Party may independently develop information similar to Confidential Information.']
for r in ex.find(doc, claims):
    print(r.claim_id, r.status, [(s.label, s.text) for s in r.spans])

from_pretrained also downloads the frozen Laya checkpoint at its pinned revision. find accepts a list of claims (ids c1, c2, …) or a dict of id → claim, and the keyword arguments threshold (default τ), segments (your own sorted (start, end) character spans instead of the built-in segmenter), max_claims_per_pass, top_k and debug. find_batch takes lists of documents and claims. For every span, document[span.start_char:span.end_char] == span.text (end_char is exclusive).

Each kept span also has evidence_probability: its evidence score calibrated on the dev data, the probability that the segment is gold evidence. The decisions stay on the raw evidence_score, and the calibration never reorders segments. calibration= picks the map: "abstract", "contract", "default" (by default, the one named by threshold, else "default"), or False to leave the field out.

LayaEvidence exposes the same model as a fourth Laya request type, next to Laya's own:

from laya_evidence import LayaEvidence
agent = LayaEvidence.from_pretrained("subhamoybasu/laya-evidence-v2")
doc = ("3. The Receiving Party may disclose Confidential Information to its employees and advisers who "
       "need to know it for the Purpose. 4. The Receiving Party shall not reverse engineer, decompile or "
       "disassemble any samples or software provided under this Agreement.")
out = agent.predict(doc, {
    "shares": {"type": "noul", "instructions": "Does the agreement allow sharing with employees?"},
    "grounding": {"type": "evidence", "claims": ['Receiving Party may share some Confidential Information with some of its employees.']},
})
print(out["answers"]["shares"])
print(out["answers"]["grounding"]["claims"]["c1"])

Interpreting the output

  • evidence_found: at least one segment passed the evidence-score threshold. It is the model's prediction, not a verification of the claim.
  • insufficient_evidence: no segment passed the threshold. Relevant evidence may have been missed.
  • supports / contradicts: the predicted direction of a returned segment. It can be wrong, and some evidence means something only together with other segments.
  • evidence_score: the raw score the threshold applies to, P(supports) + P(contradicts). It is not a probability.
  • evidence_probability: the evidence score calibrated on dev data, an estimate of the probability that the segment is gold evidence. It is not the probability that the claim is true or that the label is right.
  • polarity_confidence: the share of the evidence score that the returned label holds. It is not calibrated.
  • Offsets: document[start_char:end_char] == text always holds, so a span can be located exactly. That does not make it relevant, complete or correctly labelled.
  • The result depends on the threshold, on the segmentation and on the other claims packed into the same encoder pass.

Training data

The weights were trained on ContractNLI, SciFact, CUAD and FEVER. Documents are split into segments, and every (claim, segment) pair gets one label: none, supports or contradicts.

Dataset Use Train Dev (model selection, τ) Test (final evaluation only)
ContractNLI training and evaluation 423 docs / 7,191 claims 61 docs / 1,037 claims 123 docs / 2,091 claims
SciFact training and evaluation 491 docs / 788 claims 74 docs / 131 claims 283 docs / 339 claims
CUAD training only (no dev or test split here) 510 docs / 15,810 claims — —
FEVER training only (no dev or test split here) 75,000 docs / 75,000 claims — —

Claims are counted once per document they are asked of, so every ContractNLI contract contributes its 17 hypotheses. More detail on sources, conversion and splits is in docs/data.md.

  • ContractNLI (Koreeda and Manning, 2021) asks the same 17 hypotheses of every non-disclosure agreement. The official splits are used, and each official span is one segment. Evidence spans take the (contract, hypothesis) label: Entailment → supports, Contradiction → contradicts. Every other span, and every span of a NotMentioned pair, is none.
  • SciFact (Wadden et al., 2020) pairs scientific claims with research abstracts. Each abstract sentence is one segment (the title is not used). Rationale sentences take the rationale's label, and all other sentences are none. The official dev claims are the test split, and an internal dev split (about 15% of the official train claims, split by connected groups of claims and abstracts) is held out from training.
  • CUAD (Hendrycks et al., 2021) annotates 510 commercial contracts with 41 clause categories. 31 of the categories became claims (for example "The contract includes a cap on liability for the breach of a party's obligation"), written by Qwen3-32B-AWQ from the category definitions and reviewed by the coding assistant. A segment overlapping an annotated clause of the category supports its claim; a contract without such a clause does not mention it. Used for training only.
  • FEVER (Thorne et al., 2018) checks claims against Wikipedia introductions; the data comes from MultiVerS's pretraining files, which mark the evidence sentences. 25,000 claims of each label (supported, refuted, not enough information) were sampled. Each sentence is one segment, and the evidence sentences support or contradict the claim. Used for training only.
  • Reworded hypotheses. In training, each ContractNLI hypothesis is replaced, with probability 0.5 per contract and epoch, by one of its paraphrases or its negation, whose labels are reversed. Qwen3-32B-AWQ wrote them; they were filtered with an NLI model and reviewed by the coding assistant, and the evaluation-only ones also by the author. Two paraphrases and one negation per hypothesis are kept for evaluation only.

Attribution

Training data attribution.

ContractNLI: "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Yuta Koreeda and Christopher D. Manning; licensor Hitachi America, Ltd.; https://stanfordnlp.github.io/contract-nli/; licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: document-level labels and evidence-span indices were converted to one label per span (supports, contradicts or none).

SciFact: "Fact or Fiction: Verifying Scientific Claims", David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan and Hannaneh Hajishirzi; https://github.com/allenai/scifact; claims and evidence annotations licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: rationale annotations were converted to one label per abstract sentence. Contains information from the SciFact corpus (https://github.com/allenai/scifact), whose abstracts are drawn from S2ORC (https://github.com/allenai/s2orc), which is made available under the ODC Attribution License (https://opendatacommons.org/licenses/by/1-0/).

CUAD: "CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review", Dan Hendrycks, Collin Burns, Anya Chen and Spencer Ball; licensor The Atticus Project; https://huggingface.co/datasets/theatticusproject/cuad; licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: 31 of its 41 clause categories were turned into claims, written by the generator model Qwen3-32B-AWQ from the category questions and reviewed by the coding assistant (two rewritten), and each segment overlapping an annotated clause of a category was labelled as supporting its claim.

FEVER: "FEVER: a Large-scale Dataset for Fact Extraction and VERification", James Thorne, Andreas Vlachos, Christos Christodoulopoulos and Arpit Mittal; https://fever.ai/; the annotations incorporate material from Wikipedia and are made available under the license terms of the Wikipedia articles or, where those are unavailable, CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/), and are provided without warranties; taken from the MultiVerS pretraining data (https://github.com/dwadden/multivers). Modified: 75,000 claims with their Wikipedia introductions were sampled (25,000 per label), each sentence became one segment, and evidence sentences were labelled supports or contradicts. These weights were trained on FEVER and are therefore released under CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/).

Base model: Laya by Convai Innovations (https://huggingface.co/convaiinnovations/laya), Apache-2.0, loaded at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Laya's weights are not included in this repository. This model adds trained layers on top of Laya's frozen encoder; Laya itself is not modified. Laya Evidence is an independent project, not affiliated with or endorsed by Convai Innovations. LoRA adapters on Laya's encoder (lora_adapter/) were also trained; they modify the encoder only on the evidence path.

Evaluation

All numbers are on the test splits: ContractNLI's official test set, and SciFact's official dev claims, which this project uses as its test split. They were opened only after the shipped run and its settings were frozen in configs/final_v2.lock.json; every access is logged in results/v2/final/TEST_ACCESS.log. Threshold metrics use each model's own τ (τ = 0.88 here). The first release was evaluated on the same test splits and its results set this release's goals, so these are not results on an untouched test set (docs/evaluation.md in the code repository). mAP / P@R80 are the official ContractNLI evidence-identification metrics, over the (contract, hypothesis) pairs whose gold label is not NotMentioned; ev-F1@τ is segment-level evidence F1, and ev+pol F1 also requires the right label. The first table is on the clean test subset, the test documents that do not also appear in the training or dev data; the second is on the full official test splits, which published results use.

Model CNLI mAP (macro) CNLI P@R80 (macro) CNLI mAP (micro) CNLI P@R80 (micro) CNLI ev-F1@τ CNLI ev+pol F1 CNLI false-alarm SciFact ev-F1@τ SciFact ev+pol F1 Latency p50 GPU / CPU (s/doc)
Ours, mean ± std over 3 seeds 0.847 ± 0.001 0.779 ± 0.007 0.830 ± 0.002 0.760 ± 0.005 0.747 ± 0.007 0.715 ± 0.003 0.148 ± 0.026 0.521 ± 0.046 0.492 ± 0.030 —
Ours, shipped run (v2_joint_C_cuad_aug_fever_s13) 0.848 [0.827, 0.876] 0.775 [0.721, 0.823] 0.832 0.759 0.746 [0.722, 0.770] 0.715 [0.692, 0.739] 0.176 [0.151, 0.201] 0.560 [0.489, 0.622] 0.520 [0.442, 0.594] 0.257 / 8.64
BL-1C (one claim per pass) 0.850 0.783 0.837 0.757 0.752 0.722 0.167 0.559 0.518 0.708 / 17.4
BL-MB (plain ModernBERT) 0.696 0.519 0.689 0.511 0.616 0.589 0.346 0.208 0.135 —
BL-NLI (zero-shot DeBERTa NLI) 0.261 0.050 0.119 0.036 0.181 0.165 0.462 0.361 0.353 1.22 / 27.1
Claude Haiku 4.5 (zero-shot LLM) — — — — 0.600 0.565 0.199 0.628 0.613 —
  • Clean test subset: the official test splits without the test documents that also appear in our training or dev data (data/splits/test_clean.json, scripts/split_overlap.py): 109 of 123 ContractNLI test contracts are kept (dropped: word 8-gram Jaccard similarity >= 0.5 with a train or dev contract) and 102 of 283 SciFact test abstracts (dropped: abstract also in our SciFact train split or dev_int). The subset was defined for the first release (v1), after its test results, and is used unchanged: it was fixed before this release's test split was opened, and none of this release's added training data (CUAD, FEVER) overlaps a test document (results/v2/overlap_added_data.json). scripts/evaluate_clean.py computes these numbers from the saved test scores. Models, τ and settings are those of the official-split table.
  • Rows as in the official table: mean ± sample std over the seed runs v2_joint_C_cuad_aug_fever_s13, v2_joint_C_cuad_aug_fever_s14, v2_joint_C_cuad_aug_fever_s15, each at its own τ; the shipped run's row has its 95% document-level bootstrap interval on the clean subset and its latency; the baselines are single runs. BL-NLI keeps its dev-tuned τ = 0.93.
  • The published Span NLI BERT results are for the full test split, so they are compared in the official-split table only. The experiments E1–E5 (results/v2/final/summary.md) compare settings of the same model on the same documents and use the full split; E3 also has clean rows.
  • The rows answer different questions. BL-1C is the shipped weights at one claim per pass. BL-MB is the first release's frozen-encoder recipe on plain ModernBERT: it is not a matched control for this release's LoRA and added data (the matched comparison, plain ModernBERT with LoRA, is on dev only: docs/training.md §3.10). BL-NLI is an off-the-shelf segment-level NLI model. BL-LLM is a prompted hosted LLM. The published Span NLI BERT numbers (official-split table) are copied from the paper and were not rerun here.
  • Zero-shot NLI: an existing natural language inference model (DeBERTa-v3-large, fine-tuned for NLI on MNLI, FEVER-NLI, ANLI, LingNLI and WANLI), used as it is: this project did not fine-tune it on ContractNLI or SciFact. It reads one document segment and one claim at a time and says whether the segment entails or contradicts the claim. It is a reference point, not a model trained on the same data as Laya Evidence.
  • Zero-shot LLM (claude-haiku-4-5-20251001, one frozen prompt, temperature 0): one API call per document; its picks are 0/1, so it has no mAP and its decisions are at τ = 0.5. API latency over the 50 E5 contracts, one call each, measured from a laptop over the internet, so it includes network and service time (not comparable with the GPU/CPU column): p50 2.97 s, p95 4.55 s.
  • Zero-shot means that this project did not fine-tune the baseline on these datasets and gave it no worked examples. It does not mean that the public benchmarks were absent from the baseline's own training: for the LLM that is unknown. The clean subset removes overlap with this project's training and dev data only.
  • Released model minus the zero-shot LLM on the same documents (paired document bootstrap, 95% interval): ContractNLI evidence F1 +0.146 [+0.120, +0.171], ContractNLI false alarms -0.022 [-0.059, +0.009], SciFact evidence F1 -0.068 [-0.136, +0.007]. The intervals reflect the choice of test documents for these two sets of predictions, not retraining or repeated API calls.

Official test splits (overlapping the training data):

Model CNLI mAP (macro) CNLI P@R80 (macro) CNLI mAP (micro) CNLI P@R80 (micro) CNLI ev-F1@τ CNLI ev+pol F1 CNLI false-alarm SciFact ev-F1@τ SciFact ev+pol F1 Latency p50 GPU / CPU (s/doc)
Ours, mean ± std over 3 seeds 0.855 ± 0.005 0.788 ± 0.005 0.838 ± 0.005 0.783 ± 0.002 0.758 ± 0.005 0.729 ± 0.002 0.147 ± 0.025 0.673 ± 0.024 0.601 ± 0.013 —
Ours, shipped run (v2_joint_C_cuad_aug_fever_s13) 0.855 [0.837, 0.880] 0.784 [0.744, 0.832] 0.840 0.783 0.757 [0.736, 0.777] 0.730 [0.706, 0.751] 0.173 [0.151, 0.195] 0.701 [0.659, 0.742] 0.591 [0.538, 0.642] 0.257 / 8.64
BL-1C (one claim per pass) 0.860 0.797 0.846 0.788 0.762 0.735 0.167 0.700 0.592 0.708 / 17.4
BL-MB (plain ModernBERT) 0.705 0.558 0.702 0.532 0.622 0.595 0.357 0.560 0.211 —
BL-NLI (zero-shot DeBERTa NLI) 0.253 0.049 0.118 0.035 0.178 0.163 0.466 0.434 0.422 1.22 / 27.1
Claude Haiku 4.5 (zero-shot LLM) — — — — 0.597 0.565 0.210 0.653 0.624 —
BL-PUB: Span NLI BERT (BERT-base), published — — 0.885 ± 0.025 0.663 ± 0.093 — — — — — —
BL-PUB: Span NLI BERT (BERT-large), published — — 0.922 ± 0.006 0.793 ± 0.018 — — — — — —
  • Ours: mean ± sample std (ddof=1) over the seed runs v2_joint_C_cuad_aug_fever_s13, v2_joint_C_cuad_aug_fever_s14, v2_joint_C_cuad_aug_fever_s15 of the shipped configuration, each at its own tau_default (0.88, 0.85, 0.86); max_claims_per_pass = 4. The shipped run v2_joint_C_cuad_aug_fever_s13 alone has its own row, with its 95% document-level bootstrap interval [lo, hi] (evaluate.py --bootstrap) where present and its latency; the baselines below are single runs too.

  • Macro = macro_label_micro_doc (our headline metric); micro = micro_label_micro_doc (the averaging of the ContractNLI paper, see BL-PUB); model selection uses macro alone. ev-F1@τ and ev+pol F1 are over (document, claim, segment) pairs; false-alarm is over (document, claim) with no gold evidence (docs/evaluation.md).

  • BL-1C: the shipped weights (v2_joint_C_cuad_aug_fever_s13) at max_claims_per_pass = 1, τ = 0.88.

  • BL-MB: ablation_modernbert_s13 at τ = 0.68; the same head over a frozen plain ModernBERT-large, trained on ContractNLI and SciFact without LoRA or the added data (the first release's recipe), carried unchanged from configs/final.lock.json and not re-evaluated.

  • BL-NLI: MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli@b3546ea6b0346eb6f8d5d68b13c7dc6d0376b3d7; τ = 0.93 maximises evidence F1 on the dev union of both sources (dev F1 0.204).

  • BL-PUB: Koreeda and Manning, ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts, Findings of EMNLP 2021 (arXiv:2110.01799), Table 4 (Main results), rows 'Ours (BERTbase)' and 'Ours (BERTlarge)'; the same numbers appear in Table 3 (backbone 'None' rows) (test split). The paper averages as micro_label_micro_doc.span.map, so only those columns are filled; ± is the paper's spread: Mean over the 3 of 10 hyperparameter runs with the best development macro-average NLI accuracy; '_std' is the standard deviation over those three runs (the paper's subscripts).

  • BL-LLM: claude-haiku-4-5-20251001 through the Anthropic API, temperature 0.0, no extended thinking, structured JSON output; prompt configs/baselines/llm_prompt.txt (sha256 78322db9d1a8…, 2 dev revisions). One call per document with all of its claims; the model names segments as supports or contradicts, so its scores are 0/1, decisions are at τ = 0.5 and it has no mAP, P@R80 or AP. Per contract 3757 input and 437 output tokens ($0.0059 at list price as of 2026-10-04); per abstract 1097 and 39 ($0.0013). 0 of 406 calls failed and count as no evidence. API latency over the 50 E5 contracts, one call each, measured from a laptop over the internet, so it includes network and service time (not comparable with the GPU/CPU column): p50 2.97 s, p95 4.55 s.

  • BL-MAJ (per-hypothesis majority polarity in train; polarity accuracy only): 0.886 over 2404 gold-evidence segments of ContractNLI test (0.886 restricted to doc_label ≠ none pairs, the same segments). Ours: polarity accuracy TP′/TP 0.962 ± 0.004 at each run's τ_default, over the gold-evidence segments each run also predicts, so the denominators differ. On those same segments BL-MAJ's label is right for 0.896 ± 0.004 (further analyses).

  • Latency: p50 seconds per document, from E5 (results/v2/final/summary.md).

  • E5 results/v2/final/e5_latency_test_api.json departs from the E5 protocol (n_docs=50, warmup_docs=5, repeats=3): warmup_docs=1, repeats=1.

  • Latency hardware: device=api, platform=macOS-14.4.1-arm64-arm-64bit

  • Latency hardware: cpu_count=40, cpu_model=unknown, device=cpu, machine=x86_64, platform=Linux-4.19.0-gvisor-x86_64-with-glibc2.36, sched_affinity_cpus=40, torch_num_interop_threads=40, torch_num_threads=40

  • Latency hardware: bf16_supported=True, cpu_count=17, cpu_model=unknown, cuda=13.0, device=cuda, gpu=NVIDIA L40S, gpu_total_mem_bytes=47665709056, machine=x86_64, platform=Linux-4.19.0-gvisor-x86_64-with-glibc2.36, sched_affinity_cpus=17, torch_num_interop_threads=9, torch_num_threads=1

E1: claims per encoder pass

Claims of each ContractNLI test contract are shuffled (seed 13) and packed n per pass. The shipped model packs up to 4 claims per pass, chosen on dev as the largest n whose mAP is within 0.01 of one claim per pass.

n mAP (macro) P@R80 (macro) ev-F1@τ Forwards/doc Tokens/doc
1 0.860 0.797 0.763 50.5 42613
2 0.857 0.794 0.763 27.0 22968
4 0.856 0.788 0.757 15.3 13218
8 0.855 0.789 0.752 9.6 8396
17 0.841 0.763 0.738 4.0 3729

E2: interference between claims

Each ContractNLI test (contract, claim) is scored alone and with k random companion claims from the same contract (seed 13). The flip rates are the fractions of the claim's segments whose keep decision (ev ≥ τ), or whose label among segments kept both times, differs from scoring the claim alone.

k mean |Δev| Decision flip rate Polarity flip rate
1 0.002 0.001 0.006
4 0.005 0.002 0.009
8 0.006 0.003 0.011
16 0.007 0.003 0.014

Almost every segment scores far below τ either way, so the decision flip rate over all segments is small. The same flips counted over the segments that matter:

k Segments kept alone or with companions: flipped Gold evidence segments: flipped
1 0.078 0.027
4 0.137 0.054
8 0.179 0.076
16 0.198 0.080

E3: unseen hypotheses

A separate model (the shipped recipe) was trained with 4 of the 17 ContractNLI hypotheses removed from its training and dev data; the held-out and the seen hypotheses are scored separately on the test split.

Model Hypotheses mAP (macro) P@R80 (macro) mAP (micro) P@R80 (micro)
E3 run v2_holdout_C_cuad_aug_fever_s13 (held-out hypotheses not trained on) held-out 0.485 0.331 0.526 0.281
E3 run v2_holdout_C_cuad_aug_fever_s13 (held-out hypotheses not trained on) seen 0.863 0.829 0.848 0.809
E3 run v2_holdout_C_cuad_aug_fever_s13, clean test subset held-out 0.469 0.302 0.514 0.265
E3 run v2_holdout_C_cuad_aug_fever_s13, clean test subset seen 0.856 0.823 0.841 0.789

Full results: results/v2/final/summary.md.

Limitations

  • Segment level only. The model returns whole segments (sentences or list items), never shorter spans. Its output depends on the segmentation; pass segments= to use your own.
  • Calibration. evidence_score is a raw score, and it is overconfident. evidence_probability is calibrated on the dev data (Platt scaling, one map per preset). What is calibrated is the evidence event, whether the segment is gold evidence: not the supports/contradicts label, and not the truth of the claim. On the clean test subset, among segments with a raw score of at least 0.1, the expected calibration error falls from 0.238 to 0.019 on contracts and falls from 0.205 to 0.069 on abstracts (post hoc, mean over 3 seeds). That cut-off of 0.1 defines the analysis; it is not the decision threshold. The maps were fitted on contract and abstract dev data and are not guaranteed on other domains, claim types, segmentations or rates of evidence. polarity_confidence is not calibrated. The default τ = 0.88 maximises evidence F1 on the dev data; choose your own τ for the precision and recall you need.
  • Spans inherit document-level labels. The segment labels are converted from coarser annotations: every ContractNLI evidence span takes the label of its (contract, hypothesis) pair, and every sentence of a SciFact rationale takes the rationale's label. A segment can therefore be labelled contradicts although it contradicts the claim only together with other segments.
  • 17 fixed ContractNLI hypotheses. ContractNLI asks the same 17 hypotheses of every contract, so claims unlike them may be handled worse. In E3, a model trained without 4 of the hypotheses reached a test mAP (macro) of 0.485 on them, against 0.863 on the hypotheses it was trained on. By evidence F1 on the clean test subset it reached 0.311 on them, where a zero-shot LLM (Claude Haiku 4.5) reached 0.622.
  • Domain-specific. Trained only on non-disclosure agreements (ContractNLI), scientific research abstracts (SciFact), commercial contracts (CUAD's clause categories) and claims about Wikipedia introductions (FEVER), English only. It works best on documents like its training data. On SciFact it does well mainly on abstracts it was trained on: test evidence F1 is 0.765 on the 157 test abstracts that are also in its training data, but 0.521 on the 102 that are in neither its training nor its dev data (mean over 3 seeds), against a zero-shot NLI baseline's 0.361 there. A zero-shot LLM (Claude Haiku 4.5) reaches 0.628 there. On a fixed probe (scripts/probe_out_of_domain.py in the code repository; short documents written by the coding assistant) it keeps the evidence segment of 14 of the 18 claims that have one, at τ = 0.88: in-domain clauses as a bare three-clause snippet, 3 of 3; in-domain: the same three clauses inside a full NDA frame, 3 of 3; in-domain agreement (fictional): claim phrasing and polarity, 5 of 5; near-domain (biomedical abstract written for this probe), 2 of 3; out-of-domain (a two-sentence drug description), 0 of 2; out-of-domain (news-style paragraph), 1 of 2. Check it on your own data before relying on it.
  • Polarity errors remain. A contradicting clause can be returned as supports. Among the gold evidence segments it returns (its true positives), its label is right for 0.958 on ContractNLI. On SciFact abstracts it has not seen it is right for 0.946 (always answering supports: 0.697, on about 59 segments per seed); on the official split, where most abstracts were seen in training, it is right for 0.893, above always answering supports (0.683). These rates are conditional: they leave out the evidence it misses and the segments it returns wrongly. The results table's ev+pol F1 counts both.
  • Claim interference. Claims packed into one encoder pass attend to each other, so a claim's scores can change with the other claims in the request. In E2, with k = 16 companion claims, the keep decisions of a claim's segments flipped at a rate of 0.003 (mean |Δev| 0.007) compared with scoring the claim alone; among the segments kept alone or with the companions, 0.198 flipped. Pass max_claims_per_pass=1 to score every claim on its own, at the cost of more encoder passes.

Base model

Laya Evidence runs on Laya by Convai Innovations (Apache-2.0): convaiinnovations/laya (root checkpoint) at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851, loaded with the laya==0.3.21 package. Laya's English checkpoints are fine-tuned from ModernBERT-large (Apache-2.0).

Laya's weights are not included in this repository. from_pretrained downloads them from the Hub at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 and attaches this repository's weights with a strict state-dict load, so later changes to the base repository do not change this model. This is a configuration C model: lora_adapter/ holds LoRA adapters (peft format) for Laya's encoder. They are applied only to evidence requests; choice, score and noul requests run with the adapters disabled, so they get Laya's own answers. The adapters are switched off and on again around those requests on the one shared encoder, so a model object must serve one call at a time: do not call it from several threads at once.

Laya's own capabilities are unchanged. On Laya's published benchmark (49 suites, 17,416 choice, score and noul questions, rebuilt with Laya's own code), this model answers exactly as vanilla Laya (accuracy):

Benchmark Type Laya, published Laya (vanilla, our run) Laya inside Laya Evidence
typed-decisions, all all three 0.3615 0.3615 0.3615
typed-decisions choice 0.2883 0.2883 0.2883
typed-decisions score 0.3225 0.3225 0.3225
typed-decisions noul 0.4867 0.4867 0.4867
SST-5 score 0.3717 0.3717 0.3717
BoolQ noul 0.8300 0.8300 0.8300
prompt-injections noul 0.6983 0.6983 0.6983
AG News choice 0.9467 0.9467 0.9467
DAIR Emotion choice 0.5733 0.5733 0.5733
MASSIVE intent, 14 languages choice 0.3405 0.3402 0.3402
MASSIVE scenario, 14 languages choice 0.3036 0.3036 0.3036
XNLI, 15 languages choice 0.5436 0.5438 0.5438

Laya's published numbers: typed-decisions from Laya's CPU run, the rest from its T4 run. Our runs: fp32 on an NVIDIA L40S GPU; 40 of the 49 suites match the published T4 accuracy exactly. Comparing every full answer as predict returns it (probabilities, confidence and action probability, which Laya rounds to 4 decimals) with vanilla Laya's: 17,416 of 17,416 answers identical in fp32; 17,416 of 17,416 in bf16; 17,416 of 17,416 with an evidence request before every call; 2,000 of 2,000 with an evidence question in the same call.

Details: docs/evaluation.md.

Citation

If you use this model, please cite Laya Evidence, and also ContractNLI, SciFact (and its S2ORC corpus), CUAD, FEVER, Laya and ModernBERT:

@software{basu-2026-laya-evidence,
    title = "{Laya Evidence}",
    author = "Basu, Subhamoy",
    year = "2026",
    version = "2.1.1",
    url = "https://github.com/subhamoy-basu/laya-evidence",
    note = "Model weights: \url{https://huggingface.co/subhamoybasu/laya-evidence-v2}. Apache License 2.0; model weights: CC BY-SA 4.0"
}

@inproceedings{koreeda-manning-2021-contractnli-dataset,
    title = "{C}ontract{NLI}: A Dataset for Document-level Natural Language Inference for Contracts",
    author = "Koreeda, Yuta  and
      Manning, Christopher",
    editor = "Moens, Marie-Francine  and
      Huang, Xuanjing  and
      Specia, Lucia  and
      Yih, Scott Wen-tau",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
    month = nov,
    year = "2021",
    address = "Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.findings-emnlp.164/",
    doi = "10.18653/v1/2021.findings-emnlp.164",
    pages = "1907--1919"
}

@inproceedings{wadden-etal-2020-fact,
    title = "Fact or Fiction: Verifying Scientific Claims",
    author = "Wadden, David  and
      Lin, Shanchuan  and
      Lo, Kyle  and
      Wang, Lucy Lu  and
      van Zuylen, Madeleine  and
      Cohan, Arman  and
      Hajishirzi, Hannaneh",
    editor = "Webber, Bonnie  and
      Cohn, Trevor  and
      He, Yulan  and
      Liu, Yang",
    booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)",
    month = nov,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2020.emnlp-main.609/",
    doi = "10.18653/v1/2020.emnlp-main.609",
    pages = "7534--7550"
}

@inproceedings{lo-etal-2020-s2orc,
    title = "{S}2{ORC}: The Semantic Scholar Open Research Corpus",
    author = "Lo, Kyle  and
      Wang, Lucy Lu  and
      Neumann, Mark  and
      Kinney, Rodney  and
      Weld, Daniel",
    editor = "Jurafsky, Dan  and
      Chai, Joyce  and
      Schluter, Natalie  and
      Tetreault, Joel",
    booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
    month = jul,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2020.acl-main.447/",
    doi = "10.18653/v1/2020.acl-main.447",
    pages = "4969--4983"
}

@misc{convai-2026-laya,
    title = "{Laya}",
    author = "{Convai Innovations}",
    year = "2026",
    howpublished = "\url{https://github.com/NandhaKishorM/laya}",
    note = "Python package laya 0.3.21; model weights at \url{https://huggingface.co/convaiinnovations/laya}. Apache License 2.0"
}

@misc{warner-etal-2024-modernbert,
    title = "Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference",
    author = {Warner, Benjamin  and
      Chaffin, Antoine  and
      Clavi{\'e}, Benjamin  and
      Weller, Orion  and
      Hallstr{\"o}m, Oskar  and
      Taghadouini, Said  and
      Gallagher, Alexis  and
      Biswas, Raja  and
      Ladhak, Faisal  and
      Aarsen, Tom  and
      Cooper, Nathan  and
      Adams, Griffin  and
      Howard, Jeremy  and
      Poli, Iacopo},
    year = "2024",
    eprint = "2412.13663",
    archivePrefix = "arXiv",
    primaryClass = "cs.CL",
    url = "https://arxiv.org/abs/2412.13663",
    note = "Also published at ACL 2025, doi:10.18653/v1/2025.acl-long.127, with a different author list"
}

@inproceedings{hendrycks-etal-2021-cuad,
    title = "{CUAD}: An Expert-Annotated {NLP} Dataset for Legal Contract Review",
    author = "Hendrycks, Dan  and
      Burns, Collin  and
      Chen, Anya  and
      Ball, Spencer",
    editor = "Vanschoren, Joaquin  and
      Yeung, Serena",
    booktitle = "Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks",
    volume = "1",
    year = "2021",
    eprint = "2103.06268",
    archivePrefix = "arXiv",
    primaryClass = "cs.CL",
    url = "https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/6ea9ab1baa0efb9e19094440c317e21b-Abstract-round1.html"
}

@inproceedings{thorne-etal-2018-fever,
    title = "{FEVER}: a Large-scale Dataset for Fact Extraction and {VER}ification",
    author = "Thorne, James  and
      Vlachos, Andreas  and
      Christodoulopoulos, Christos  and
      Mittal, Arpit",
    editor = "Walker, Marilyn  and
      Ji, Heng  and
      Stent, Amanda",
    booktitle = "Proceedings of the 2018 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)",
    month = jun,
    year = "2018",
    address = "New Orleans, Louisiana",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/N18-1074/",
    doi = "10.18653/v1/N18-1074",
    pages = "809--819"
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for subhamoybasu/laya-evidence-v2

Adapter
(19)
this model

Papers for subhamoybasu/laya-evidence-v2