Instructions to use subhamoybasu/laya-evidence-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use subhamoybasu/laya-evidence-v1 with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya Evidence
Laya Evidence finds the sentences (or list items) of a document that support or contradict
each of one or more claims. It adds a fourth request type, evidence, to
Laya, next to choice, score and noul, and packs the
claims into one encoder input per document window.
It is a specialist. Its strongest evaluated setting is the contract questions it was trained on (ContractNLI's 17 hypothesis types), asked of contracts it has not seen. On claim types and kinds of document unlike its training data it is much weaker (see Limitations). It is not a general-purpose fact checker.
Summary
For each claim the model returns a status (evidence_found or insufficient_evidence) and the
segments it keeps. Each segment comes with its character offsets in the document, its exact text, a
label (supports or contradicts), an evidence score and a polarity confidence.
Each encoder row is [CLS] head [SEP] [MASK] claim 1 … [MASK] claim n [SEP] document window [SEP].
The row runs through Laya's encoder; Laya's head layers are not used on this path. The rows are centred, a trained evidence type vector marking the request type is added to every token vector, and the rows are layer-normalised and passed through 2 adapter layers. An evidence head then scores every (claim, segment) pair as none, supports or
contradicts, from the claim's [MASK] vector and the segment's token vectors. A segment is kept when
its evidence score, P(supports) + P(contradicts), is at least the threshold τ.
This repository holds only the parameters Laya Evidence trains: the evidence type vector, a LayerNorm, 2 adapter layers and the evidence head. The Laya checkpoint is downloaded from the Hub at a pinned revision when the model is loaded (see Base model).
| Developed by | Subhamoy Basu (subhamoy.basu@berkeley.edu) |
| Code | https://github.com/subhamoy-basu/laya-evidence |
| License | Apache-2.0 (see LICENSE and NOTICE) |
| Base checkpoint | convaiinnovations/laya (root checkpoint) at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 |
| Configuration | B: all of Laya's weights frozen and unmodified |
| Encoder budget | 1,024 tokens per row |
| Claims per encoder pass | up to 17, within 448 claim tokens |
| Default threshold τ | 0.74, the value that maximises evidence F1 on the dev data of the training datasets (almost all of it ContractNLI's) |
| Segmenter | pysbd 0.3.4 (or the caller's own segments) |
| Training run | joint_B_s13, seed 13, code at commit 5dcea41ea6e524639e472f31b8212639dcfb6459 |
| Weights | evidence.safetensors, sha256 c44d3cce7a1d2c504f3537cbab9aacbe6d968749fe80524671c7fcd1e0365b81 |
Intended use
- Finding the evidence for recurring contract questions like the ones it was evaluated on, with a person reviewing the output, after you have checked it on documents like yours.
- Adding evidence extraction to an existing Laya workflow, as an experimental request type.
- Research on scoring several claims per encoder pass, calibration, parameter-efficient adaptation, and where such a model stops working.
Scientific abstracts are experimental (see Limitations). Other kinds of document and claim have not been evaluated on a benchmark.
Out-of-scope uses
- Deciding whether a claim is true. The model judges a claim only against the document it is given, and it gives no document-level verdict.
- Reading an empty result as a negative.
insufficient_evidencemeans that no segment passed the threshold. It does not mean that the claim is false or that the document has no relevant evidence: evidence can be missed. - Reading one returned segment as the whole evidence. A segment can be evidence only together with other segments.
- General-purpose verification. Checking arbitrary generated answers or summaries, or claim types and domains unlike the training data, without your own evaluation on representative data.
- Decisions without review. Legal, medical, scientific or other consequential decisions. The output is an unreviewed model prediction, not a legal review of a contract.
- Trusting the scores on new domains. The thresholds and probabilities were fitted on contract and abstract dev data.
- Languages other than English. The model was trained and evaluated on English text only.
How to use
Requires Python 3.11 or later.
pip install git+https://github.com/subhamoy-basu/laya-evidence
from laya_evidence import EvidenceExtractor
ex = EvidenceExtractor.from_pretrained("subhamoybasu/laya-evidence-v1")
doc = ("3. The Receiving Party may disclose Confidential Information to its employees and advisers who "
"need to know it for the Purpose. 4. The Receiving Party shall not reverse engineer, decompile or "
"disassemble any samples or software provided under this Agreement.")
claims = ['Receiving Party may share some Confidential Information with some of its employees.', 'Receiving Party may independently develop information similar to Confidential Information.']
for r in ex.find(doc, claims):
print(r.claim_id, r.status, [(s.label, s.text) for s in r.spans])
from_pretrained also downloads the frozen Laya checkpoint at its pinned revision. find accepts a list
of claims (ids c1, c2, …) or a dict of id → claim, and the keyword arguments threshold (default
τ), segments (your own sorted (start, end) character spans instead of the built-in segmenter),
max_claims_per_pass, top_k and debug. find_batch takes lists of documents and claims. For every
span, document[span.start_char:span.end_char] == span.text (end_char is exclusive).
LayaEvidence exposes the same model as a fourth Laya request type, next to Laya's own:
from laya_evidence import LayaEvidence
agent = LayaEvidence.from_pretrained("subhamoybasu/laya-evidence-v1")
doc = ("3. The Receiving Party may disclose Confidential Information to its employees and advisers who "
"need to know it for the Purpose. 4. The Receiving Party shall not reverse engineer, decompile or "
"disassemble any samples or software provided under this Agreement.")
out = agent.predict(doc, {
"shares": {"type": "noul", "instructions": "Does the agreement allow sharing with employees?"},
"grounding": {"type": "evidence", "claims": ['Receiving Party may share some Confidential Information with some of its employees.']},
})
print(out["answers"]["shares"])
print(out["answers"]["grounding"]["claims"]["c1"])
Interpreting the output
evidence_found: at least one segment passed the evidence-score threshold. It is the model's prediction, not a verification of the claim.insufficient_evidence: no segment passed the threshold. Relevant evidence may have been missed.supports/contradicts: the predicted direction of a returned segment. It can be wrong, and some evidence means something only together with other segments.evidence_score: the raw score the threshold applies to, P(supports) + P(contradicts). It is not a probability.polarity_confidence: the share of the evidence score that the returned label holds. It is not calibrated.- Offsets:
document[start_char:end_char] == textalways holds, so a span can be located exactly. That does not make it relevant, complete or correctly labelled. - The result depends on the threshold, on the segmentation and on the other claims packed into the same encoder pass.
Training data
The weights were trained on ContractNLI and SciFact. Documents are split into segments, and every (claim, segment) pair gets one label: none, supports or contradicts.
| Dataset | Use | Train | Dev (model selection, τ) | Test (final evaluation only) |
|---|---|---|---|---|
| ContractNLI | training and evaluation | 423 docs / 7,191 claims | 61 docs / 1,037 claims | 123 docs / 2,091 claims |
| SciFact | training and evaluation | 491 docs / 788 claims | 74 docs / 131 claims | 283 docs / 339 claims |
Claims are counted once per document they are asked of, so every ContractNLI contract contributes its
17 hypotheses. More detail on sources, conversion and splits is in
docs/data.md.
- ContractNLI (Koreeda and Manning, 2021) asks the same 17 hypotheses of every non-disclosure agreement. The official splits are used, and each official span is one segment. Evidence spans take the (contract, hypothesis) label: Entailment → supports, Contradiction → contradicts. Every other span, and every span of a NotMentioned pair, is none.
- SciFact (Wadden et al., 2020) pairs scientific claims with research abstracts. Each abstract sentence is one segment (the title is not used). Rationale sentences take the rationale's label, and all other sentences are none. The official dev claims are the test split, and an internal dev split (about 15% of the official train claims, split by connected groups of claims and abstracts) is held out from training.
Attribution
Training data attribution.
ContractNLI: "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Yuta Koreeda and Christopher D. Manning; licensor Hitachi America, Ltd.; https://stanfordnlp.github.io/contract-nli/; licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: document-level labels and evidence-span indices were converted to one label per span (supports, contradicts or none).
SciFact: "Fact or Fiction: Verifying Scientific Claims", David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan and Hannaneh Hajishirzi; https://github.com/allenai/scifact; claims and evidence annotations licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) and provided without warranties. Modified: rationale annotations were converted to one label per abstract sentence. Contains information from the SciFact corpus (https://github.com/allenai/scifact), whose abstracts are drawn from S2ORC (https://github.com/allenai/s2orc), which is made available under the ODC Attribution License (https://opendatacommons.org/licenses/by/1-0/).
Base model: Laya by Convai Innovations (https://huggingface.co/convaiinnovations/laya), Apache-2.0, loaded at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Laya's weights are not included in this repository. This model adds trained layers on top of Laya's frozen encoder; Laya itself is not modified. Laya Evidence is an independent project, not affiliated with or endorsed by Convai Innovations.
Evaluation
All numbers are on the test splits: ContractNLI's official test set, and SciFact's official dev claims, which this project uses as its test split. They were opened only after the shipped run and its settings were frozen in configs/final.lock.json; every access is logged in results/final/TEST_ACCESS.log. Threshold metrics use each model's own τ (τ = 0.74 here). mAP / P@R80 are the official ContractNLI evidence-identification metrics, over the (contract, hypothesis) pairs whose gold label is not NotMentioned; ev-F1@τ is segment-level evidence F1, and ev+pol F1 also requires the right label. The first table is on the clean test subset, the test documents that do not also appear in the training or dev data; the second is on the full official test splits, which published results use.
| Model | CNLI mAP (macro) | CNLI P@R80 (macro) | CNLI mAP (micro) | CNLI P@R80 (micro) | CNLI ev-F1@τ | CNLI ev+pol F1 | CNLI false-alarm | SciFact ev-F1@τ | SciFact ev+pol F1 | Latency p50 GPU / CPU (s/doc) |
|---|---|---|---|---|---|---|---|---|---|---|
| Ours, mean ± std over 3 seeds | 0.784 ± 0.005 | 0.660 ± 0.013 | 0.765 ± 0.009 | 0.628 ± 0.028 | 0.684 ± 0.006 | 0.647 ± 0.005 | 0.244 ± 0.036 | 0.234 ± 0.033 | 0.193 ± 0.025 | — |
| Ours, shipped run (joint_B_s13) | 0.778 [0.756, 0.811] | 0.646 [0.577, 0.704] | 0.756 | 0.596 | 0.677 [0.650, 0.701] | 0.643 [0.615, 0.670] | 0.216 [0.184, 0.251] | 0.220 [0.133, 0.304] | 0.178 [0.103, 0.250] | 0.0453 / 2.48 |
| BL-1C (one claim per pass) | 0.767 | 0.624 | 0.751 | 0.592 | 0.662 | 0.631 | 0.201 | 0.213 | 0.160 | 0.413 / 33.8 |
| BL-MB (plain ModernBERT) | 0.696 | 0.519 | 0.689 | 0.511 | 0.616 | 0.589 | 0.346 | 0.208 | 0.135 | — |
| BL-NLI (zero-shot DeBERTa NLI) | 0.261 | 0.050 | 0.119 | 0.036 | 0.181 | 0.165 | 0.462 | 0.361 | 0.353 | 1.18 / 54 |
- Clean test subset: the official test splits without the test documents that also appear in our training or dev data (
data/splits/test_clean.json,scripts/split_overlap.py): 109 of 123 ContractNLI test contracts are kept (dropped: word 8-gram Jaccard similarity >= 0.5 with a train or dev contract) and 102 of 283 SciFact test abstracts (dropped: abstract also in our SciFact train split or dev_int). The subset was defined after the test results were first computed;scripts/evaluate_clean.pycomputes these numbers from the saved test scores. Models, τ and settings are those of the official-split table. - Rows as in the official table: mean ± sample std over the seed runs joint_B_s13, joint_B_s14, joint_B_s15, each at its own τ; the shipped run's row has its 95% document-level bootstrap interval on the clean subset and its latency; the baselines are single runs. BL-NLI keeps its dev-tuned τ = 0.93.
- The published Span NLI BERT results are for the full test split, so they are compared in the official-split table only. The experiments E1–E5 (
results/final/summary.md) compare settings of the same model on the same documents and use the full split; E3 also has clean rows.
Official test splits (overlapping the training data):
| Model | CNLI mAP (macro) | CNLI P@R80 (macro) | CNLI mAP (micro) | CNLI P@R80 (micro) | CNLI ev-F1@τ | CNLI ev+pol F1 | CNLI false-alarm | SciFact ev-F1@τ | SciFact ev+pol F1 | Latency p50 GPU / CPU (s/doc) |
|---|---|---|---|---|---|---|---|---|---|---|
| Ours, mean ± std over 3 seeds | 0.797 ± 0.003 | 0.686 ± 0.016 | 0.774 ± 0.008 | 0.660 ± 0.024 | 0.698 ± 0.007 | 0.664 ± 0.005 | 0.238 ± 0.037 | 0.562 ± 0.017 | 0.316 ± 0.031 | — |
| Ours, shipped run (joint_B_s13) | 0.793 [0.773, 0.821] | 0.668 [0.606, 0.722] | 0.765 | 0.633 | 0.690 [0.667, 0.713] | 0.659 [0.633, 0.684] | 0.212 [0.182, 0.243] | 0.548 [0.482, 0.613] | 0.281 [0.222, 0.341] | 0.0453 / 2.48 |
| BL-1C (one claim per pass) | 0.779 | 0.648 | 0.758 | 0.618 | 0.679 | 0.651 | 0.197 | 0.569 | 0.307 | 0.413 / 33.8 |
| BL-MB (plain ModernBERT) | 0.705 | 0.558 | 0.702 | 0.532 | 0.622 | 0.595 | 0.357 | 0.560 | 0.211 | — |
| BL-NLI (zero-shot DeBERTa NLI) | 0.253 | 0.049 | 0.118 | 0.035 | 0.178 | 0.163 | 0.466 | 0.434 | 0.422 | 1.18 / 54 |
| BL-PUB: Span NLI BERT (BERT-base), published | — | — | 0.885 ± 0.025 | 0.663 ± 0.093 | — | — | — | — | — | — |
| BL-PUB: Span NLI BERT (BERT-large), published | — | — | 0.922 ± 0.006 | 0.793 ± 0.018 | — | — | — | — | — | — |
Ours: mean ± sample std (ddof=1) over the seed runs joint_B_s13, joint_B_s14, joint_B_s15 of the shipped configuration, each at its own tau_default (0.74, 0.73, 0.48); max_claims_per_pass = 17. The shipped run joint_B_s13 alone has its own row, with its 95% document-level bootstrap interval [lo, hi] (evaluate.py --bootstrap) where present and its latency; the baselines below are single runs too.
Macro = macro_label_micro_doc (our headline metric); micro = micro_label_micro_doc (the averaging of the ContractNLI paper, see BL-PUB); model selection averages macro with SciFact dev_int AP_all. ev-F1@τ and ev+pol F1 are over (document, claim, segment) pairs; false-alarm is over (document, claim) with no gold evidence (docs/evaluation.md).
BL-1C: the shipped weights (joint_B_s13) at max_claims_per_pass = 1, τ = 0.74.
BL-MB: ablation_modernbert_s13 at τ = 0.68.
BL-NLI: MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli@b3546ea6b0346eb6f8d5d68b13c7dc6d0376b3d7; τ = 0.93 maximises evidence F1 on the dev union of both sources (dev F1 0.204).
BL-PUB: Koreeda and Manning, ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts, Findings of EMNLP 2021 (arXiv:2110.01799), Table 4 (Main results), rows 'Ours (BERTbase)' and 'Ours (BERTlarge)'; the same numbers appear in Table 3 (backbone 'None' rows) (test split). The paper averages as micro_label_micro_doc.span.map, so only those columns are filled; ± is the paper's spread: Mean over the 3 of 10 hyperparameter runs with the best development macro-average NLI accuracy; '_std' is the standard deviation over those three runs (the paper's subscripts).
BL-LLM, the optional prompted-LLM baseline, was not run (no results/baselines/llm_contractnli_test.json), so it has no row.
BL-MAJ (per-hypothesis majority polarity in train; polarity accuracy only): 0.886 over 2404 gold-evidence segments of ContractNLI test (0.886 restricted to doc_label ≠ none pairs, the same segments). Ours: polarity accuracy TP′/TP 0.952 ± 0.003 at each run's τ_default, over the gold-evidence segments each run also predicts, so the denominators differ. On those same segments BL-MAJ's label is right for 0.912 ± 0.004 (further analyses).
Latency: p50 seconds per document, from E5 (
results/final/summary.md).Latency hardware: cpu_count=24, cpu_model=unknown, device=cpu, machine=x86_64, platform=Linux-4.19.0-gvisor-x86_64-with-glibc2.36, sched_affinity_cpus=24, torch_num_interop_threads=12, torch_num_threads=24
Latency hardware: bf16_supported=True, cpu_count=21, cpu_model=unknown, cuda=13.0, device=cuda, gpu=NVIDIA L40S, gpu_total_mem_bytes=47665709056, machine=x86_64, platform=Linux-4.19.0-gvisor-x86_64-with-glibc2.36, sched_affinity_cpus=21, torch_num_interop_threads=11, torch_num_threads=8
E1: claims per encoder pass
Claims of each ContractNLI test contract are shuffled (seed 13) and packed n per pass. The shipped model packs up to 17 claims per pass, chosen on dev as the largest n whose mAP is within 0.01 of one claim per pass.
| n | mAP (macro) | P@R80 (macro) | ev-F1@τ | Forwards/doc | Tokens/doc |
|---|---|---|---|---|---|
| 1 | 0.779 | 0.649 | 0.679 | 50.5 | 42613 |
| 2 | 0.784 | 0.663 | 0.685 | 27.0 | 22968 |
| 4 | 0.789 | 0.674 | 0.687 | 15.3 | 13218 |
| 8 | 0.792 | 0.668 | 0.685 | 9.6 | 8396 |
| 17 | 0.791 | 0.668 | 0.690 | 4.0 | 3729 |
E2: interference between claims
Each ContractNLI test (contract, claim) is scored alone and with k random companion claims from the same contract (seed 13). The flip rates are the fractions of the claim's segments whose keep decision (ev ≥ τ), or whose label among segments kept both times, differs from scoring the claim alone.
| k | mean |Δev| | Decision flip rate | Polarity flip rate |
|---|---|---|---|
| 1 | 0.002 | 0.002 | 0.006 |
| 4 | 0.004 | 0.003 | 0.010 |
| 8 | 0.005 | 0.004 | 0.013 |
| 16 | 0.006 | 0.004 | 0.014 |
Almost every segment scores far below τ either way, so the decision flip rate over all segments is small. The same flips counted over the segments that matter:
| k | Segments kept alone or with companions: flipped | Gold evidence segments: flipped |
|---|---|---|
| 1 | 0.140 | 0.053 |
| 4 | 0.204 | 0.083 |
| 8 | 0.248 | 0.098 |
| 16 | 0.276 | 0.116 |
E3: unseen hypotheses
A separate model (ContractNLI only, otherwise the shipped settings) was trained with 4 of the 17 ContractNLI hypotheses removed from its training and dev data; the held-out and the seen hypotheses are scored separately on the test split.
| Model | Hypotheses | mAP (macro) | P@R80 (macro) | mAP (micro) | P@R80 (micro) |
|---|---|---|---|---|---|
| E3 run holdout_B_s13 (held-out hypotheses not trained on) | held-out | 0.305 | 0.198 | 0.346 | 0.066 |
| E3 run holdout_B_s13 (held-out hypotheses not trained on) | seen | 0.807 | 0.708 | 0.780 | 0.681 |
| E3 run holdout_B_s13, clean test subset | held-out | 0.300 | 0.197 | 0.357 | 0.076 |
| E3 run holdout_B_s13, clean test subset | seen | 0.794 | 0.688 | 0.768 | 0.654 |
Full results: results/final/summary.md.
Limitations
- Segment level only. The model returns whole segments (sentences or list items), never shorter
spans. Its output depends on the segmentation; pass
segments=to use your own. - Uncalibrated scores.
evidence_scoreandpolarity_confidenceare not calibrated probabilities. The default τ = 0.74 maximises evidence F1 on the dev data; choose your own τ for the precision and recall you need. - Spans inherit document-level labels. The segment labels are converted from coarser annotations:
every ContractNLI evidence span takes the label of its (contract, hypothesis) pair, and every sentence
of a SciFact rationale takes the rationale's label. A segment can therefore be labelled
contradictsalthough it contradicts the claim only together with other segments. - 17 fixed ContractNLI hypotheses. ContractNLI asks the same 17 hypotheses of every contract, so claims unlike them may be handled worse. In E3, a model trained without 4 of the hypotheses reached a test mAP (macro) of 0.305 on them, against 0.807 on the hypotheses it was trained on.
- Domain-specific and hypothesis-bound. Trained only on non-disclosure agreements (ContractNLI) and scientific research abstracts (SciFact), English only. It works
best on documents like its training data with claims worded like ContractNLI's hypotheses ("Receiving
Party shall …"). On SciFact it does well mainly on abstracts it was trained on: test evidence F1 is 0.750 on the 157 test abstracts that are also in its training data, but 0.234 on the 102 that are in neither its training nor its dev data (mean over 3 seeds), below a zero-shot NLI baseline's 0.361 there. On a fixed probe (
scripts/probe_out_of_domain.pyin the code repository) a free paraphrase scores lower than the hypothesis-style wording of the same claim, two of three clauses score far lower as a bare snippet than inside a full agreement, and short general-domain texts get no evidence even for verbatim claims. Check it on your own data before relying on it. - Polarity errors remain. A contradicting clause can be returned as
supports. Among the gold evidence segments it returns (its true positives), its label is right for 0.952 on ContractNLI. On SciFact it is right for 0.562, where always answering supports would be right for 0.631 of the same segments. These rates are conditional: they leave out the evidence it misses and the segments it returns wrongly. The results table's ev+pol F1 counts both. - Claim interference. Claims packed into one encoder pass attend to each other, so a claim's scores
can change with the other claims in the request. In E2, with k = 16 companion claims, the keep decisions of a claim's segments flipped at a rate of 0.004 (mean |Δev| 0.006) compared with scoring the claim alone; among the segments kept alone or with the companions, 0.276 flipped. Pass
max_claims_per_pass=1to score every claim on its own, at the cost of more encoder passes.
Base model
Laya Evidence runs on Laya by Convai Innovations
(Apache-2.0): convaiinnovations/laya (root checkpoint) at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851, loaded with the laya==0.3.21 package. Laya's English
checkpoints are fine-tuned from ModernBERT-large
(Apache-2.0).
Laya's weights are not included in this repository. from_pretrained downloads them from the Hub
at revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851 and attaches this repository's weights with a strict state-dict load, so
later changes to the base repository do not change this model. This is a configuration B model: Laya's weights are used frozen and unmodified, and choice, score and noul requests go to Laya unchanged.
Laya's own capabilities are unchanged. On Laya's published benchmark (49 suites, 17,416 choice, score and noul questions, rebuilt with Laya's own code), this model answers exactly as vanilla Laya (accuracy):
| Benchmark | Type | Laya, published | Laya (vanilla, our run) | Laya inside Laya Evidence |
|---|---|---|---|---|
| typed-decisions, all | all three | 0.3615 | 0.3615 | 0.3615 |
| typed-decisions | choice | 0.2883 | 0.2883 | 0.2883 |
| typed-decisions | score | 0.3225 | 0.3225 | 0.3225 |
| typed-decisions | noul | 0.4867 | 0.4867 | 0.4867 |
| SST-5 | score | 0.3717 | 0.3717 | 0.3717 |
| BoolQ | noul | 0.8300 | 0.8300 | 0.8300 |
| prompt-injections | noul | 0.6983 | 0.6983 | 0.6983 |
| AG News | choice | 0.9467 | 0.9467 | 0.9467 |
| DAIR Emotion | choice | 0.5733 | 0.5733 | 0.5733 |
| MASSIVE intent, 14 languages | choice | 0.3405 | 0.3402 | 0.3402 |
| MASSIVE scenario, 14 languages | choice | 0.3036 | 0.3036 | 0.3036 |
| XNLI, 15 languages | choice | 0.5436 | 0.5438 | 0.5438 |
Laya's published numbers: typed-decisions from Laya's CPU run, the rest from its T4 run. Our runs: fp32 on an NVIDIA L40S GPU; 40 of the 49 suites match the published T4 accuracy exactly. Comparing every full answer as predict returns it (probabilities, confidence and action probability, which Laya rounds to 4 decimals) with vanilla Laya's: 17,416 of 17,416 answers identical in fp32; 17,416 of 17,416 in bf16; 17,416 of 17,416 with an evidence request before every call; 2,000 of 2,000 with an evidence question in the same call.
Details: docs/evaluation.md.
Citation
If you use this model, please cite Laya Evidence, and also ContractNLI, SciFact (and its S2ORC corpus), Laya and ModernBERT:
@software{basu-2026-laya-evidence,
title = "{Laya Evidence}",
author = "Basu, Subhamoy",
year = "2026",
version = "1.0.0",
url = "https://github.com/subhamoy-basu/laya-evidence",
note = "Model weights: \url{https://huggingface.co/subhamoybasu/laya-evidence-v1}. Apache License 2.0"
}
@inproceedings{koreeda-manning-2021-contractnli-dataset,
title = "{C}ontract{NLI}: A Dataset for Document-level Natural Language Inference for Contracts",
author = "Koreeda, Yuta and
Manning, Christopher",
editor = "Moens, Marie-Francine and
Huang, Xuanjing and
Specia, Lucia and
Yih, Scott Wen-tau",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
month = nov,
year = "2021",
address = "Punta Cana, Dominican Republic",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2021.findings-emnlp.164/",
doi = "10.18653/v1/2021.findings-emnlp.164",
pages = "1907--1919"
}
@inproceedings{wadden-etal-2020-fact,
title = "Fact or Fiction: Verifying Scientific Claims",
author = "Wadden, David and
Lin, Shanchuan and
Lo, Kyle and
Wang, Lucy Lu and
van Zuylen, Madeleine and
Cohan, Arman and
Hajishirzi, Hannaneh",
editor = "Webber, Bonnie and
Cohn, Trevor and
He, Yulan and
Liu, Yang",
booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)",
month = nov,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.emnlp-main.609/",
doi = "10.18653/v1/2020.emnlp-main.609",
pages = "7534--7550"
}
@inproceedings{lo-etal-2020-s2orc,
title = "{S}2{ORC}: The Semantic Scholar Open Research Corpus",
author = "Lo, Kyle and
Wang, Lucy Lu and
Neumann, Mark and
Kinney, Rodney and
Weld, Daniel",
editor = "Jurafsky, Dan and
Chai, Joyce and
Schluter, Natalie and
Tetreault, Joel",
booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.acl-main.447/",
doi = "10.18653/v1/2020.acl-main.447",
pages = "4969--4983"
}
@misc{convai-2026-laya,
title = "{Laya}",
author = "{Convai Innovations}",
year = "2026",
howpublished = "\url{https://github.com/NandhaKishorM/laya}",
note = "Python package laya 0.3.21; model weights at \url{https://huggingface.co/convaiinnovations/laya}. Apache License 2.0"
}
@misc{warner-etal-2024-modernbert,
title = "Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference",
author = {Warner, Benjamin and
Chaffin, Antoine and
Clavi{\'e}, Benjamin and
Weller, Orion and
Hallstr{\"o}m, Oskar and
Taghadouini, Said and
Gallagher, Alexis and
Biswas, Raja and
Ladhak, Faisal and
Aarsen, Tom and
Cooper, Nathan and
Adams, Griffin and
Howard, Jeremy and
Poli, Iacopo},
year = "2024",
eprint = "2412.13663",
archivePrefix = "arXiv",
primaryClass = "cs.CL",
url = "https://arxiv.org/abs/2412.13663",
note = "Also published at ACL 2025, doi:10.18653/v1/2025.acl-long.127, with a different author list"
}
Model tree for subhamoybasu/laya-evidence-v1
Base model
convaiinnovations/laya