Instructions to use Pranshurs/groundcheck-modernbert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Pranshurs/groundcheck-modernbert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Pranshurs/groundcheck-modernbert")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Pranshurs/groundcheck-modernbert") model = AutoModelForSequenceClassification.from_pretrained("Pranshurs/groundcheck-modernbert", device_map="auto") - Notebooks
- Google Colab
- Kaggle
GroundCheck v2 (ModernBERT-base)
A ~150M-parameter encoder that checks whether a RAG answer is supported by the source it
was given. Given an answer and a source (and optionally the question), it returns
grounded or hallucinated with P(grounded). It runs on CPU.
- Labels:
0 = grounded,1 = hallucinated, andP(grounded) = softmax(logits)[0]. - Code, evaluation harness and provenance: https://github.com/Pranshurs/groundcheck
- Every number below is traceable to a committed report in that repository.
Which weights
| Recommended revision | 998cec35563d6b90947409d1c7510adac7f7c80c (the groundcheck-rag package pins this) |
| Weights committed in | 0c7dd0636c0d58b5e3865db8f1deee5a5acfeb6c (later commits change only this card) |
model.safetensors SHA-256 |
9ec331ba6d8a9d93236dd72b239df518b07e61241323876f8aafc223d959461a |
| Base model | answerdotai/ModernBERT-base (Apache-2.0) |
Hashes for every file are in the repository's MODEL_PROVENANCE.json.
python -m eval.provenance --verify re-checks them.
Results
All measured on CPU against the pinned weights, using test sets rebuilt from pinned public
datasets. F1 is for the hallucinated class, and the 95% CIs come from a 2,000-resample
bootstrap.
| Suite | n | max_length 512 (training protocol) | max_length 2048 (groundcheck-rag default) |
|---|---|---|---|
| RAGTruth test (first 2,500 of 2,700 responses) | 2,500 | F1 0.682 [0.658, 0.705], acc 0.746 | F1 0.696 [0.672, 0.718], acc 0.758 |
| VitaminC test | 2,000 | acc 0.850, F1 0.845 | identical (all inputs are short) |
| One-fact flips caught (regenerated holdout) | 500 | 78.0% | 87.2% |
| Same answers unflipped, kept grounded | 500 | 76.0% | 74.0% |
- 512 tokens reproduces the training run's published numbers (F1 0.6824, acc 0.7468; VitaminC acc 0.8495) to within one prediction in 2,500.
- On the same rows, 2048 tokens scores a paired ΔF1 of +0.014 (95% CI −0.001 to +0.028). That's a small gain, at about 2.5× the CPU time on long documents.
- The flipped-fact holdout is regenerated. The original run's sample depended on Python's per-process hash seed and can't be rebuilt. The original sample reported 80.4% caught and 76.2% kept.
- RAGTruth uses the first 2,500 of the 2,700 test responses, the same subset the training run evaluated on.
External published reference (not a controlled comparison)
The RAGTruth paper (Niu et al., 2024; tabulated in LettuceDetect, 2025) reports F1 0.634 for a zero-shot GPT-4-turbo prompt judge. That figure comes from a different protocol: a prompted judge scored on all 2,700 test responses. It was not run head-to-head with GroundCheck, so it's there for orientation and doesn't support a "beats GPT-4" claim.
Latency (one machine, not a guarantee)
Apple M1, CPU, 4 torch threads, single requests after warm-up, at max_length 2048:
| Pair length | p50 |
|---|---|
| ≤ 128 tokens | 39 ms |
| 129–512 tokens | 172 ms |
| 513–2,048 tokens | 349 ms |
| > 2,048 tokens | ~1.45 s |
Other hardware will differ. Measure with python -m bench.latency.
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
name, rev = "Pranshurs/groundcheck-modernbert", "998cec35563d6b90947409d1c7510adac7f7c80c"
tok = AutoTokenizer.from_pretrained(name, revision=rev)
model = AutoModelForSequenceClassification.from_pretrained(name, revision=rev).eval()
source = "France's capital and largest city is Paris."
answer = "Paris is the capital of France."
enc = tok(source, answer, truncation="only_first", max_length=512, return_tensors="pt")
with torch.no_grad():
grounded = torch.softmax(model(**enc).logits, dim=-1)[0, 0].item()
print("grounded" if grounded >= 0.5 else "hallucinated", round(grounded, 3))
Or use the library (pip install "groundcheck-rag[model]"), which pins the revision,
truncates only the source, and raises an error rather than substituting a heuristic when the
model can't load:
from groundcheck import GroundCheck
print(GroundCheck().check(source="...", answer="..."))
Training
Fine-tuned from answerdotai/ModernBERT-base (pinned at 8949b909) as a sequence-pair
classifier: the premise is the optional question plus the source, and the hypothesis is the
answer. The training run was on a single Kaggle P100 with torch 2.4.1 and transformers 4.49.0:
3 epochs, batch 16, learning rate 2e-5, linear schedule, warmup 0.06, weight decay 0.01,
seed 42, fp16, sequence length 512.
The 28,500 training rows break down as:
- 10,000 from RAGTruth (
wandb/RAGTruth-processed@eb4f4b9d), with the question dropped on ~50% of rows. - 16,000 from VitaminC (
tals/vitaminc@be6febb7): SUPPORTS → grounded; REFUTES and NOT ENOUGH INFO → hallucinated. - 2,500 rule-based one-fact flips (a number, date, direction word or entity) of grounded RAGTruth answers.
There is no LLM-generated augmentation and no private data. The recipe and data builders
are in the repository's training/ directory.
Intended use
Use it as a post-generation check in RAG pipelines: flag answers the retrieved source doesn't support, so they can be routed to review, regeneration or a stronger checker. It checks support against the provided source only and isn't a world-knowledge fact-checker.
Limitations
- English only. Verdicts are per answer, not per span.
- Long sources are truncated from the end. 76% of RAGTruth test pairs exceed 512 tokens and 19 of 2,500 exceed 2,048, so chunk long documents.
- RAGTruth precision is about 0.63, so roughly a third of
hallucinatedverdicts on long RAG answers are false alarms. Scores aren't calibrated probabilities; tune the threshold on your own data. - About a quarter of unedited grounded answers in the minimal-edit holdout are flagged.
- It hasn't been evaluated on adversarial or out-of-domain inputs.
License and data terms
| What | Terms |
|---|---|
| These model weights | MIT |
Base model, answerdotai/ModernBERT-base |
Apache-2.0 |
| The GroundCheck code | Apache-2.0 |
| Training data | Keeps its upstream terms, which the MIT license on the weights does not change |
The upstream data terms (detailed in the repository's DATA_LICENSES.md):
- RAGTruth is MIT. Its source passages come
from:
- MS MARCO: Microsoft's terms allow non-commercial research use only.
- The Yelp Open Dataset: academic and non-commercial use only.
- CNN/DailyMail: the articles are copyrighted by their publishers.
- VitaminC is CC BY-SA 3.0.
Whether a dataset's non-commercial terms extend to a model trained on it is legally unsettled. Review the data terms before any commercial use. This card is not legal advice.
- Downloads last month
- 47
Model tree for Pranshurs/groundcheck-modernbert
Base model
answerdotai/ModernBERT-base