Tessary groundedness classifier v1

This model checks an agent's response against the documents the agent retrieved, and marks the words those documents don't support. Tessary built it for its groundedness classifier, and you can also run it on its own.

Each word in the response gets one of three labels:

label meaning
O Supported by the retrieved documents.
BASELESS Not in the documents. The agent either made it up or got it from somewhere you didn't pass in, such as a tool result.
CONFLICT Contradicts the documents.

The model reads the documents and the response together in one pass of up to 8,192 tokens. It runs on CPU, taking about 1 second for a typical response on 4 vCPUs.

Scores and thresholds

For each word, the model gives the probability that it's unsupported: P(BASELESS) + P(CONFLICT). The response score is the highest of those probabilities, so one unsupported claim is enough to flag the whole response.

Pick a threshold for the response score:

  • 0.975 if false alarms are expensive. It flags about 2% of supported responses.
  • 0.5 for the best balance between catching problems and avoiding false alarms.

Both thresholds come from a public benchmark of news and question-answering responses. Check them on your own traffic before relying on them.

Run the model

pip install "transformers>=4.48" torch

The model expects its input in the exact layout it was trained on. Build that layout with context_for below, and pass each retrieved document as its own passage. Don't join several documents into one.

import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer

MODEL = "tessaryai/groundedness-classifier-v1"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForTokenClassification.from_pretrained(MODEL).eval()


def context_for(passages: list[str], question: str | None) -> str:
    """The context layout the model was trained on (LettuceDetect's). Do not change it."""
    ctx = "\n".join(f"passage {i + 1}: {p}" for i, p in enumerate(passages))
    if question:
        return ("Briefly answer the following question:\n" f"{question}\n"
                f"Bear in mind that your response should be strictly based on the following {len(passages)} passages:\n"
                f"{ctx}\nIn case the passages do not contain the necessary information to answer the question, "
                'please reply with: "Unable to answer based on given passages."\noutput:')
    return f"Summarize the following text:\n{ctx}\noutput:"


def score(passages: list[str], question: str | None, response: str, threshold: float = 0.5):
    """Response score, plus the response spans whose score is at or above `threshold`."""
    enc = tok(context_for(passages, question), response, truncation="only_first", max_length=8192,
              return_offsets_mapping=True, return_tensors="pt")
    offsets = enc.pop("offset_mapping")[0].tolist()
    seq = enc.sequence_ids()
    with torch.no_grad():
        p = model(**enc).logits[0].softmax(-1)
    resp = [i for i, s in enumerate(seq) if s == 1]
    unsupported = p[:, 1] + p[:, 2]                      # P(BASELESS) + P(CONFLICT) per token
    spans = []                                           # merge adjacent flagged tokens
    for i in resp:
        if unsupported[i] < threshold:
            continue
        label = model.config.id2label[int(p[i, 1:].argmax()) + 1]
        start, end = offsets[i]
        if spans and spans[-1][2] == label and start - spans[-1][1] <= 1:
            spans[-1][1] = end
        else:
            spans.append([start, end, label])
    return float(unsupported[resp].max()), [(response[a:b].strip(), label) for a, b, label in spans]


passages = [
    "The Pro plan costs $49 per month and includes 10 seats.",
    "Annual Pro plans are billed once a year at a 20% discount.",
]
question = "How much is the Pro plan?"
response = "The Pro plan costs $59 per month and includes 10 seats. It comes with a free 14-day trial."

response_score, spans = score(passages, question, response)
print(f"response score: {response_score:.3f}")
for text, label in spans:
    print(f"{label:9} {text}")

Output:

response score: 1.000
CONFLICT  The Pro plan costs $59 per month
CONFLICT  10
BASELESS  It comes with a free 14-day trial.

The model catches the wrong price and the made-up trial. It also flags the correct 10: spans tend to cover more of the sentence than the error itself.

If your agent summarizes rather than answers a question, pass question=None. To keep results stable across updates to this repository, pin the revision: from_pretrained(MODEL, revision="6746fa25").

Run it with ONNX Runtime

onnx/model.onnx is the same model exported to ONNX (fp32, 1.58 GB), for ONNX Runtime and transformers.js. Its scores match the PyTorch model. It takes input_ids and attention_mask and returns logits. Build the input the same way as above:

# pip install onnxruntime huggingface_hub
import onnxruntime as ort
from huggingface_hub import hf_hub_download

sess = ort.InferenceSession(hf_hub_download(MODEL, "onnx/model.onnx"))
enc = tok(context_for(passages, question), response, truncation="only_first", max_length=8192,
          return_tensors="np")
logits = sess.run(["logits"], {"input_ids": enc["input_ids"], "attention_mask": enc["attention_mask"]})[0]

How well it works

On RAGTruth, a public benchmark of news summaries and question answering with human-labeled errors:

threshold unsupported responses caught flags that are correct
0.975 39% 83%
0.5 63% 70%

We fine-tuned this model from LettuceDetect, an open-source detector trained on RAGTruth. It matches LettuceDetect on RAGTruth overall, and it's much better at the two things Tessary needs:

at a 2% false-positive rate this model LettuceDetect
Contradicted sentences caught, RAGTruth 24% 8%
Unsupported sentences caught, product and support responses 57% 4%

The product and support responses come from a set Tessary built and never trained on. We don't publish it.

Limitations

  • It misses many contradictions. A wrong number or name is usually caught. A policy stated the wrong way round often isn't. With a document saying "Monthly plans can be canceled at any time but are not refunded", the response "Yes, monthly plans are refunded within 30 days" scores 0.02.
  • It only knows what you pass in. A fact the agent got from a tool is BASELESS unless the tool's output is one of the passages. The model doesn't check anything against world knowledge.
  • It drops documents past 8,192 tokens without warning. The text is cut from the end of the documents, so a claim supported only by the dropped text is flagged as BASELESS.
  • Heavy paraphrase lowers accuracy. When a response rewords the documents a lot, correct sentences can score high. On a set built this way, it catches only 5% of unsupported sentences at a 2% false-positive rate. Set thresholds for each domain you run it on.
  • English text only. It hasn't been tested on other languages or on structured data such as tables and JSON.

How we trained it

We started from LettuceDetect large (MIT license), which is built on ModernBERT large (Apache 2.0 license). We gave it a third label so it can tell CONFLICT from BASELESS, then fine-tuned it on 8,683 responses from the RAGTruth training set (MIT license) and 1,370 product, support, and policy examples that Tessary generated. training_provenance.json records the training settings.

Feedback

Open a discussion on this page's Community tab, or open an issue on GitHub.

Acknowledgments

This model builds on:

Downloads last month
116
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tessaryai/groundedness-classifier-v1

Dataset used to train tessaryai/groundedness-classifier-v1

Papers for tessaryai/groundedness-classifier-v1