TextSight detector v23

DeBERTa-v3-large fine-tuned for binary machine-generated-text detection on English prose. 435M parameters, 512-token limit, labels {0: Human, 1: AI}.

This card reports measured numbers with the measurement code attached. It also documents where the model fails, because those failures determine whether it is safe to use for your task.

Everything below is reproducible with textsight/textsight-detector-benchmark.

Use the logit margin, not the softmax probability

This is the single most important thing on this page.

The model is confident enough that softmax saturates: it returns 0.99999999, and any rounding for display collapses that to 1.0. On 2,520 RAID documents, 77.1% of scores became exact ties. A threshold cannot separate documents that have been rounded onto the same value.

score used distinct values exact ties accuracy@5%FPR AUROC
rounded probability 578 / 2520 77.1% 0.5273 0.8336
log-odds margin 2520 / 2520 0.0% 0.7144 0.8426

AUROC barely moves, which is the tell: the ranking was always there, and rounding was hiding it. In two domains accuracy went from zero to ~0.72, purely because their thresholds had been pinned at the ceiling.

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

MODEL = "textsightai/textsight-detector-v23-custom"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForSequenceClassification.from_pretrained(MODEL).eval()

def score(texts):
    """Returns log-odds. Higher = more likely machine-generated.
    Unbounded and does not saturate - use this for ranking and thresholds."""
    enc = tok(texts, return_tensors="pt", truncation=True,
              max_length=512, padding=True)
    with torch.no_grad():
        logits = model(**enc).logits.float()
    return (logits[:, 1] - logits[:, 0]).tolist()   # 1 = AI, 0 = Human

If you need a probability for display, derive it from the margin with a calibration fitted on your own data. Do not use the raw softmax as a calibrated confidence β€” it is not one.

Normalise the input before scoring

These are only weights. They carry no input handling, and a case-sensitive subword tokenizer is fragile to perturbations that leave text looking identical to a reader: a Cyrillic "Π°" for a Latin "a", a zero-width space between letters, a flipped capital mid-word. Each produces tokens the model never saw in training.

Normalising first β€” NFKC, Cyrillic/Greek confusable folding, invisible-character removal, whitespace and case repair β€” recovers most of that. Measured across nine mechanical perturbations, mean accuracy 0.6087 β†’ 0.7014, with clean accuracy unchanged. Four of the nine drop to exactly zero effect.

A ready implementation is in benchmark/normalize.py. Benchmarking these weights on raw input measures tokenizer brittleness as much as model judgement.

False positives are not one number

Fit a single threshold at 5% false positives across human text, then look at what it does per domain (RAID human documents, 150 per domain):

domain FPR domain FPR
abstracts 0.0% reviews 0.0%
books 0.0% wiki 15.3%
news 0.0% recipes 24.7%
poetry 0.0% pooled 5.0%
reddit 0.0%

The pooled figure is 5.0% by construction. Behind it, one domain has a quarter of its genuinely human documents flagged. Formulaic, list-like prose β€” recipes, reference entries, structured instructions β€” is close to what models like this learn to call machine-written, and human authors of such text pay for it.

If your inputs are formulaic, this model's false-positive rate for you is several times any headline number.

Limitations

English only. Measured on non-English human text, this model assigns high machine-generated scores across the board (61–97% on the samples tested). Do not use it on non-English input; it does not abstain, it answers confidently and wrongly.

Short text is unreliable. False-positive rate rises sharply as input shortens. Treat anything under a few hundred characters as uninformative.

Adversarially fragile without normalisation. See above. With normalisation, perturbations that alter real words rather than their encoding still degrade accuracy, and synonym substitution and paraphrase β€” the two strongest attacks β€” have not been measured here at all. Treat published robustness figures for this model as an upper bound.

Single generator families. RAID covers 11 generators; performance on models outside that set is unmeasured.

Intended use

Suitable as one signal among several in content triage, moderation pipelines, and research on machine-generated text detection.

Not suitable as sole or primary evidence in any consequential decision about a person β€” academic misconduct, employment, admissions, or publication. The per-domain false-positive rates above are the reason: a system that flags a quarter of human-written documents in some genres cannot carry that weight, and no detector currently can. Anyone acting on an output should be able to see the score, the input length, and the genre, and should have a route to contest it.

Reproducing these numbers

git clone https://github.com/textsight/textsight-detector-benchmark
cd textsight-detector-benchmark
pip install -r requirements.txt
python -m benchmark.run --fetch
python -m benchmark.run --detector textsight-v23-normalised

Figures above: 2,520 documents stratified from RAID's labelled train_none split, seed 0, 150 human documents per domain, per-domain thresholds at 5% FPR. Thresholds are in-sample and therefore optimistic. RAID's human text is web-scraped and does not represent any particular product's users. The perturbations are independent implementations of RAID's published attack ideas, not RAID's code, so the numbers are not comparable to the RAID leaderboard.

model.safetensors is 1,740,304,440 bytes, sha256 9177c456a91e41afc6e03583d81a63a5b81d5d0a96ecfd783cd5b8c234418443.

Related

Downloads last month
98
Safetensors
Model size
0.4B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for textsightai/textsight-detector-v23-custom

Finetuned
(309)
this model