Instructions to use textsightai/textsight-detector-v23-custom with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use textsightai/textsight-detector-v23-custom with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="textsightai/textsight-detector-v23-custom")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("textsightai/textsight-detector-v23-custom") model = AutoModelForSequenceClassification.from_pretrained("textsightai/textsight-detector-v23-custom", device_map="auto") - Notebooks
- Google Colab
- Kaggle
TextSight detector v23
DeBERTa-v3-large fine-tuned for binary machine-generated-text detection on
English prose. 435M parameters, 512-token limit, labels {0: Human, 1: AI}.
This card reports measured numbers with the measurement code attached. It also documents where the model fails, because those failures determine whether it is safe to use for your task.
Everything below is reproducible with textsight/textsight-detector-benchmark.
Use the logit margin, not the softmax probability
This is the single most important thing on this page.
The model is confident enough that softmax saturates: it returns 0.99999999, and any rounding for display collapses that to 1.0. On 2,520 RAID documents, 77.1% of scores became exact ties. A threshold cannot separate documents that have been rounded onto the same value.
| score used | distinct values | exact ties | accuracy@5%FPR | AUROC |
|---|---|---|---|---|
| rounded probability | 578 / 2520 | 77.1% | 0.5273 | 0.8336 |
| log-odds margin | 2520 / 2520 | 0.0% | 0.7144 | 0.8426 |
AUROC barely moves, which is the tell: the ranking was always there, and rounding was hiding it. In two domains accuracy went from zero to ~0.72, purely because their thresholds had been pinned at the ceiling.
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
MODEL = "textsightai/textsight-detector-v23-custom"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForSequenceClassification.from_pretrained(MODEL).eval()
def score(texts):
"""Returns log-odds. Higher = more likely machine-generated.
Unbounded and does not saturate - use this for ranking and thresholds."""
enc = tok(texts, return_tensors="pt", truncation=True,
max_length=512, padding=True)
with torch.no_grad():
logits = model(**enc).logits.float()
return (logits[:, 1] - logits[:, 0]).tolist() # 1 = AI, 0 = Human
If you need a probability for display, derive it from the margin with a calibration fitted on your own data. Do not use the raw softmax as a calibrated confidence β it is not one.
Normalise the input before scoring
These are only weights. They carry no input handling, and a case-sensitive subword tokenizer is fragile to perturbations that leave text looking identical to a reader: a Cyrillic "Π°" for a Latin "a", a zero-width space between letters, a flipped capital mid-word. Each produces tokens the model never saw in training.
Normalising first β NFKC, Cyrillic/Greek confusable folding, invisible-character removal, whitespace and case repair β recovers most of that. Measured across nine mechanical perturbations, mean accuracy 0.6087 β 0.7014, with clean accuracy unchanged. Four of the nine drop to exactly zero effect.
A ready implementation is in
benchmark/normalize.py.
Benchmarking these weights on raw input measures tokenizer brittleness as much
as model judgement.
False positives are not one number
Fit a single threshold at 5% false positives across human text, then look at what it does per domain (RAID human documents, 150 per domain):
| domain | FPR | domain | FPR | |
|---|---|---|---|---|
| abstracts | 0.0% | reviews | 0.0% | |
| books | 0.0% | wiki | 15.3% | |
| news | 0.0% | recipes | 24.7% | |
| poetry | 0.0% | pooled | 5.0% | |
| 0.0% |
The pooled figure is 5.0% by construction. Behind it, one domain has a quarter of its genuinely human documents flagged. Formulaic, list-like prose β recipes, reference entries, structured instructions β is close to what models like this learn to call machine-written, and human authors of such text pay for it.
If your inputs are formulaic, this model's false-positive rate for you is several times any headline number.
Limitations
English only. Measured on non-English human text, this model assigns high machine-generated scores across the board (61β97% on the samples tested). Do not use it on non-English input; it does not abstain, it answers confidently and wrongly.
Short text is unreliable. False-positive rate rises sharply as input shortens. Treat anything under a few hundred characters as uninformative.
Adversarially fragile without normalisation. See above. With normalisation, perturbations that alter real words rather than their encoding still degrade accuracy, and synonym substitution and paraphrase β the two strongest attacks β have not been measured here at all. Treat published robustness figures for this model as an upper bound.
Single generator families. RAID covers 11 generators; performance on models outside that set is unmeasured.
Intended use
Suitable as one signal among several in content triage, moderation pipelines, and research on machine-generated text detection.
Not suitable as sole or primary evidence in any consequential decision about a person β academic misconduct, employment, admissions, or publication. The per-domain false-positive rates above are the reason: a system that flags a quarter of human-written documents in some genres cannot carry that weight, and no detector currently can. Anyone acting on an output should be able to see the score, the input length, and the genre, and should have a route to contest it.
Reproducing these numbers
git clone https://github.com/textsight/textsight-detector-benchmark
cd textsight-detector-benchmark
pip install -r requirements.txt
python -m benchmark.run --fetch
python -m benchmark.run --detector textsight-v23-normalised
Figures above: 2,520 documents stratified from RAID's labelled train_none
split, seed 0, 150 human documents per domain, per-domain thresholds at 5% FPR.
Thresholds are in-sample and therefore optimistic. RAID's human text is
web-scraped and does not represent any particular product's users. The
perturbations are independent implementations of RAID's published attack ideas,
not RAID's code, so the numbers are not comparable to the RAID leaderboard.
model.safetensors is 1,740,304,440 bytes, sha256
9177c456a91e41afc6e03583d81a63a5b81d5d0a96ecfd783cd5b8c234418443.
Related
- Benchmark and measurement code: textsight/textsight-detector-benchmark
- RAID: raid-bench.xyz
- Downloads last month
- 98
Model tree for textsightai/textsight-detector-v23-custom
Base model
microsoft/deberta-v3-large