Ouroboros

Ouroboros detector (v7)

A token-level detector of AI-written text. For a document it returns character-aligned segments labelled human, ai-assisted or ai-generated with confidences, a weighted AI fraction, a human / mixed / ai verdict, and a 4-way "was this humanized?" probability (human, ai-generated, humanized-ai, mixed-authorship).

Unofficial. This is an independent, open reimplementation of the method of the Pangram 4 technical report (arXiv:2607.27183). It is not affiliated with or endorsed by Pangram, and it is not their model, weights or data.

Built end to end by Claude Opus 5.5. The architecture implementation, data pipeline, training, evaluation, interpretability study, bias audit and this model card were written and run by Claude Opus 5.5 (Anthropic), working as an autonomous coding agent, at the direction of the author, who set the goals and reviewed the results.

  • Code, training pipeline, dataset builder, demo: https://github.com/Jourdelune/ouroboros-detector
  • Backbone: Qwen3-1.7B-Base with a LoRA adapter (r = 32) merged into the weights (bf16, 3.4 GB), plus four single-layer heads.
  • Languages: evaluated on English and French; other languages appear in the training data but are not benchmarked.

Use

pip install "git+https://github.com/Jourdelune/ouroboros-detector.git"
from ouroboros.infer.predict import Predictor

detector = Predictor.from_pretrained("Jour/ouroboros-detector")
result = detector.predict(text)                  # texts of at least ~50 words
print(result.document_label, result.weighted_ai_fraction)
for seg in result.segments:
    print(seg.start, seg.end, seg.label_name, seg.confidence)
print(detector.humanizer(text))

The model needs the ouroboros package: it implements Repeat2 inputs, the four heads, sliding windows, calibration and CRF decoding. AutoModel.from_pretrained alone loads only the backbone.

Architecture

One causal backbone fed the window twice (Repeat2, so every supervised token has seen the whole window) and four heads: segment (D→15, weighted AI fraction), tokenwise provenance (D→3), mixed-authorship (D→2), humanizer probe (D→4, stop-gradient). Windows of 512 tokens with stride 256, then calibration and a linear-chain CRF.

Architecture

Evaluation

Held-out test split of the project's corpus: 12,000 documents (4,531 human, 6,821 AI, 648 mixed), split by source document. Operating point chosen on a separate calibration split for a 0.5 % false-positive target.

metric v7
False-positive rate 1.35 %
False-negative rate (AI or mixed missed) 0.94 %
AUROC 0.9915
TPR @ 1 % FPR 98.9 %
Token accuracy 98.4 %
Token F1 — human / AI-assisted / AI-generated 0.986 / 0.714 / 0.988
Mixed-document accuracy 77.2 %
4-way humanizer head accuracy / recall on humanized-ai 95.0 % / 56.8 %
Benchmarks

These numbers are in-distribution (same corpus families as training) and not comparable to the Pangram report's, which uses a different, proprietary corpus and backbone.

Known limitations and biases

Known biases
  • Higher false-positive rate on French human text (7.0 % vs 1.0 % in English), on very short text (< 80 words: 6.5 %), and on human text formatted like an assistant answer (bullets / bold / headers: 8.8 %).
  • Format and encoding shortcuts: human text split into paragraphs is flagged 14.5 % of the time; inserting four Cyrillic look-alike letters into a human text makes it read as AI in 86 % of cases. Normalise confusable characters before inference.
  • Not robust to adaptive adversaries: a rewriting loop with access to the detector's score made 70–88 % of a small held-out set (50 French + 50 English passages) read as human while passing automatic fidelity checks.
  • Not supported: texts under ~50 words. ai-assisted is the weakest class (F1 0.71).
  • Do not use a verdict as the sole basis for a decision about a person. AI-text detectors can wrongly flag non-native writers and formal prose.

What it looks at (interpretability)

Measured on held-out passages with 95 % intervals: a linear probe on the untrained Qwen3-1.7B-Base already separates classes at 95 %; fine-tuning raises it to 98.5 % and makes it language-independent (French→English transfer 99.5 % vs 91.5 %). The verdict is written in layers 18–27, mostly by two attention heads and the last MLP; sentence-final periods (2.4 % of tokens) carry 28 % of the evidence and paragraph breaks (0.8 %) 21 %; shuffling words inside sentences removes most of the signal while surface edits barely move it; the score is uncorrelated with a base language model's surprisal.

Interpretability

Training data

Public sources only: RAID, MAGE, COLING-2025 MGT, ai-text-detection-pile, dmitva/human_ai_generated_text, Cosmopedia, WildChat, UltraChat, OpenHermes-2.5, arena preference sets, Amazon/Yelp/IMDB reviews, arXiv abstracts, FineWeb-2 (French) and FineWeb-Edu, Aya, Wikipedia, French instruction sets and several public distillation sets from recent frontier models (≈ 1.09 M training documents including augmentations). The dataset builder is in the GitHub repository.

Licences. The weights are released under Apache-2.0 and the backbone is Apache-2.0, but the training data comes from many datasets with their own terms (some research-only or non-commercial). Check them before any commercial use. No dataset is redistributed here.

Files

model.safetensors, config.json (merged backbone) · heads.pt (the four heads) · calibrator.json (decoder calibration, 0.5 % FPR target) · ouroboros_config.yaml (inference settings) · tokenizer files.

Citation

@software{ouroboros_detector,
  title  = {Ouroboros: an open token-level AI-text detector},
  author = {Jourdelune},
  year   = {2026},
  note   = {Unofficial reimplementation of the Pangram 4 technical report (arXiv:2607.27183); built end to end with Claude Opus 5.5},
  url    = {https://github.com/Jourdelune/ouroboros-detector}
}
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jour/ouroboros-detector

Adapter
(78)
this model

Space using Jour/ouroboros-detector 1

Paper for Jour/ouroboros-detector