Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,104 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
language: en
|
| 4 |
+
base_model: microsoft/deberta-v3-large
|
| 5 |
+
tags:
|
| 6 |
+
- ai-generated-text-detection
|
| 7 |
+
- machine-generated-text
|
| 8 |
+
- deberta
|
| 9 |
+
- raid
|
| 10 |
+
- adversarial-robustness
|
| 11 |
+
pipeline_tag: text-classification
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# RawGuard (DeBERTa-ConPara)
|
| 15 |
+
|
| 16 |
+
RawGuard detects machine-generated English text and stays accurate under the
|
| 17 |
+
adversarial edits that defeat most detectors: homoglyph substitution, zero-width
|
| 18 |
+
character insertion, whitespace and typographic attacks.
|
| 19 |
+
|
| 20 |
+
It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The
|
| 21 |
+
finding behind it: normalising the **training** corpus deduplicates it — 35.4% of
|
| 22 |
+
RAID rows collapse into byte-identical copies of their clean siblings, deleting
|
| 23 |
+
the adversarial supervision — while normalising at **inference** is an effective
|
| 24 |
+
defence. RawGuard trains on raw text and normalises only at inference.
|
| 25 |
+
|
| 26 |
+
- **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU →
|
| 27 |
+
Dropout(0.1) → Linear(512, 2). No feature branch.
|
| 28 |
+
- **Training data:** 1.56M documents, leakage-free stratified splits over RAID,
|
| 29 |
+
HC3 Plus, MAGE and M4, grouped by source id.
|
| 30 |
+
- **Output:** logit margin, `logit[AI] − logit[human]`. Higher is more
|
| 31 |
+
machine-like. Not a probability.
|
| 32 |
+
|
| 33 |
+
## Results
|
| 34 |
+
|
| 35 |
+
RAID hidden test (672,000 documents, 11 attacks): **AUROC 99.61**,
|
| 36 |
+
**TPR@5% FPR 99.01**, **TPR@1% FPR 96.57**.
|
| 37 |
+
|
| 38 |
+
Cross-dataset, balanced accuracy at one threshold calibrated on our source
|
| 39 |
+
validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23,
|
| 40 |
+
M4 98.27.
|
| 41 |
+
|
| 42 |
+
The README of the [GitHub repository](https://github.com/MohamedMady19/deberta-conpara)
|
| 43 |
+
compares RawGuard with every RAID leaderboard system that publishes a checkpoint,
|
| 44 |
+
and states the two caveats those numbers need: MELD scores higher than RawGuard
|
| 45 |
+
on RAID itself, and HC3/MAGE/M4 are training sources for RawGuard while being
|
| 46 |
+
external data for the other systems.
|
| 47 |
+
|
| 48 |
+
## Usage
|
| 49 |
+
|
| 50 |
+
The checkpoint is a plain PyTorch state dict with a small custom head, so it does
|
| 51 |
+
not load through `AutoModelForSequenceClassification`. Use the loader from the
|
| 52 |
+
repository:
|
| 53 |
+
|
| 54 |
+
```python
|
| 55 |
+
from src.rawguard import RawGuard # pip install -r requirements.txt
|
| 56 |
+
|
| 57 |
+
det = RawGuard.from_pretrained() # pulls rawguard.pt from this repo
|
| 58 |
+
det.score(["a document to check"]) # logit margin
|
| 59 |
+
det.predict(["a document to check"]) # bool at the stored threshold
|
| 60 |
+
```
|
| 61 |
+
|
| 62 |
+
Inference-time Unicode normalisation is applied by default; it is what makes the
|
| 63 |
+
detector robust to homoglyph and zero-width attacks.
|
| 64 |
+
|
| 65 |
+
## Intended use and limits
|
| 66 |
+
|
| 67 |
+
Intended as **supporting evidence for a human decision** — flagging text for
|
| 68 |
+
review, studying detector behaviour, benchmarking. Not intended as a verdict, and
|
| 69 |
+
not suitable for disciplinary or hiring decisions about individuals.
|
| 70 |
+
|
| 71 |
+
Measured limits:
|
| 72 |
+
|
| 73 |
+
- **Short text:** below ~60 words errors are enriched 4–5×. Treat under 75 words
|
| 74 |
+
with caution; under 25 words the score is meaningless.
|
| 75 |
+
- **Academic prose:** false-positive rates between 13% and 67% depending on the
|
| 76 |
+
subcorpus. A flag on a student essay is not evidence of misconduct.
|
| 77 |
+
- **Calibration:** the score is not a probability and saturates at the extremes.
|
| 78 |
+
Prefer coarse bands (`det.band()`) over percentages.
|
| 79 |
+
- **Language:** English only.
|
| 80 |
+
- **Drift:** trained against generators available in 2026; newer ones are
|
| 81 |
+
untested, and detectors decay as generators improve.
|
| 82 |
+
- **Attacks not covered:** heavy paraphrasing by a strong model, and mixed
|
| 83 |
+
human/machine documents, remain hard.
|
| 84 |
+
|
| 85 |
+
## Training and evaluation protocol
|
| 86 |
+
|
| 87 |
+
Full corpus construction, the leakage-free splitting, the 2×2×2 factorial and the
|
| 88 |
+
fixed-threshold evaluation protocol are in the repository. Every number above is
|
| 89 |
+
reproducible from `evaluation/`.
|
| 90 |
+
|
| 91 |
+
## Citation
|
| 92 |
+
|
| 93 |
+
```bibtex
|
| 94 |
+
@inproceedings{mady2026rawguard,
|
| 95 |
+
title = {Where You Normalise Matters: Unicode Preprocessing and the
|
| 96 |
+
Robustness of AI-Generated Text Detection},
|
| 97 |
+
author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
|
| 98 |
+
booktitle = {Proceedings of AACL-IJCNLP 2026},
|
| 99 |
+
year = {2026},
|
| 100 |
+
note = {To appear}
|
| 101 |
+
}
|
| 102 |
+
```
|
| 103 |
+
|
| 104 |
+
License: MIT, matching the DeBERTa-v3-large backbone.
|