mohamedmady commited on
Commit
7d26533
·
verified ·
1 Parent(s): 79e0d67

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +104 -0
README.md ADDED
@@ -0,0 +1,104 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language: en
4
+ base_model: microsoft/deberta-v3-large
5
+ tags:
6
+ - ai-generated-text-detection
7
+ - machine-generated-text
8
+ - deberta
9
+ - raid
10
+ - adversarial-robustness
11
+ pipeline_tag: text-classification
12
+ ---
13
+
14
+ # RawGuard (DeBERTa-ConPara)
15
+
16
+ RawGuard detects machine-generated English text and stays accurate under the
17
+ adversarial edits that defeat most detectors: homoglyph substitution, zero-width
18
+ character insertion, whitespace and typographic attacks.
19
+
20
+ It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The
21
+ finding behind it: normalising the **training** corpus deduplicates it — 35.4% of
22
+ RAID rows collapse into byte-identical copies of their clean siblings, deleting
23
+ the adversarial supervision — while normalising at **inference** is an effective
24
+ defence. RawGuard trains on raw text and normalises only at inference.
25
+
26
+ - **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU →
27
+ Dropout(0.1) → Linear(512, 2). No feature branch.
28
+ - **Training data:** 1.56M documents, leakage-free stratified splits over RAID,
29
+ HC3 Plus, MAGE and M4, grouped by source id.
30
+ - **Output:** logit margin, `logit[AI] − logit[human]`. Higher is more
31
+ machine-like. Not a probability.
32
+
33
+ ## Results
34
+
35
+ RAID hidden test (672,000 documents, 11 attacks): **AUROC 99.61**,
36
+ **TPR@5% FPR 99.01**, **TPR@1% FPR 96.57**.
37
+
38
+ Cross-dataset, balanced accuracy at one threshold calibrated on our source
39
+ validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23,
40
+ M4 98.27.
41
+
42
+ The README of the [GitHub repository](https://github.com/MohamedMady19/deberta-conpara)
43
+ compares RawGuard with every RAID leaderboard system that publishes a checkpoint,
44
+ and states the two caveats those numbers need: MELD scores higher than RawGuard
45
+ on RAID itself, and HC3/MAGE/M4 are training sources for RawGuard while being
46
+ external data for the other systems.
47
+
48
+ ## Usage
49
+
50
+ The checkpoint is a plain PyTorch state dict with a small custom head, so it does
51
+ not load through `AutoModelForSequenceClassification`. Use the loader from the
52
+ repository:
53
+
54
+ ```python
55
+ from src.rawguard import RawGuard # pip install -r requirements.txt
56
+
57
+ det = RawGuard.from_pretrained() # pulls rawguard.pt from this repo
58
+ det.score(["a document to check"]) # logit margin
59
+ det.predict(["a document to check"]) # bool at the stored threshold
60
+ ```
61
+
62
+ Inference-time Unicode normalisation is applied by default; it is what makes the
63
+ detector robust to homoglyph and zero-width attacks.
64
+
65
+ ## Intended use and limits
66
+
67
+ Intended as **supporting evidence for a human decision** — flagging text for
68
+ review, studying detector behaviour, benchmarking. Not intended as a verdict, and
69
+ not suitable for disciplinary or hiring decisions about individuals.
70
+
71
+ Measured limits:
72
+
73
+ - **Short text:** below ~60 words errors are enriched 4–5×. Treat under 75 words
74
+ with caution; under 25 words the score is meaningless.
75
+ - **Academic prose:** false-positive rates between 13% and 67% depending on the
76
+ subcorpus. A flag on a student essay is not evidence of misconduct.
77
+ - **Calibration:** the score is not a probability and saturates at the extremes.
78
+ Prefer coarse bands (`det.band()`) over percentages.
79
+ - **Language:** English only.
80
+ - **Drift:** trained against generators available in 2026; newer ones are
81
+ untested, and detectors decay as generators improve.
82
+ - **Attacks not covered:** heavy paraphrasing by a strong model, and mixed
83
+ human/machine documents, remain hard.
84
+
85
+ ## Training and evaluation protocol
86
+
87
+ Full corpus construction, the leakage-free splitting, the 2×2×2 factorial and the
88
+ fixed-threshold evaluation protocol are in the repository. Every number above is
89
+ reproducible from `evaluation/`.
90
+
91
+ ## Citation
92
+
93
+ ```bibtex
94
+ @inproceedings{mady2026rawguard,
95
+ title = {Where You Normalise Matters: Unicode Preprocessing and the
96
+ Robustness of AI-Generated Text Detection},
97
+ author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
98
+ booktitle = {Proceedings of AACL-IJCNLP 2026},
99
+ year = {2026},
100
+ note = {To appear}
101
+ }
102
+ ```
103
+
104
+ License: MIT, matching the DeBERTa-v3-large backbone.