File size: 4,983 Bytes
7d26533
 
 
 
 
 
 
 
 
 
 
 
 
dc01729
7d26533
dc01729
7d26533
 
 
 
a5f896f
7d26533
a5f896f
dc01729
7d26533
 
 
dc01729
7d26533
 
 
 
 
 
 
 
 
 
 
 
 
a5f896f
dc01729
 
 
7d26533
 
 
 
 
 
 
 
 
dc01729
7d26533
dc01729
7d26533
 
 
 
 
 
 
 
 
a5f896f
7d26533
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dc01729
a5f896f
 
 
 
 
 
 
 
 
 
 
 
7d26533
 
 
 
dc01729
 
 
a5f896f
dc01729
a5f896f
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
---
license: mit
language: en
base_model: microsoft/deberta-v3-large
tags:
  - ai-generated-text-detection
  - machine-generated-text
  - deberta
  - raid
  - adversarial-robustness
pipeline_tag: text-classification
---

# DeBERTa-ConPara

DeBERTa-ConPara detects machine-generated English text and stays accurate under the
adversarial edits that defeat most detectors: homoglyph substitution, zero-width
character insertion, whitespace and typographic attacks.

It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The
finding behind it: normalising the **training** corpus deduplicates it (35.4% of
RAID rows collapse into byte-identical copies of their clean siblings, deleting
the adversarial supervision), while normalising at **inference** is an effective
defence. DeBERTa-ConPara trains on raw text and normalises only at inference.

- **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU →
  Dropout(0.1) → Linear(512, 2). No feature branch.
- **Training data:** 1.55M documents, leakage-free stratified splits over RAID,
  HC3 Plus, MAGE and M4, grouped by source id.
- **Output:** logit margin, `logit[AI] − logit[human]`. Higher is more
  machine-like. Not a probability.

## Results

RAID hidden test (672,000 documents, 11 attacks): **AUROC 99.61**,
**TPR@5% FPR 99.01**, **TPR@1% FPR 96.57**.

Cross-dataset, balanced accuracy at one threshold calibrated on our source
validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23,
M4 98.27.

The README of the [GitHub repository](https://github.com/SES-Lab-OTH/deberta-conpara)
compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint,
and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara
on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being
external data for the other systems.

## Usage

The checkpoint is a plain PyTorch state dict with a small custom head, so it does
not load through `AutoModelForSequenceClassification`. Use the loader from the
repository:

```python
from src.conpara import ConPara          # pip install -r requirements.txt

det = ConPara.from_pretrained()           # pulls rawguard.pt from this repo
det.score(["a document to check"])         # logit margin
det.predict(["a document to check"])       # bool at the stored threshold
```

Inference-time Unicode normalisation is applied by default; it is what makes the
detector robust to homoglyph and zero-width attacks.

## Intended use and limits

Intended as **supporting evidence for a human decision**: flagging text for
review, studying detector behaviour, benchmarking. Not intended as a verdict, and
not suitable for disciplinary or hiring decisions about individuals.

Measured limits:

- **Short text:** below ~60 words errors are enriched 4–5×. Treat under 75 words
  with caution; under 25 words the score is meaningless.
- **Academic prose:** false-positive rates between 13% and 67% depending on the
  subcorpus. A flag on a student essay is not evidence of misconduct.
- **Calibration:** the score is not a probability and saturates at the extremes.
  Prefer coarse bands (`det.band()`) over percentages.
- **Language:** English only.
- **Drift:** trained against generators available in 2026; newer ones are
  untested, and detectors decay as generators improve.
- **Attacks not covered:** heavy paraphrasing by a strong model, and mixed
  human/machine documents, remain hard.

## Training and evaluation protocol

Full corpus construction, the leakage-free splitting, the 2×2×2 factorial and the
fixed-threshold evaluation protocol are in the repository. Every number above is
reproducible from `evaluation/`.

## Citation

```bibtex
@inproceedings{mady2026conpara,
  title         = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic
                   Detection of {AI}-Generated Text},
  author        = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
  booktitle     = {Proceedings of the 14th International Joint Conference on Natural
                   Language Processing and the 4th Conference of the Asia-Pacific Chapter
                   of the Association for Computational Linguistics (AACL-IJCNLP 2026)},
  year          = {2026},
  publisher     = {Association for Computational Linguistics},
  eprint        = {2610.00883},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2610.00883}
}
```

License: MIT, matching the DeBERTa-v3-large backbone.

## Links

- Code and evaluation scripts: https://github.com/SES-Lab-OTH/deberta-conpara
- Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara
- Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026, [arXiv:2610.00883](https://arxiv.org/abs/2610.00883)
- Lab: [Smart Embedded Systems Lab, OTH Regensburg](https://github.com/SES-Lab-OTH)