deberta-conpara / README.md
mohamedmady's picture
Upload README.md
a5f896f verified
|
Raw History Blame Contribute Delete
4.98 kB
---
license: mit
language: en
base_model: microsoft/deberta-v3-large
tags:
- ai-generated-text-detection
- machine-generated-text
- deberta
- raid
- adversarial-robustness
pipeline_tag: text-classification
---
# DeBERTa-ConPara
DeBERTa-ConPara detects machine-generated English text and stays accurate under the
adversarial edits that defeat most detectors: homoglyph substitution, zero-width
character insertion, whitespace and typographic attacks.
It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The
finding behind it: normalising the **training** corpus deduplicates it (35.4% of
RAID rows collapse into byte-identical copies of their clean siblings, deleting
the adversarial supervision), while normalising at **inference** is an effective
defence. DeBERTa-ConPara trains on raw text and normalises only at inference.
- **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU →
Dropout(0.1) → Linear(512, 2). No feature branch.
- **Training data:** 1.55M documents, leakage-free stratified splits over RAID,
HC3 Plus, MAGE and M4, grouped by source id.
- **Output:** logit margin, `logit[AI] − logit[human]`. Higher is more
machine-like. Not a probability.
## Results
RAID hidden test (672,000 documents, 11 attacks): **AUROC 99.61**,
**TPR@5% FPR 99.01**, **TPR@1% FPR 96.57**.
Cross-dataset, balanced accuracy at one threshold calibrated on our source
validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23,
M4 98.27.
The README of the [GitHub repository](https://github.com/SES-Lab-OTH/deberta-conpara)
compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint,
and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara
on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being
external data for the other systems.
## Usage
The checkpoint is a plain PyTorch state dict with a small custom head, so it does
not load through `AutoModelForSequenceClassification`. Use the loader from the
repository:
```python
from src.conpara import ConPara # pip install -r requirements.txt
det = ConPara.from_pretrained() # pulls rawguard.pt from this repo
det.score(["a document to check"]) # logit margin
det.predict(["a document to check"]) # bool at the stored threshold
```
Inference-time Unicode normalisation is applied by default; it is what makes the
detector robust to homoglyph and zero-width attacks.
## Intended use and limits
Intended as **supporting evidence for a human decision**: flagging text for
review, studying detector behaviour, benchmarking. Not intended as a verdict, and
not suitable for disciplinary or hiring decisions about individuals.
Measured limits:
- **Short text:** below ~60 words errors are enriched 4–5×. Treat under 75 words
with caution; under 25 words the score is meaningless.
- **Academic prose:** false-positive rates between 13% and 67% depending on the
subcorpus. A flag on a student essay is not evidence of misconduct.
- **Calibration:** the score is not a probability and saturates at the extremes.
Prefer coarse bands (`det.band()`) over percentages.
- **Language:** English only.
- **Drift:** trained against generators available in 2026; newer ones are
untested, and detectors decay as generators improve.
- **Attacks not covered:** heavy paraphrasing by a strong model, and mixed
human/machine documents, remain hard.
## Training and evaluation protocol
Full corpus construction, the leakage-free splitting, the 2×2×2 factorial and the
fixed-threshold evaluation protocol are in the repository. Every number above is
reproducible from `evaluation/`.
## Citation
```bibtex
@inproceedings{mady2026conpara,
title = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic
Detection of {AI}-Generated Text},
author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
booktitle = {Proceedings of the 14th International Joint Conference on Natural
Language Processing and the 4th Conference of the Asia-Pacific Chapter
of the Association for Computational Linguistics (AACL-IJCNLP 2026)},
year = {2026},
publisher = {Association for Computational Linguistics},
eprint = {2610.00883},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.00883}
}
```
License: MIT, matching the DeBERTa-v3-large backbone.
## Links
- Code and evaluation scripts: https://github.com/SES-Lab-OTH/deberta-conpara
- Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara
- Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026, [arXiv:2610.00883](https://arxiv.org/abs/2610.00883)
- Lab: [Smart Embedded Systems Lab, OTH Regensburg](https://github.com/SES-Lab-OTH)