|
Download README.md from SES-Lab-OTH/deberta-conpara: direct link, hf CLI and curl.
- Browser
- Download file 4.98 kB
-
https://huggingface.co/SES-Lab-OTH/deberta-conpara/resolve/main/README.md
- Command line
-
hf download hf://SES-Lab-OTH/deberta-conpara/README.md
-
curl -L -o README.md https://huggingface.co/SES-Lab-OTH/deberta-conpara/resolve/main/README.md
4.98 kB
| license: mit | |
| language: en | |
| base_model: microsoft/deberta-v3-large | |
| tags: | |
| - ai-generated-text-detection | |
| - machine-generated-text | |
| - deberta | |
| - raid | |
| - adversarial-robustness | |
| pipeline_tag: text-classification | |
| # DeBERTa-ConPara | |
| DeBERTa-ConPara detects machine-generated English text and stays accurate under the | |
| adversarial edits that defeat most detectors: homoglyph substitution, zero-width | |
| character insertion, whitespace and typographic attacks. | |
| It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The | |
| finding behind it: normalising the **training** corpus deduplicates it (35.4% of | |
| RAID rows collapse into byte-identical copies of their clean siblings, deleting | |
| the adversarial supervision), while normalising at **inference** is an effective | |
| defence. DeBERTa-ConPara trains on raw text and normalises only at inference. | |
| - **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU → | |
| Dropout(0.1) → Linear(512, 2). No feature branch. | |
| - **Training data:** 1.55M documents, leakage-free stratified splits over RAID, | |
| HC3 Plus, MAGE and M4, grouped by source id. | |
| - **Output:** logit margin, `logit[AI] − logit[human]`. Higher is more | |
| machine-like. Not a probability. | |
| ## Results | |
| RAID hidden test (672,000 documents, 11 attacks): **AUROC 99.61**, | |
| **TPR@5% FPR 99.01**, **TPR@1% FPR 96.57**. | |
| Cross-dataset, balanced accuracy at one threshold calibrated on our source | |
| validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23, | |
| M4 98.27. | |
| The README of the [GitHub repository](https://github.com/SES-Lab-OTH/deberta-conpara) | |
| compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint, | |
| and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara | |
| on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being | |
| external data for the other systems. | |
| ## Usage | |
| The checkpoint is a plain PyTorch state dict with a small custom head, so it does | |
| not load through `AutoModelForSequenceClassification`. Use the loader from the | |
| repository: | |
| ```python | |
| from src.conpara import ConPara # pip install -r requirements.txt | |
| det = ConPara.from_pretrained() # pulls rawguard.pt from this repo | |
| det.score(["a document to check"]) # logit margin | |
| det.predict(["a document to check"]) # bool at the stored threshold | |
| ``` | |
| Inference-time Unicode normalisation is applied by default; it is what makes the | |
| detector robust to homoglyph and zero-width attacks. | |
| ## Intended use and limits | |
| Intended as **supporting evidence for a human decision**: flagging text for | |
| review, studying detector behaviour, benchmarking. Not intended as a verdict, and | |
| not suitable for disciplinary or hiring decisions about individuals. | |
| Measured limits: | |
| - **Short text:** below ~60 words errors are enriched 4–5×. Treat under 75 words | |
| with caution; under 25 words the score is meaningless. | |
| - **Academic prose:** false-positive rates between 13% and 67% depending on the | |
| subcorpus. A flag on a student essay is not evidence of misconduct. | |
| - **Calibration:** the score is not a probability and saturates at the extremes. | |
| Prefer coarse bands (`det.band()`) over percentages. | |
| - **Language:** English only. | |
| - **Drift:** trained against generators available in 2026; newer ones are | |
| untested, and detectors decay as generators improve. | |
| - **Attacks not covered:** heavy paraphrasing by a strong model, and mixed | |
| human/machine documents, remain hard. | |
| ## Training and evaluation protocol | |
| Full corpus construction, the leakage-free splitting, the 2×2×2 factorial and the | |
| fixed-threshold evaluation protocol are in the repository. Every number above is | |
| reproducible from `evaluation/`. | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{mady2026conpara, | |
| title = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic | |
| Detection of {AI}-Generated Text}, | |
| author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.}, | |
| booktitle = {Proceedings of the 14th International Joint Conference on Natural | |
| Language Processing and the 4th Conference of the Asia-Pacific Chapter | |
| of the Association for Computational Linguistics (AACL-IJCNLP 2026)}, | |
| year = {2026}, | |
| publisher = {Association for Computational Linguistics}, | |
| eprint = {2610.00883}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CL}, | |
| url = {https://arxiv.org/abs/2610.00883} | |
| } | |
| ``` | |
| License: MIT, matching the DeBERTa-v3-large backbone. | |
| ## Links | |
| - Code and evaluation scripts: https://github.com/SES-Lab-OTH/deberta-conpara | |
| - Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara | |
| - Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026, [arXiv:2610.00883](https://arxiv.org/abs/2610.00883) | |
| - Lab: [Smart Embedded Systems Lab, OTH Regensburg](https://github.com/SES-Lab-OTH) | |