--- license: mit language: en base_model: microsoft/deberta-v3-large tags: - ai-generated-text-detection - machine-generated-text - deberta - raid - adversarial-robustness pipeline_tag: text-classification --- # DeBERTa-ConPara DeBERTa-ConPara detects machine-generated English text and stays accurate under the adversarial edits that defeat most detectors: homoglyph substitution, zero-width character insertion, whitespace and typographic attacks. It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The finding behind it: normalising the **training** corpus deduplicates it (35.4% of RAID rows collapse into byte-identical copies of their clean siblings, deleting the adversarial supervision), while normalising at **inference** is an effective defence. DeBERTa-ConPara trains on raw text and normalises only at inference. - **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU → Dropout(0.1) → Linear(512, 2). No feature branch. - **Training data:** 1.55M documents, leakage-free stratified splits over RAID, HC3 Plus, MAGE and M4, grouped by source id. - **Output:** logit margin, `logit[AI] − logit[human]`. Higher is more machine-like. Not a probability. ## Results RAID hidden test (672,000 documents, 11 attacks): **AUROC 99.61**, **TPR@5% FPR 99.01**, **TPR@1% FPR 96.57**. Cross-dataset, balanced accuracy at one threshold calibrated on our source validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23, M4 98.27. The README of the [GitHub repository](https://github.com/SES-Lab-OTH/deberta-conpara) compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint, and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being external data for the other systems. ## Usage The checkpoint is a plain PyTorch state dict with a small custom head, so it does not load through `AutoModelForSequenceClassification`. Use the loader from the repository: ```python from src.conpara import ConPara # pip install -r requirements.txt det = ConPara.from_pretrained() # pulls rawguard.pt from this repo det.score(["a document to check"]) # logit margin det.predict(["a document to check"]) # bool at the stored threshold ``` Inference-time Unicode normalisation is applied by default; it is what makes the detector robust to homoglyph and zero-width attacks. ## Intended use and limits Intended as **supporting evidence for a human decision**: flagging text for review, studying detector behaviour, benchmarking. Not intended as a verdict, and not suitable for disciplinary or hiring decisions about individuals. Measured limits: - **Short text:** below ~60 words errors are enriched 4–5×. Treat under 75 words with caution; under 25 words the score is meaningless. - **Academic prose:** false-positive rates between 13% and 67% depending on the subcorpus. A flag on a student essay is not evidence of misconduct. - **Calibration:** the score is not a probability and saturates at the extremes. Prefer coarse bands (`det.band()`) over percentages. - **Language:** English only. - **Drift:** trained against generators available in 2026; newer ones are untested, and detectors decay as generators improve. - **Attacks not covered:** heavy paraphrasing by a strong model, and mixed human/machine documents, remain hard. ## Training and evaluation protocol Full corpus construction, the leakage-free splitting, the 2×2×2 factorial and the fixed-threshold evaluation protocol are in the repository. Every number above is reproducible from `evaluation/`. ## Citation ```bibtex @inproceedings{mady2026conpara, title = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic Detection of {AI}-Generated Text}, author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.}, booktitle = {Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL-IJCNLP 2026)}, year = {2026}, publisher = {Association for Computational Linguistics}, eprint = {2610.00883}, archivePrefix = {arXiv}, primaryClass = {cs.CL}, url = {https://arxiv.org/abs/2610.00883} } ``` License: MIT, matching the DeBERTa-v3-large backbone. ## Links - Code and evaluation scripts: https://github.com/SES-Lab-OTH/deberta-conpara - Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara - Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026, [arXiv:2610.00883](https://arxiv.org/abs/2610.00883) - Lab: [Smart Embedded Systems Lab, OTH Regensburg](https://github.com/SES-Lab-OTH)