mohamedmady commited on
Commit
a5f896f
·
verified ·
1 Parent(s): 4d528ea

Upload README.md

Browse files
Files changed (1) hide show
  1. README.md +19 -12
README.md CHANGED
@@ -18,9 +18,9 @@ adversarial edits that defeat most detectors: homoglyph substitution, zero-width
18
  character insertion, whitespace and typographic attacks.
19
 
20
  It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The
21
- finding behind it: normalising the **training** corpus deduplicates it — 35.4% of
22
  RAID rows collapse into byte-identical copies of their clean siblings, deleting
23
- the adversarial supervision — while normalising at **inference** is an effective
24
  defence. DeBERTa-ConPara trains on raw text and normalises only at inference.
25
 
26
  - **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU →
@@ -39,7 +39,7 @@ Cross-dataset, balanced accuracy at one threshold calibrated on our source
39
  validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23,
40
  M4 98.27.
41
 
42
- The README of the [GitHub repository](https://github.com/MohamedMady19/deberta-conpara)
43
  compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint,
44
  and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara
45
  on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being
@@ -64,7 +64,7 @@ detector robust to homoglyph and zero-width attacks.
64
 
65
  ## Intended use and limits
66
 
67
- Intended as **supporting evidence for a human decision** — flagging text for
68
  review, studying detector behaviour, benchmarking. Not intended as a verdict, and
69
  not suitable for disciplinary or hiring decisions about individuals.
70
 
@@ -92,12 +92,18 @@ reproducible from `evaluation/`.
92
 
93
  ```bibtex
94
  @inproceedings{mady2026conpara,
95
- title = {Where You Normalise Matters: Unicode Preprocessing and the
96
- Robustness of AI-Generated Text Detection},
97
- author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
98
- booktitle = {Proceedings of AACL-IJCNLP 2026},
99
- year = {2026},
100
- note = {To appear}
 
 
 
 
 
 
101
  }
102
  ```
103
 
@@ -105,6 +111,7 @@ License: MIT, matching the DeBERTa-v3-large backbone.
105
 
106
  ## Links
107
 
108
- - Code and evaluation scripts: https://github.com/MohamedMady19/deberta-conpara
109
  - Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara
110
- - Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026 (arXiv link to follow)
 
 
18
  character insertion, whitespace and typographic attacks.
19
 
20
  It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The
21
+ finding behind it: normalising the **training** corpus deduplicates it (35.4% of
22
  RAID rows collapse into byte-identical copies of their clean siblings, deleting
23
+ the adversarial supervision), while normalising at **inference** is an effective
24
  defence. DeBERTa-ConPara trains on raw text and normalises only at inference.
25
 
26
  - **Architecture:** DeBERTa-v3-large → CLS token → Linear(1024, 512) → GELU →
 
39
  validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23,
40
  M4 98.27.
41
 
42
+ The README of the [GitHub repository](https://github.com/SES-Lab-OTH/deberta-conpara)
43
  compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint,
44
  and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara
45
  on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being
 
64
 
65
  ## Intended use and limits
66
 
67
+ Intended as **supporting evidence for a human decision**: flagging text for
68
  review, studying detector behaviour, benchmarking. Not intended as a verdict, and
69
  not suitable for disciplinary or hiring decisions about individuals.
70
 
 
92
 
93
  ```bibtex
94
  @inproceedings{mady2026conpara,
95
+ title = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic
96
+ Detection of {AI}-Generated Text},
97
+ author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
98
+ booktitle = {Proceedings of the 14th International Joint Conference on Natural
99
+ Language Processing and the 4th Conference of the Asia-Pacific Chapter
100
+ of the Association for Computational Linguistics (AACL-IJCNLP 2026)},
101
+ year = {2026},
102
+ publisher = {Association for Computational Linguistics},
103
+ eprint = {2610.00883},
104
+ archivePrefix = {arXiv},
105
+ primaryClass = {cs.CL},
106
+ url = {https://arxiv.org/abs/2610.00883}
107
  }
108
  ```
109
 
 
111
 
112
  ## Links
113
 
114
+ - Code and evaluation scripts: https://github.com/SES-Lab-OTH/deberta-conpara
115
  - Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara
116
+ - Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026, [arXiv:2610.00883](https://arxiv.org/abs/2610.00883)
117
+ - Lab: [Smart Embedded Systems Lab, OTH Regensburg](https://github.com/SES-Lab-OTH)