DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
Organizations: CHI, Chair of Health Informatics, Technical University of Munich, Germany · Smart Embedded Systems Lab, OTH Regensburg, Germany · GLAM, Group on Language, Audio & Music, Imperial College London, UK
Abstract
Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Feature | Description |
| Vocabulary (8) | |
| vocab_zipf_coef | Zipf law coefficient |
| vocab_yules_k | Yule’s K richness measure |
| vocab_hapax | Hapax legomena ratio |
| vocab_dis | Dis legomena ratio |
| vocab_heaps | Heaps’ law exponent |
| Version | TPR@5 % | TPR@1 % | AUROC | Change |
| Reported (Sec. 5.2 ) | 99.01 % | 96.57 % | 99.61 % | deberta-v3-large , no features, raw training, normalised inference, 1.55M rows |
| Full RAID (Sec. 5.4 ) | 99.13 % | 97.12 % | 99.63 % | Same configuration on the complete RAID split (9.38M rows); loses 3.42 pp TPR@1 % FPR on MAGE |
| v1 | 94.29 % | 87.36 % | 98.58 % | deberta-v3-base , feature fusion |
| v2 | 97.11 % | 93.34 % | 97.38 % | Attack-diverse training data |
| v2.2 | 97.65 % | 93.99 % | 97.87 % | Cohere balancing, revised RAID validation |
| v2.2-Seed42 | 97.29 % | 93.55 % | 96.86 % | v2.2, second seed |