DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
Authors: Mohamed Mady, Yupei Li, Johannes Reschke, Björn W. Schuller
Organizations: CHI, Chair of Health Informatics, Technical University of Munich, Germany · Smart Embedded Systems Lab, OTH Regensburg, Germany · GLAM, Group on Language, Audio & Music, Imperial College London, UK
Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.
Figures & tables
Figure 1: DeBERTa-ConPara architecture. Solid components form the reported configuration: attack-aware preprocessing applied at inference only, a DeBERTa-v3-large encoder, and a two-layer head on the classification embedding. Dashed components are the linguistic feature branch and its gated fusion, one factor of the 2×2×2 ablation in Section 5 , which is negative in every condition and is not part of the reported model.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Feature
Description
Vocabulary (8)
vocab_zipf_coef
Zipf law coefficient
vocab_yules_k
Yule’s K richness measure
vocab_hapax
Hapax legomena ratio
vocab_dis
Dis legomena ratio
vocab_heaps
Heaps’ law exponent
Appendix
Table 10: Complete list of all 62 handcrafted features with descriptions, organised by category. Left: Perplexity, Entropy, Burstiness, Repetition (38 features). Right: Vocabulary, Coherence, Readability, Stylometric (28 features). The top k=30 are selected via Mutual Information (Section A.2.2 ).
Figure 2: Oracle gap, the balanced accuracy given up by fixing τ⋆ in advance rather than tuning per target, across 24 cell-benchmark pairs. The cost is under one point everywhere except HC3-SI. Per-cell sweep curves are in Figure 3 .
Figure 3: Balanced accuracy as the decision threshold is swept across its range, for each cell on each benchmark. Markers show where the frozen τ⋆ falls. HC3-QA and M4 give broad plateaux, so the frozen threshold lands almost anywhere without cost; HC3-SI is sharply peaked for every cell, which is where the oracle gap comes from.
Version
TPR@5 %
TPR@1 %
AUROC
Change
Reported (Sec. 5.2 )
99.01 %
96.57 %
99.61 %
deberta-v3-large , no features, raw training, normalised inference, 1.55M rows
Full RAID (Sec. 5.4 )
99.13 %
97.12 %
99.63 %
Same configuration on the complete RAID split (9.38M rows); loses 3.42 pp TPR@1 % FPR on MAGE
v1
94.29 %
87.36 %
98.58 %
deberta-v3-base , feature fusion
v2
97.11 %
93.34 %
97.38 %
Attack-diverse training data
v2.2
97.65 %
93.99 %
97.87 %
Cohere balancing, revised RAID validation
v2.2-Seed42
97.29 %
93.55 %
96.86 %
v2.2, second seed
Appendix
Table 17: RAID hidden-test results of the configuration we report, of the same configuration trained on the complete RAID split, and of the development sequence behind the submitted version (all on deberta-v3-base ). TPR at 5 % and 1 % FPR. Bold marks the configuration we report and, within the development sequence, the best value in each column. A1 and A3 differ only by a train/inference mismatch; the controlled comparison of the two preprocessing placements is the eight-cell factorial of Section 5.2 .
Recent AI-generated text detection work often introduces a new benchmark together with a specialized detector tailored to it. We revisit this practice from a baseline-first perspective. Across several benchmarks, we show that a plain, fully fine-tuned RoBERTa matches or exceeds the specialized detectors those benchmarks are built around. This suggests that much of the recent architectural complexity is not what drives strong in-distribution detection. The remaining challenge is the distribution shift. The same strong baseline degrades sharply when the topic domain or generating model changes at test time, and simply adding more source data does not close the gap. We identify a key failure mode: under distribution shift, the detector can assign high-confidence machine labels to human-written text from unseen domains. We then study two lightweight domain adaptation methods to address this problem: K-shot adaptation with first-order MAML over LoRA adapters, and a per-sample confidence-weighted ensemble built on top of the adapted detector. Overall, our results suggest that progress in AI-generated text detection should be measured not only by in-distribution performance, but also by robustness under distribution shift.
Zhuoer Shen, Mingyi Wang, Shaofeng Zou +1
Department of Computer Science, University of California, Santa Barbara · School of ECEE, Arizona State University
AI-generated text is nowadays produced at scale across domains and heterogeneous generation pipelines, making robustness to distribution shift a central requirement for supervised binary detectors. We train transformer-based detectors on HC3 PLUS and calibrate a single decision threshold by maximising balanced accuracy on held-out validation; this threshold is then kept fixed for all downstream test distributions, revealing domain- and generator-dependent error asymmetries under shift. We evaluate in-domain on HC3 PLUS, under cross-dataset transfer to the multi-domain, multi-generator M4 benchmark, and on the external AI-Text-Detection-Pile. Although base models achieve near-ceiling in-domain performance (up to 99.5% balanced accuracy), performance under shift is brittle and strongly model-dependent. Feature augmentation via attention-based linguistic feature fusion improves transfer, with our best model (DeBERTa-v3-base+FeatAttn) achieving 85.9% balanced accuracy on M4. Multi-seed experiments confirm high stability. Under the same fixed-threshold protocol, our model outperforms strong zero-shot baselines by up to +7.22 points. Category-level ablations further show that readability and vocabulary features contribute most to robustness under shift. Overall, these results demonstrate that feature augmentation and a modern DeBERTa backbone significantly outperform earlier BERT/RoBERTa models, while the fixed-threshold protocol provides a more realistic and informative assessment of practical detector robustness.
Mohamed Mady, Johannes Reschke, Björn Schuller
Chair of Health Informatics, Technical University of Munich (TUM), Germany · OTH Regensburg, Germany · GLAM – Group on Language, Audio & Music, Imperial College London, United Kingdom
Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel Triospect Detection Framework by using additional perspectives of content (core ideas) and expression (stylistic elements) within a given text. Experiments on two benchmarks involving 17 attacks, 12 domains, and 17 source models demonstrate that Triospect is robust against these attacks. It improves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K after-attack subset, and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAID. This framework marks a pioneering effort in statistical methods to enhance detection reliability against attacks. We release our data and code at https://github.com/baoguangsheng/triospect.
Guangsheng Bao, Lihua Rong, Yanbin Zhao +3
Zhejiang University · Zhejiang University of Technology · Shanghai Polytechnic University +2