cs.CLOct 1, 2026

DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text

Authors: Mohamed Mady, Yupei Li, Johannes Reschke, Björn W. Schuller

Organizations: CHI, Chair of Health Informatics, Technical University of Munich, Germany · Smart Embedded Systems Lab, OTH Regensburg, Germany · GLAM, Group on Language, Audio & Music, Imperial College London, UK

Abstract

Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking AI-Generated Text Detection: A Strong Baseline and the Distribution-Shift Problem That Remains

    Jul 4, 2026Zhuoer Shen, Mingyi Wang, Shaofeng Zou +1Machine-Generated Text DetectionBert-Based Models

  2. Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators

    May 5, 2026Mohamed Mady, Johannes Reschke, Björn SchullerMachine-Generated Text DetectionBert-Based Models

  3. Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks

    Jun 30, 2026Guangsheng Bao, Lihua Rong, Yanbin Zhao +3Diverse Attacks