cs.CLSep 5, 2026

SinoGlyphBench: A Diagnostic Benchmark for Chinese Glyph-Level Obfuscation in Language-Model Moderation

Authors: Yifan Wang, Zimu Wang, Suliu Qin, Changyu Zeng, Tong Chen, Siqi Chen, Yijie Lin, Lingyu Jiang, +5 more

Organizations: East China Normal University · New York University Shanghai · Xi’an Jiaotong-Liverpool University · University of Liverpool · Singapore University of Technology and Design · Eastern Institute of Technology (Ningbo)

Abstract

Glyph-level obfuscation can leave harmful Chinese content readable to humans while degrading automated moderation. We introduce SinoGlyphBench, a diagnostic benchmark that identifies label-critical semantic anchors and creates matched original and glyph-obfuscated inputs in text and image modalities. By perturbing anchors, background context, or both, this design distinguishes corruption of moderation-relevant evidence from general surface variation. Across 176,916 paired evaluations of 12 LLMs and MLLMs, obfuscation increases harmful false-negative and false-positive rates by 6.1 and 4.7 percentage points, respectively, and reduces four-way accuracy by 5.0 points. Models retain 75.7% of the decisions that were correct on the matched original inputs. Full-scope perturbations cause the largest degradation, anchor-only perturbations are more damaging than background-only perturbations, and cross-script substitution is particularly difficult in the text modality. Analysis of structured outputs identifies observable mismatches in visible-form reading, intended-message recovery, and final safety judgment. The evaluated models, therefore, remain brittle to Chinese content written with non-canonical glyphs. Resources are available at https://github.com/fengshun124/SinoGlyphBench.

Explore similar work

CardsList
  1. Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting

    May 21, 2026Jingyi Kang, Junyu Lu, Bo Xu +4Language Model Safety EvaluationToxicity Detection