cs.CVSep 29, 2026

Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering

Authors: Yazhen Xie, Xingsong Ye, Zhineng Chen

Organizations: Institute of Trustworthy Embodied AI, Fudan University · Shanghai Key Laboratory of Multimodal Embodied AI

Abstract

Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.

Figures & tables

Explore similar work

CardsList
  1. TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

    Date pendingMingxuan Cui, Jingpu Yang, Fengxian Ji +7Image-Text AlignmentFew-Shot Font Generation

  2. Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers

    Jul 4, 2026Shuyang Xiang, Hao GuanToken EmbeddingsTransformer Architectures

  3. Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering

    Jun 1, 2026Dongxing Mao, Jinpeng Wang, Jiahao Tang +6Autoregressive Image GenerationVisual Tokenizers