Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.
Figures & tables
Figure 1 : Local reward responses to a malformed ideograph. Character-level rewards produce discrete decisions, whereas IDSpect retains partial structural credit via IDS matching.
Figure 2 : Framework of IDSpect. An OCR detector extracts text crops from each generated image. The semantic branch compares OCR transcripts with the target text, while the IDS branch compares visually predicted and target IDS tokens for intra-character feedback. Their weighted sum guides GRPO post-training of Qwen-Image.
Rewards
LongText
GenTextEval
Avg.
Qua.
Sem.
Qua.
Sem.
Base
0.920
0.924
0.834
0.933
0.810
OCR
0.967
0.956
0.886
0.953
0.874
TextPecker
0.974
0.969
0.908
0.973
0.897
IDSpect
0.972
0.975
0.928
0.979
0.911
Table 1: Quantitative comparison of Qwen-Image variants on Chinese visual text rendering benchmarks. Base denotes the frozen model. OCR, TextPecker [ 32 ] , and IDSpect denote GRPO post-training with the corresponding rewards. Avg. : original benchmark text score, Qua. : structural quality, Sem. : semantic alignment. Qua. and Sem. are evaluated by TextPecker.
Figure 3 : Qualitative comparison of Qwen-Image and OCR-, TextPecker-, and IDSpect-guided GRPO on Chinese text-rendering prompts. The rightmost column lists the target text.
IDS reference
Comparison
Qua. ↑
Sem. ↑
OCR prediction
Per-crop consistency
0.980
0.901
Target text
Direct concatenation
0.976
0.886
Target text
Character-wise matching
0.971
0.899
Target text
Crop-wise alignment (ours)
0.979
0.911
Table 2: Ablation of IDS reward construction on GenTextEval. All variants use the same balanced IDS recognizer, semantic OCR term, and training configuration.
Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level structure. Prior methods often improve this ability through architecture-specific modules or encoder modifications, which complicate deployment across foundation models. We study text rendering as a post-training preference-alignment problem and propose TextAlign, a non-invasive framework that keeps the generator architecture unchanged. The key component is a hierarchical vision-language model (VLM)-based reward that decomposes rendering errors into global, word, and glyph levels, then converts binary defect judgments into a scalar preference signal. The resulting signal supports both Group Relative Policy Optimization (GRPO) and Direct Preference Optimization (DPO). Experiments on FLUX.1-dev and Z-Image-Turbo show consistent gains in OCR-based text accuracy without degrading general generation quality. Compared with strong foundation and text-rendering baselines, including SD3.5, Qwen-Image, AnyText, and TextDiffuser, these results indicate that reward design offers a scalable alternative to model redesign for improving text rendering.
Mingxuan Cui, Jingpu Yang, Fengxian Ji +7
Mohamed bin Zayed University of Artificial Intelligence · Chinese Academy of Sciences Institute of Automation · Northeastern University +1
Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice but rarely questioned. We challenge this representation, especially for Chinese, by replacing index-based token embeddings entirely with a single rasterized image of the character sequence, processed by a vision encoder composed of a shared ResNet and a shallow Vision Transformer. To isolate the role of input representation, we construct a dual-branch controlled framework in which both a Vision-based model and an index-based baseline share an identical decoder backbone, training objective, optimizer, and data curriculum. Any performance difference is therefore attributable to the input modality only. Across all tested decoder backbones, the Vision-based model consistently outperforms the baseline, reaching a peak accuracy of 0.429 versus 0.355 for the index-based baseline,that is, a 21% relative improvement, while converging in about half the number of training epochs. The advantage emerges especially within the first five epochs (under 21% of total data) and persists under moderate character corruption: the corrupted Vision model matches the clean index-based baseline. Ablation studies reveal that the advantage requires both spatially coherent input and a ViT encoder with 2D positional encodings. A cross-script comparison on English shows the advantage does not transfer directly to alphabetic writing systems, suggesting that the uniform visual density and radical structure of Chinese characters are enabling conditions. These findings suggest that transformers are more modality-agnostic than commonly assumed, and that discrete tokenization is not a fundamental requirement for Chinese language modeling.
Shuyang Xiang, Hao Guan
Independent Researcher. · Institute of Software, Chinese Academy of Sciences.
Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation ability, they still underperform on text rendering with blur strokes and disrupt letter shapes. In this work, we trace this limitation to the visual tokenizer, which struggles to reconstruct fine-grained detail. Improving the tokenizer is straightforward but expensive, as it necessitates retraining both the tokenizer and the AR model. Can we improve text rendering performance of AR models without retraining the existing tokenizer and AR model? To achieve this, we propose the Residual Decoder Adapter(RDA) that upgrades an existing tokenizer post-hoc without changing its token space. Specifically, it refines the decoder output of the visual tokenizer by introducing two novel components: (i) a paired codebook that shares the token distribution with the original one; (ii) a parallel branch to learn the tiny differences (residual) between the reconstructed image and the ground-truth images in the pixel space. This residual design allows us to enhance the tokenizer non-invasively while preserving compatibility with prior AR models. RDA substantially improves text rendering significantly by a large margin. For instance, we boost finetuned Janus-Pro OCR accuracy rises from 24.52% to 58.26% (TextVisionBlend), from 12.75% to 36.81% (StyledTextSynth) on competitive TextAtlas benchmark. The code is available at https://github.com/CSU-JPG/RDA
Dongxing Mao, Jinpeng Wang, Jiahao Tang +6
Central South University · University of Oxford · Microsoft Research