Organizations: National Yang Ming Chiao Tung University, Taiwan · E.SUN Financial Holding Co., Ltd., Taiwan · National Kaohsiung Normal University, Taiwan
Zero-shot Chinese character recognition (ZS-CCR) aims to recognize characters whose categories are never observed during training, and typically relies on the compositional structure shared between seen and unseen characters. Recent CLIP-style methods represent this structure with the Ideographic Description Sequence (IDS) and align it with glyph images in a shared embedding space. However, they rely on a single global image--IDS similarity that discards the spatial layout of radicals and, being learned only implicitly from seen classes, generalizes poorly to unseen ones; moreover, global matching often retrieves the correct character within the top candidates yet fails to rank it first when characters differ only in subtle local radicals. To address these issues, we propose a global-to-local two-stage framework. In the first stage, STG-CLIP augments the IDS with explicit tree-position and radical-level geometric priors, yielding a spatial-aware prototype that provides a consistent spatial description across seen and unseen categories for high-recall global retrieval. In the second stage, the Radical Verification Module (RVM) uses the radical instances of each retrieved candidate as queries to verify whether the corresponding radicals can be matched to spatially compatible regions in the input glyph. A margin-based gating rule activates the RVM only when the leading global candidates receive similar similarity scores. Experiments on the ICDAR2013 benchmark demonstrate that our method achieves state-of-the-art performance under the character-level zero-shot setting, obtaining 83.06% top-1 accuracy with 2,755 seen classes. Ablation studies further show that the explicit geometric priors and radical-level verification provide complementary improvements.
Figures & tables
Figure 1 : Overview of our framework. STG-CLIP (Left): the first stage learns a shared image–text space by constructing tree- and geometry-aware IDS prototypes. RVM (Top-right): the RVM uses the radical occurrences of each candidate as queries to verify whether they are grounded in the input image. Two-stage Inference (Bottom-right): STG-CLIP first retrieves the Top- K candidates by global similarity; the RVM then verifies the radicals of each candidate against the input glyph and selects the final prediction when the leading candidates are globally ambiguous.
Character Zero-shot
Radical Zero-shot
Method
500
1000
1500
2000
2755
50
40
30
20
10
DenseRAN [ 9 ]
1.70%
8.44%
14.71%
19.51%
30.68%
0.21%
0.29%
0.25%
0.42%
0.69%
HDE [ 2 ]
4.90%
12.77%
19.25%
25.13%
33.49%
3.26%
4.29%
6.33%
7.64%
9.33%
SD [ 3 ]
5.60%
13.85%
22.88%
25.73%
37.91%
5.28%
6.87%
9.02%
14.67%
15.83%
STAR [ 12 ]
7.54%
19.47%
27.79%
35.53%
43.86%
6.95%
12.28%
14.74%
18.37%
23.23%
FaRE [ 13 ]
7.21%
21.78%
36.58%
47.33%
57.17%
-
-
-
-
-
Table 1 : Character-level and radical-level zero-shot recognition accuracy on handwritten datasets under different numbers of seen training classes.
Cumulative setting
Top-1 Acc.
Δ
Baseline
76.27
–
+ IDS Prototype Encoding
77.78
+1.51
+ Radical Geometry Augmentation
82.26
+4.48
+ RVM
83.06
+0.80
Table 2 : Cumulative ablation study of the proposed components under the character zero-shot setting with 2,755 seen classes. Each row adds one component to the previous setting.
Figure 2 : Hyperparameter selection for shortlist size K and gating threshold δ on the validation set (character zero-shot setting). The optimal configuration ( K=2,δ=0.03 ) is fixed at test time.
Chinese character categories are extremely large, and unseen characters frequently arise in open-world scenarios, making zero-shot Chinese character recognition an important yet challenging problem. Existing IDS-based retrieval methods usually encode a character image and its ideographic description sequence into a single global vector for matching. Although efficient, such holistic alignment often under-models local component differences. Moreover, directly introducing patch-token level fine-grained interaction suffers from both the noise of structural operators in IDS and the high cost of full-candidate retrieval.To address these issues, we propose a Global-Local Hierarchical Perception Network (GL-HPN), which jointly learns global and local representations of character images and IDS sequences within a unified cross-modal alignment framework. The global branch supports efficient coarse recall, while the local branch improves component-level discrimination through patch-token interaction. We further introduce a structure filtering mask to suppress structurally meaningful but visually non-entity IDS operators in local similarity aggregation. On top of this, we design a coarse-to-fine hierarchical inference strategy that performs global retrieval over the full candidate set and local reranking only on Top-K candidates, followed by parameter-free multiplicative fusion of normalized posterior scores. Experimental results show that GL-HPN achieves competitive performance across multiple zero-shot splits, performs especially well under low-resource settings, and substantially reduces the inference cost of large-scale candidate retrieval.
Wei Cao, Hao Xu, Xiaolei Diao
1Jilin University, China · University College London, UK
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.
Yazhen Xie, Xingsong Ye, Zhineng Chen
Institute of Trustworthy Embodied AI, Fudan University · Shanghai Key Laboratory of Multimodal Embodied AI
Zero-shot recognition aims to classify an image by selecting the most compatible label description from a set of candidate classes without any task-specific supervision. In fine-grained settings, however, the relevant evidence often lies in localized parts, attributes, or textures rather than in the full image, making whole-image alignment suboptimal. Recent localized visual-text alignment methods address this by comparing class descriptions with multiple image regions, but they typically rely on large sets of random or redundant crops, increasing inference cost and introducing many highly redundant or weakly relevant candidates. Moreover, introducing semantic guidance too early can create an error-amplifying feedback process in which inaccurate intermediate predictions bias later localization and reinforce subsequent mistakes; we refer to this failure mode as the prediction loop. We propose LAGO (LAnguage-Guided adaptive Object-region focus), a framework for efficient and robust zero-shot localized visual-text alignment. LAGO first performs class-agnostic object-centric candidate discovery to obtain a stable visual initialization, and then applies adaptive language-guided refinement with the strength of semantic guidance controlled by intermediate confidence. It further combines object-level, contextual, and full-image evidence through an effective object-context dual-channel aggregation strategy. Extensive experiments show that LAGO consistently achieves state-of-the-art performance on standard zero-shot benchmarks and challenging distribution-shift settings, while requiring substantially fewer candidate regions at inference time.
Junyi Hu, Qiji Zhou, Lei Zhang +1
Beijing Jiaotong University · Westlake University · Rochester Institute of Technology