Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Existing digital typography methods mainly focus on single-character scenarios. They suffer from a lack of legibility constraints and insufficient local deformation when extended to multi-character words, as the intricate structures among multiple characters are hardly preserved during the typography process. In this paper, we propose a global-to-local typography framework for multi-character scenarios. It performs mask-driven silhouette approximation at the global level, while semantic-guided refinement at the local level, with a culling step in between to improve efficiency. To preserve word legibility, we designed structural losses (including explicit collision constraints and implicit Jacobian singular value constraints) and an OCR constraint for character-level readability. To enhance the object recognizability, we leverage semantic guidance with diffusion priors, which drives the character glyph toward the target concept while preserving its structural integrity. To the best of our knowledge, this is the first multi-character semantic typography method that effectively balances word legibility and object recognizability. Evaluations on five representative languages (English, Chinese, Japanese, Korean, Arabic) demonstrate superiority over SOTA methods. Codes will be open-sourced.
Figures & tables
Figure 1: Results of our method with five different languages. From top to bottom are initial masks, intermediate results with global deformation and final typography results. From left to right are foxes in Chinese and Korean, sharks in English, bunnies in Japanese and camels in Arabic.
Figure 2: The artists’ work, (a) English ( Osotspa Co. 1998 ) , (b) Chinese ( Zhu 2016 ) , (c) Japanese ( Hudejii 2020 ) and (d) Korean ( Kyriazi 2021 ) .
Figure 3: The overview of our global-to-local semantic typography framework. At the global level, PCA-based initial placement, linear/Bézier deformations optimized by mask-filling, layout, and Jacobian losses. A culling step then filters candidates by concave-hull IoU. At the local level, semantic guidance for detail refinement, with OCR, collision detection, and ARAP constraints to preserve legibility and local rigidity.
Figure 4: Comparison with SOTA methods (WAI for Word-As-Image, DT for Dynamic Typography, OBI for OBI-Designer, NB-sp for Neural B - splines, GPT for GPT-Image-1.5) across 5 languages (English, Chinese, Japanese, Korean, Arabic).
Table 1: Quantitative evaluation of CLIP error ↓ (left) and OCR error ↓ ( ×10−3 , right).
Figure 5: Ablation study. (a)–(h) Global-level ablation: top row after global deformation, bottom row after full pipeline optimization. (i)–(n) Local-level ablation, showing final results.
Figure 6: User study results: mean scores and standard deviations across five methods and three criteria.