cs.CLDec 3, 2024

TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity

Authors: Xi Cao, Quzong Gesang, Yuan Sun, Nuo Qun, Tashi Nyima

Organizations: Minzu University of China, Beijing, China · National Language Resource Monitoring & Research Center Minority Languages Branch, Beijing, China · Tibet University, Lhasa, China · Collaborative Innovation Center for Tibet Informatization by Ministry of Education & Tibet Autonomous Region, Lhasa, China

Abstract

Language models based on deep neural networks are vulnerable to textual adversarial attacks. While rich-resource languages like English are receiving focused attention, Tibetan, a cross-border language, is gradually being studied due to its abundant ancient literature and critical language strategy. Currently, there are several Tibetan adversarial text generation methods, but they do not fully consider the textual features of Tibetan script and overestimate the quality of generated adversarial texts. To address this issue, we propose a novel Tibetan adversarial text generation method called TSCheater, which considers the characteristic of Tibetan encoding and the feature that visually similar syllables have similar semantics. This method can also be transferred to other abugidas, such as Devanagari script. We utilize a self-constructed Tibetan syllable visual similarity database called TSVSDB to generate substitution candidates and adopt a greedy algorithm-based scoring mechanism to determine substitution order. After that, we conduct the method on eight victim language models. Experimentally, TSCheater outperforms existing methods in attack effectiveness, perturbation magnitude, semantic similarity, visual similarity, and human acceptance. Finally, we construct the first Tibetan adversarial robustness evaluation benchmark called AdvTS, which is generated by existing methods and proofread by humans.

Figures & tables

Explore similar work

CardsList
  1. Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

    May 4, 2026Jiaxu He, Chao Wang, Jie Lian +4Large Language Model AdaptationSynthesis

  2. FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

    May 26, 2026Guixian Xu, Yide Liang, Zeli Su +5Recent Vision-Language ModelsMultimodal Benchmarks

  3. destroR: A Benchmark and Adversarial-Training Defense for Bangla Transfer Models under Meaning-Preserving Attacks

    Nov 13, 2025Saadat Rafid Ahmed, Rubayet Shareen, Radoan Sharkar +3BanglaSentiment