TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity
Authors: Xi Cao, Quzong Gesang, Yuan Sun, Nuo Qun, Tashi Nyima
Organizations: Minzu University of China, Beijing, China · National Language Resource Monitoring & Research Center Minority Languages Branch, Beijing, China · Tibet University, Lhasa, China · Collaborative Innovation Center for Tibet Informatization by Ministry of Education & Tibet Autonomous Region, Lhasa, China
Language models based on deep neural networks are vulnerable to textual adversarial attacks. While rich-resource languages like English are receiving focused attention, Tibetan, a cross-border language, is gradually being studied due to its abundant ancient literature and critical language strategy. Currently, there are several Tibetan adversarial text generation methods, but they do not fully consider the textual features of Tibetan script and overestimate the quality of generated adversarial texts. To address this issue, we propose a novel Tibetan adversarial text generation method called TSCheater, which considers the characteristic of Tibetan encoding and the feature that visually similar syllables have similar semantics. This method can also be transferred to other abugidas, such as Devanagari script. We utilize a self-constructed Tibetan syllable visual similarity database called TSVSDB to generate substitution candidates and adopt a greedy algorithm-based scoring mechanism to determine substitution order. After that, we conduct the method on eight victim language models. Experimentally, TSCheater outperforms existing methods in attack effectiveness, perturbation magnitude, semantic similarity, visual similarity, and human acceptance. Finally, we construct the first Tibetan adversarial robustness evaluation benchmark called AdvTS, which is generated by existing methods and proofread by humans.
Figures & tables
Syllable
CCOEFF_NORMED
Syllable
CCOEFF_NORMED
0.9806
0.9370
0.9705
0.9352
0.9690
0.9312
0.9510
0.9168
0.9388
…
TABLE I: Tibetan Syllables Visually Similar to
Metric
Method
TNCC-title
TU_SA
Tibetan-BERT
CINO-small
CINO-base
CINO-large
Tibetan-BERT
CINO-small
CINO-base
CINO-large
ADV ↑
TSAttacker
0.3420
0.3592
0.3646
0.3430
0.1570
0.2260
0.2240
0.2660
TSTricker-s
0.5124
0.5685
0.5414
0.5426
0.3080
0.4300
0.4730
0.5060
TSTricker-w
0.5124
0.5588
0.5566
0.5286
0.2870
0.4050
0.4200
0.5100
TSCheater-s
0.4714
0.5717
0.5620
0.5329
0.2810
0.3790
0.4390
0.4280
TSCheater-w
0.5027
0.5696
0.5696
0.5405
0.2810
0.3770
0.4540
0.4260
TABLE II: Experimental Results bold and underlined values represent the best performance bold values represent the second best performance
TABLE III: Composition Information of AdvTS bold and underlined values represent the largest part
Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work presents, to the best of our knowledge, the first large-model-based Tibetan TTS system in the industry, built upon a large speech synthesis model developed by Xingchen AGI Lab. The proposed system integrates data quality enhancement, Tibetan-oriented text representation and tokenizer adaptation, and cross-lingual adaptive training for low-resource Tibetan speech synthesis. Experimental results show that the system can generate stable, natural, and intelligible Tibetan speech under low-resource conditions. In subjective evaluation, the MOS scores of the syllable-level and BPE-based systems reach 4.28 and 4.35, while their pronunciation accuracies reach 97.6% and 96.6%, respectively, outperforming an external commercial Tibetan TTS interface. These results demonstrate that combining a large-model backbone with Tibetan-oriented text representation adaptation and cross-lingual adaptive training enables highly usable low-resource Tibetan speech synthesis, and also provides a technical foundation for future unified multi-dialect Tibetan speech synthesis.
Jiaxu He, Chao Wang, Jie Lian +4
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology Co. Ltd · Qinghai Normal University, Xining, China · University of Electronic Science and Technology of China,Chengdu,China +1
Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastructure. To fill this gap, we introduce FTibSuite, a comprehensive resource suite for Tibetan vision-language research, consisting of FTibData (human-verified multimodal training corpora spanning continual pretraining, image-text alignment, and instruction tuning data), FTibBench (Tibetan adaptations of five mainstream multimodal benchmarks with a hierarchical quality-control workflow to reduce translation noise), and FTibVLM, a reproducible baseline built on Qwen3-VL-8B-Instruct via a three-stage adaptation pipeline. Experiments on FTibBench show FTibVLM delivers consistent performance gains across all tasks, such as improving MMBench accuracy from 42.97 to 67.78 and POPE-random accuracy from 47.53 to 80.56, while retaining the backbone's original Chinese capabilities with minimal degradation, providing the first standardized foundation for Tibetan multimodal research.
Guixian Xu, Yide Liang, Zeli Su +5
Hainan International College, Minzu University of China · School of Information Engineering, Minzu University of China · Institute of Automation, Chinese Academy of Sciences +1
Transformer-based transfer models now dominate Bangla sentiment classification, yet their adversarial robustness remains largely unexamined, and no prior study pairs a Bangla attack suite with a defense that measurably recovers robustness. We address this gap with destroR, a unified pipeline for evaluating and hardening Bangla text classifiers. First, we introduce three meaning-preserving Bangla attack recipes a paraphrase attack, a back-translation attack, and a one-hot word-swap attack that perturb inputs while regenerating fluent, semantically faithful sentences, inducing model prediction perplexity rather than input noise. Second, we construct a robustness benchmark that evaluates five transfer models (BanglaBERT, BanglishBERT, XLM-RoBERTa, MuRIL, and IndicBERTv2) across four datasets against five attacks, placing our recipes against two strong word-substitution baselines, TextFooler and BAE, under an identical protocol. Third, we harden every model through adversarial training and report a full robustness matrix. Our analysis yields three findings: word-substitution baselines are more potent than semantically constrained recipes (BAE reaches a 54.2% attack success rate); adversarial training on the union of all attack families lowers residual attack success for every attack; and, contrary to expectation, the Indic-multilingual MuRIL backbone is markedly more robust than the Bangla-dedicated models. All models, adversarial data, and code are released for full reproducibility.