cs.CLAug 10, 2026

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

Authors: Ghazal KalhorZahra JafariAmirarsalan ShahbaziBehnam Bahrak

Organizations: School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran · School of Engineering Science, College of Engineering, University of Tehran, Tehran, Iran · Tehran Institute for Advanced Studies, Khatam University, Tehran, Iran

Abstract

Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code-mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.

Explore similar work

CardsList
  1. MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

    Feb 25, 2026Kazi Samin Yasar Alam, Md Tanbir Chowdhury, Tamim Ahmed +2BanglaHumor