cs.CLSep 30, 2025

VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text

Authors: Trieu Hai Nguyen, Sivaswamy Akilesh

Organizations: Faculty of Information Technology, Nha Trang University, 02 Nguyen Dinh Chieu Street, North Nha Trang 57100, Vietnam · Faculty of Business Administration, Swiss School of Business and Management Geneva, 12 Avenue des Morgines, Geneva 1213, Switzerland

Abstract

The rapid proliferation of Large Language Models has intensified the challenge of distinguishing LLM-generated text from human writing in non-English languages. This study introduces VietBinoculars, a zero-shot detection framework coupling PhoGPT-4B observer and performer models with calibrated global decision thresholds. By utilizing specialized Vietnamese BPE tokenization, the method eliminates byte-level fragmentation and probability dilution common in massive multilingual backbones. Evaluated across multi-domain benchmarks, VietBinoculars achieves an area under the ROC curve exceeding 0.99. Under optimal Youden's J thresholds and greedy decoding, detection accuracy reaches at least 98.78%, while significantly outperforming baseline Binoculars, zero-shot detectors, and commercial tools on creative Capybara prompts. Even under a strict false positive rate constraint of 0.06%, the detector maintains F1-scores between 83.15% and 94.70%. Detection performance consistently improves with sequence length, stabilizing at optimal accuracy for passages containing 450 to 550 tokens. Extended stress testing across 48 distinct model-decoding configurations and three post-generation rewriting strategies delineates practical operational boundaries. VietBinoculars exhibits robust resilience against single-pass paraphrasing and human-style revisions, but experiences notable performance degradation under high-entropy sampling and iterative double paraphrasing.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model

    Apr 23, 2026Runheng Liu, Heyan Huang, Xingchen Xiao +1Machine-Generated Text DetectionZero-Shot

  2. Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration

    Aug 6, 2026Hongrui Bao, Yubing Ren, Yanan Cao +3Machine-Generated Text DetectionText Analysis and Detection

  3. EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

    Aug 11, 2026Hongrui Bao, Hangyu Rong, Zhuoshang Wang +2Machine-Generated Text DetectionLarge Language Model Evaluation