VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
Authors: Trieu Hai Nguyen, Sivaswamy Akilesh
Organizations: Faculty of Information Technology, Nha Trang University, 02 Nguyen Dinh Chieu Street, North Nha Trang 57100, Vietnam · Faculty of Business Administration, Swiss School of Business and Management Geneva, 12 Avenue des Morgines, Geneva 1213, Switzerland
The rapid proliferation of Large Language Models has intensified the challenge of distinguishing LLM-generated text from human writing in non-English languages. This study introduces VietBinoculars, a zero-shot detection framework coupling PhoGPT-4B observer and performer models with calibrated global decision thresholds. By utilizing specialized Vietnamese BPE tokenization, the method eliminates byte-level fragmentation and probability dilution common in massive multilingual backbones. Evaluated across multi-domain benchmarks, VietBinoculars achieves an area under the ROC curve exceeding 0.99. Under optimal Youden's J thresholds and greedy decoding, detection accuracy reaches at least 98.78%, while significantly outperforming baseline Binoculars, zero-shot detectors, and commercial tools on creative Capybara prompts. Even under a strict false positive rate constraint of 0.06%, the detector maintains F1-scores between 83.15% and 94.70%. Detection performance consistently improves with sequence length, stabilizing at optimal accuracy for passages containing 450 to 550 tokens. Extended stress testing across 48 distinct model-decoding configurations and three post-generation rewriting strategies delineates practical operational boundaries. VietBinoculars exhibits robust resilience against single-pass paraphrasing and human-style revisions, but experiences notable performance degradation under high-entropy sampling and iterative double paraphrasing.
Figures & tables
Figure 1: Illustration of log(perplexity) and log(cross-perplexity) . The x-axis represents the average next-token prediction probability Yˉ for the same input string s . The solid magenta curve represents logPPLM2(s) and the dashed curves represent logX-PPLM1,M2(s) for different average next-token prediction probabilities of the observer and performer models, denoted YM1 and YM2 , respectively. The dashed blue curve illustrates logX-PPLM1,M2(s) when YM1≈YM2 , i.e., the observer and performer models are nearly identical. When the two models differ substantially, i.e., YM1≫YM2 , the cross-perplexity rises sharply, as depicted by the dash-dotted red curve.
Domain
Dataset
Target
# Docs
Human Tokens/Doc
LLM Tokens/Doc
Mean ± SD
Median [IQR]
Mean ± SD
Median [IQR]
News articles 6
Sailor2-8B-OptiThreshold-News
In-Domain
82,690
466.0 ± 142.1
467 [391, 524]
575.8 ± 114.0
512 [512, 685]
Sailor2-8B-Validation-News
In-Domain
38,258
487.8 ± 159.1
469 [395, 594]
627.4 ± 121.4
640 [528, 768]
Gemma-3-12B-News
Out-of-Domain
14,444
499.8 ± 163.4
483 [396, 611]
592.0 ± 103.3
592 [488, 690]
Literary works 7
Sailor2-8B-VuTrongPhung
Out-of-Domain
738
478.0 ± 35.6
491 [462, 502]
531.8 ± 14.3
532 [522, 541]
Gemma-3-12B-VuTrongPhung
Out-of-Domain
738
478.0 ± 35.6
491 [462, 502]
575.3 ± 15.8
576 [566, 584]
Table 1: Summary and token-length distributions of Vietnamese datasets from two domains used in this study. The datasets are categorized into In-Domain and Out-of-Domain types. The number of documents (#Docs) is evenly distributed between the Human-written and LLM-generated classes. Token counts are determined using the BPE tokenizer of the PhoGPT model, reporting mean, sample standard deviation (SD), median, and interquartile range [Q1, Q3] per class.
Figure 2: Global optimal threshold points represented on the ROC curve based on VietBinoculars scores for the Sailor2-8B-OptiThreshold-News dataset, showing (1) Youden’s J point, (2) Closest point, and (3) optimal point at 0.06% FPR. In subfigure (a), Youden’s J and Closest thresholds are marked by a red circle [ ∙ ] and a blue square [ ■ ], respectively. The magenta diamond [ ⧫ ] denotes the TPR@0.06%FPR threshold in subfigure (b).
Dataset
Class/
Precision
Recall
F1-score
Support
Metric
Youden’s J
Closest Point
TPR@ 0.06%
Youden’s J
Closest Point
TPR@ 0.06%
Youden’s J
Closest Point
TPR@ 0.06%
OptiThreshold News
Human
0.94
0.95
0.65
0.98
0.97
1.00
0.96
0.96
0.78
41345
LLM
0.98
0.97
1.00
0.94
0.95
0.45
0.96
0.96
0.62
41345
accuracy
0.96
0.96
0.73
82690
macro avg
0.96
0.96
0.82
0.96
0.96
0.73
0.96
0.96
0.70
82690
weighted avg
0.96
0.96
0.82
0.96
0.96
0.73
0.96
0.96
0.70
82690
Table 2: Comparison of detection results using global thresholds of VietBinoculars on the Sailor2-8B-OptiThreshold-News and Sailor2-8B-Validation-News datasets. Results are reported for Youden’s J, Closest Point, and TPR@0.06%FPR thresholds. Zero-shot transfer yields consistent performance across both datasets, with both achieving the same highest F1-score of 0.96. Bold text denotes the highest F1-score within each column for the in-domain datasets.
Figure 3: Confusion matrices of VietBinoculars on the Sailor2-8B-Validation-News dataset for (a) Youden’s J threshold, (b) Closest Point threshold, and (c) TPR@0.06%FPR threshold.
Dataset
Threshold
TP
FP
TN
FN
Accuracy (95% Wilson CI)
Macro F1-score
McNemar p -value
Sailor2-8B- VuTrongPhung ( N=738 )
Youden’s J
369
0
369
0
1.0000 [0.9948, 1.0000]
1.0000
<0.001
Closest Point
369
0
369
0
1.0000 [0.9948, 1.0000]
1.0000
<0.001
TPR@0.06%FPR
248
0
369
121
0.8360 [0.8076, 0.8610]
0.8315
<0.001
Gemma-3-12B- VuTrongPhung ( N=738 )
Youden’s J
369
1
368
0
0.9986 [0.9924, 0.9998]
0.9986
<0.001
Closest Point
369
1
368
0
0.9986 [0.9924, 0.9998]
0.9986
<0.001
TPR@0.06%FPR
330
0
369
39
0.9472 [0.9286, 0.9611]
0.9470
<0.001
Table 3: Generalization performance of VietBinoculars across out-of-domain benchmarks, including raw confusion counts, 95% Wilson score confidence intervals for accuracy, and paired McNemar significance tests against original Binoculars. Global decision thresholds determined on the training set are Youden’s J ( τ=0.8607 ), Closest Point ( τ=0.8729 ), and TPR@0.06%FPR ( τ=0.7023 ). Macro F1-score is reported across Human and AI classes in alignment with Table 2 . Paired McNemar tests employ Edwards’ continuity correction against baseline Binoculars at its standard threshold ( τ=0.9015 ).
Figure 4: Detection performance versus BPE token length across out-of-domain benchmarks: (a) Gemma-3-12B-News, (b) Gemma-3-12B-VuTrongPhung, and (c) Sailor2-8B-VuTrongPhung. Solid and dashed curves denote Accuracy and macro F1-score, respectively. Markers show original Binoculars ( ▶ , × ) and VietBinoculars under Youden’s J threshold ( ⧫ , ■ ). Curves are computed across ten equal-frequency decile bins on length-aligned subsets: N=1,046 ( 105 docs/bin) for news in (a), N=77 (7–8 docs/bin) in (b), and N=90 (9 docs/bin) in (c). Axis spans reflect empirical token support (80–800 tokens in news; 330–550 in literature).
Figure 5: Detection AUROC for various zero-shot and supervised learning methods on Vietnamese out-of-domain datasets: (a)–Gemma-3-12B-News; (b)–Gemma-3-12B-VuTrongPhung; (c)–Sailor2-8B-VuTrongPhung. Zero-shot detectors are represented with the “ \ ” hatch pattern, while supervised learning detectors are represented with the “o” hatch pattern. DetectGPT uses top- k and top- p sampling with parameters k=40 and p=0.96 .
Figure 6: Comparison of VietBinoculars with popular detection methods on the Capybara problem in Vietnamese Language. The x-axis denotes different detection methods, while the y-axis indicates their detection accuracy. VietBinoculars was evaluated using three different global thresholds as described above. Other detection methods, such as RadarTester, which is based on four backbone models (Dolly V2-3B, Camel-5B, Dolly V1-6B, and Vicuna-7B), along with DetectGPT, GPTZero, and Ghostbuster, all apply a default prediction threshold of 0.5 for classification. Commercial tools exhibit poor detection performance on the Vietnamese Capybara problem.
Figure 7: ROC AUC across varying temperature and top- p values. Rows correspond to Sailor2-8B-Chat, Sailor2-14B-Chat, and Gemma-3-12B-it. Columns display VietBinoculars, Binoculars, and their AUC differences, where positive values indicate higher AUC for VietBinoculars. Absolute AUC panels share a common color scale, whereas difference panels use a symmetric scale centered at zero. Outlined (1,1) cells denote greedy controls for Sailor2-8B-Chat and Gemma-3-12B-it, while the corresponding Sailor2-14B-Chat cell represents a sampled configuration.
Generation model
Detector
AUC (T,p)=(1,1)
Mean AUC
Minimum AUC
Mean F1
Minimum F1
Sailor2-8B
VietBinoculars
1.000
0.837
0.145
0.587
0.000
Binoculars
0.761
0.576
0.319
0.001
0.000
Sailor2-14B
VietBinoculars
0.952
0.837
0.135
0.626
0.005
Binoculars
0.851
0.865
0.521
0.220
0.000
Gemma-3-12B
VietBinoculars
1.000
0.997
0.951
0.962
0.570
Binoculars
0.789
0.811
0.660
0.001
0.000
Table 4: Detection performance of VietBinoculars and Binoculars across 16 decoding configurations for each generation model. The reference (T,p)=(1.0,1.0) setting employs greedy decoding for Sailor2-8B and Gemma-3-12B, whereas Sailor2-14B uses sampling. All reference configurations are included in the reported grid means. Binary F1 is computed using the fixed decision thresholds described in the text, where higher values indicate superior performance.
Figure 8: Overview of the AI text rewriting pipeline. In the first stage, each batch of B source records generates a single paraphrase, an intermediate text for double paraphrasing, and a human-style rewrite under a shared temperature and top- p pair. The second stage subsequently rewrites these intermediate texts using new prompts and independently sampled parameters. The final output preserves the original texts alongside the three rewritten variants and their generation metadata.
Model
Text variant
AUC VietBinoculars
AUC Binoculars
Sailor2-8B
Original AI text
1.000
0.761
Single paraphrase
0.929
0.683
Double paraphrase
0.876
0.686
Human-style rewrite
0.919
0.679
Sailor2-14B
Original AI text
0.952
0.851
Single paraphrase
0.907
0.925
Table 5: ROC AUC of VietBinoculars and Binoculars on original and rewritten LLM-generated texts. Each variant is evaluated against human-written passages on a common subset across both detectors. Values are rounded to three decimal places.
Figure 9: Change in ROC AUC relative to the original AI texts, expressed in percentage points. Negative values indicate reduced discrimination, while positive values indicate increased discrimination. The three rewriting strategies are separate conditions. Bars show descriptive estimates without confidence intervals.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Token-length probability distributions per class across the five Vietnamese benchmarks, showing (a) Sailor2-8B-OptiThreshold-News, (b) Sailor2-8B-Validation-News, (c) Gemma-3-12B-News, (d) Sailor2-8B-VuTrongPhung, and (e) Gemma-3-12B-VuTrongPhung. Blue histograms and dashed vertical lines indicate human-written distributions and medians, respectively. Red histograms and dotted vertical lines denote LLM-generated distributions and medians, respectively.
Figure 11: Distribution of VietBinoculars scores using BPE-encoded sequence length on the Gemma-3-12B-News dataset. The plotted lines include the classification thresholds: Youden’s J, Closest Point, and TPR@0.06%FPR. The y-axis indicates the VietBinoculars scores. Each point represents a text sample, with blue squares denoting human-written texts and red circles denoting LLM-generated texts.
Figure 12: Detection accuracy of RadarTester, DetectGPT, GPTZero, Ghostbuster, and VietBinoculars on Vietnamese texts generated by various LLMs from the capybara prompt. The x-axis represents the number of tokens in text samples generated by LLMs using the capybara prompt, while the y-axis indicates the names of the LLMs used to generate the texts. The lines with markers show the detection accuracy of each method. If a detection method correctly identifies a text as LLM-generated, the corresponding marker is placed above zero; otherwise, the marker is positioned at zero.
Figure 13: Inference latency per sample of VietBinoculars across different token lengths on the Vietnamese literary dataset. Measurements reflect synchronized GPU execution after two warmup batches. The dashed line depicts the linear trend line.
Template ID
Rewriting strategy
Original Vietnamese prompt
English translation
P1
Paraphrasing
Hãy viết lại đoạn văn sau với văn phong tự nhiên hơn, giữ nguyên ý nghĩa: {}
Rewrite the following passage in a more natural style while preserving its meaning: {}
P2
Paraphrasing
Diễn đạt lại nội dung dưới đây bằng cách khác nhưng không thay đổi ý nghĩa: {}
Express the content below differently without changing its meaning: {}
P3
Paraphrasing
Viết lại đoạn văn sau theo cách dễ hiểu và tự nhiên hơn: {}
Rewrite the following passage in a clearer and more natural way: {}
P4
Paraphrasing
Hãy paraphrase đoạn văn sau: {}
Paraphrase the following passage: {}
P5
Paraphrasing
Rewrite đoạn sau với cách diễn đạt khác nhưng vẫn giữ nguyên ý nghĩa: {}
Rewrite the following passage using different wording while preserving its meaning: {}
P6
Paraphrasing
Paraphrase đoạn văn sau sao cho tự nhiên hơn: {}
Paraphrase the following passage to make it sound more natural: {}
Appendix
Table 6: Original Vietnamese rewriting prompts and English translations. The translations were not used to generate the evaluated texts.
Model Pair
Model Size
Gemma-3-12B- News
Gemma-3-12B- VuTrongPhung
Sailor2-8B- VuTrongPhung
Falcon-7B (Original Binoculars)
2 × 7B
0.784
0.789
0.761
Qwen2.5-7B (Multilingual Pair)
2 × 7B
0.969
1.000
0.994
PhoGPT-4B (VietBinoculars)
2 × 4B
0.999
1.000
1.000
Appendix
Table 7: Ablation of detector model pairs across Vietnamese out-of-domain benchmarks (ROC AUC).
Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has raised concerns about potential misuse. This underscores the need for reliable and effective methods to detect LLM-generated text. In this paper, we propose IRM, a novel zero-shot approach that leverages Implicit Reward Models for LLM-generated text detection. Such implicit reward models can be derived from publicly available instruction-tuned and base models. Previous reward-based method relies on preference construction and task-specific fine-tuning. In comparison, IRM requires neither preference collection nor additional training. We evaluate IRM on the DetectRL benchmark and demonstrate that IRM can achieve superior detection performance, outperforms existing zero-shot and supervised methods in LLM-generated text detection.
Runheng Liu, Heyan Huang, Xingchen Xiao +1
School of Computer Science and Technology, Beijing Institute of Technology
Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increasingly necessary. Recent zero-shot detectors mainly exploit probability-based statistical discrepancies, but they do not explicitly account for the training process of LLMs, which leaves a distinct generation mechanism insufficiently modeled and limits detection robustness. To address this issue, we propose EchoPrompt, a training-free detector based on latent prompt restoration. Our key intuition is that machine-generated text is typically produced conditioned on an upstream prompt, and this hidden dependency can be partially reactivated by prepending a unified generic prefix. Specifically, EchoPrompt restores a generic assistant-response context, measures the induced likelihood gain with an instruction-tuned model, calibrates it against the corresponding base model, and aggregates the resulting differences into a score that quantifies latent prompt dependency. Extensive experiments show that EchoPrompt achieves state-of-the-art performance among zero-shot detectors while maintaining strong robustness across challenging evaluation settings.
Hongrui Bao, Yubing Ren, Yanan Cao +3
Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · College of Computer Science and Technology, Zhejiang University +1
The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT). This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6. The system integrates edit-extent regression, zero-shot likelihood-contrast signals, lexical statistics, and conservative text rules. With calibrated decision boundaries and conflict-aware integration, our system improves robustness under strong out-of-distribution shifts, achieving a macro-F1 score of 0.8888 and ranking first in the official evaluation. Our code is available at https://github.com/bbbbhrrrr/evildetect.
Hongrui Bao, Hangyu Rong, Zhuoshang Wang +2
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China