We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other (F(7,14)=0.59, p=0.76), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by 3.1 points and change the top-ranked system. In place of preference scoring we propose emoji-affect decodability, a reference-based probe whose rankings are stable to ±0.003 macro-F1 across seeds.
Figures & tables
Source
MS
F
p
System s (vs. residual)
11.85
18.62
<10−16
System s (vs. s×r )
11.85
0.59
0.76
Rater r
1345.16
319.03
<10−16
System × rater
20.25
31.82
<10−16
Residual p×s×r
0.64
—
—
Table 1: Three-way ANOVA on the crossed 290×8×3 design. The system effect is significant against the residual and non-significant against the system × rater interaction, which is the correct error term when annotators are a sample from a population. The interaction mean square exceeds the system mean square.
Panel
Top system
ρ w/ full
Claude rank
All three
Gemma-3-27B
—
3
− A1
Claude-3-Haiku
0.55
1
− A2
Gemma-3-27B
0.90
4
− A3
Mistral-Large
0.62
8
Table 2: Leave-one-annotator-out leaderboards. Removing any single annotator from a three-person panel can change the winner. Claude-3-Haiku ranges from first to last.
System
mos
emojis
ead
win LM
Gemma-3-27B
3.667
4.95
0.653
0.543
Mistral-Large
3.638
4.87
0.639
0.506
Claude-3-Haiku
3.579
4.92
0.584
0.499
GPT-4.1-nano
3.516
4.55
0.562
0.428
DeepSeek-V3.2
3.454
3.45
0.560
0.541
Gemini-2.0-Flash
3.443
2.67
0.657
0.603
Table 3: Systems ordered by raw mos . Mean emoji count tracks mos almost perfectly; decodability does not. win LM is the within-item length-matched win rate, which reorders the table. ∗∗p<.01 .
Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating emotionally grounded responses across five categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. We evaluate three technically diverse LLMs across three dimensions: Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding. The framework aims to assess and compare the quality of LLM responses in both English and Ukrainian languages, as well as to explore the reliability of the LLM-as-a-jury paradigm. Our findings reveal that Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct degrade when transitioning to Ukrainian, while DeepSeek-V3 remains comparatively stable within our benchmark. We additionally find that human and AI evaluators agree weakly on empathy and naturalness but diverge on cultural grounding. We further argue that producing Ukrainian text is not equivalent to producing Ukrainian emotional support. Our findings may assist in the future development of more culturally tailored benchmark designs, as well as encourage a stronger emphasis on human-centered evaluation.
Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language can invert backbone rankings. We localize the Agent-as-a-Judge prompt stack to five typologically diverse languages (English, Arabic, Turkish, Chinese, Hindi) and evaluate 55 DevAI development tasks across three developer-agent frameworks and six judge backbones, totaling 4950 judge runs. The central finding is that backbone and language interact: GPT-4o achieves the highest satisfaction in English (44.72%), while Gemini leads in Arabic (51.72%, p<0.001 vs.\ GPT-4o) and Hindi (53.22%). No single backbone dominates across all languages, and inter-backbone agreement on individual requirement judgments is modest (Fleiss' κ≤0.231). A controlled ablation further shows that localizing judge-side instructions, not just benchmark content, can be decisive: Hindi satisfaction drops from 42.8% to 23.2% under partial localization. These results indicate that language should be treated as an explicit evaluation variable in agentic benchmarks. Full requirement-level judgments and runtime statistics are released for reproducibility.
Emotional validation - explicitly acknowledging that a user's feelings make sense - has proven therapeutic value but has received little computational attention. Emotional validation in dialogue systems can be decomposed into (i) validating response identification, (ii) validation timing detection, and (iii) validating response generation. To support research on all three subtasks, we release M-EDESConv, a 120k English-Japanese multilingual corpus created through hybrid manual and automatic annotation, and M-TESC, a multilingual spoken-dialogue test set. For timing detection, we propose MEGUMI, a Multilingual Emotion-aware Gated Unit for Mutual Integration, that fuses frozen XLM-RoBERTa semantics with language-specific emotion encoders via cross-modal attention and gated fusion. MEGUMI shows superior performance on both the M-EDESConv and M-TESC datasets, both objectively and subjectively. Finally, our EmoValidBench benchmarks of GPT-4.1 Nano and Llama-3.1 8B indicate that current LLMs generate contextually similar and diverse validating responses, but emotional understanding remains a major area for improvement. Project page: https://github.com/zihaurpang/Multilingual-Emotional-Validation