cs.CLJul 26, 2026

Do LLM Debates Repeat Arguments Differently Across Languages?

Authors: Huiqian Lai

Organizations: Syracuse University

Abstract

LLM debate is usually evaluated by final answers, yet transcripts reveal whether later turns develop new arguments or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}, which compares extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers \textit{prior-argument similarity} across languages, yet does not significantly narrow the Chinese--English gap. Multilingual debate evaluation should therefore measure argumentative development over time and report both average and gap terms.

Explore similar work

CardsList
  1. Argument Collapse: LLMs Flatten Long-Form Public Debate

    Jun 1, 2026Yekyung Kim, Yapei Chang, Chau Minh Pham +1DebateHuman Oversight