cs.CLSep 29, 2026

Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech

Authors: Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian

Organizations: BitmanagerAI, Dubai, UAE · lab260, Yerevan, Armenia · MTUCI, Moscow, Russia

Abstract

Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k >= 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.

Figures & tables

Appendix figures & tables35 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Taming Long-form Text-to-Speech

    Sep 15, 2026Rongxiang Wang, Berkin Durmus, Aysegul Orhon +2Autoregressive Text-To-SpeechVoice Conversion