cs.AINov 2, 2025

A Quantitative Study of Sustained Focus in Large Language Models via Repetitive Deterministic Prediction Tasks

Authors: Wanda Hou, Leon Zhou, Hong-Ye Hu, Yubei Chen, Yi-Zhuang You, Xiao-Liang Qi

Organizations: EdenCode Inc. · Stevenson School · Harvard University · UC Davis · Path Integral Technology, Inc.

Abstract

We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rate (SAR) scales with output length. Each such task involves the repetition of the same operation NN times. Examples of such tasks include letter replacement in letter strings following a given rule, integer addition, and multiplication of string operators in many-body quantum mechanics. If the LLM performs the task by a simple repetition algorithm, the success rate would follow an exponential decay with sequence length. In contrast, our experiments on leading LLMs reveal a crossover that is sharper than exponential: −log⁡SAR-\log\mathrm{SAR} grows super-linearly with NN, and accuracy collapses around a characteristic length N∗N_*, the accuracy cliff that separates reliable from unreliable generation. The hypothesis of independent per-step errors is rejected for every model and task we studied. The crossover is well described by a double-exponential accumulation law, SAR=exp⁡(−β0NαN−1)\mathrm{SAR}=\exp(-β_0 Nα^{N-1}), whose crossover scale N∗N_* does not depend on the functional form chosen to fit it. To interpret this behaviour we introduce a minimal effective model in which step-correctness variables interact through dense random couplings and compete with an external field set by the prompt. Solved by direct enumeration, the model reproduces the super-linear error accumulation and the accuracy cliff qualitatively, and it assigns to each model--task pair two interpretable parameters, an intrinsic error rate and an error-accumulation factor.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 30, 2026cs.LG

Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We investigate this accuracy--stability relationship in large language model (LLM) evaluation for deterministic text-conditioned generation, using programming tasks as a concrete testbed. Standard code-generation benchmarks emphasize single-run accuracy or eventual success under repeated sampling, but many deployment settings also require stability: consistent outcomes across repeated invocations under the same task description. We present a repeated-run evaluation protocol with metrics for run-level accuracy, retry-free coverage, and per-problem variability. On a recency-based benchmark of 100 LeetCode-style problems, we evaluate 16 models from five provider families under two prompt templates with five repeated runs per problem, yielding 16,000 evaluation instances. Although run-level pass rate and perfect stability rate are strongly correlated (r=0.985), pass rate consistently exceeds retry-free coverage -- a gap that reaches 17.8 percentage points and reverses model rankings even among closely matched systems. Prompt effects are model-dependent rather than uniformly beneficial. These results suggest that repeated-run stability analysis is a necessary complement to conventional accuracy reporting for deterministic text-conditioned generation tasks.
May 1, 2026cs.CL

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models

Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithfully execute the procedure specified in a prompt. We introduce a controlled diagnostic benchmark for procedural execution, where models are given a step-wise arithmetic procedure and two numeric inputs, and must return the final computed value. Complexity is varied through procedure length and look-back dependencies over intermediate variables. Average first-answer accuracy drops from 63% on 5-step procedures to 20% on 95-step procedures. Generation-level analysis shows that failures often involve missing answers, premature answers, self-correction after an initial error and under-executed traces. These findings suggest that apparent reasoning ability can mask substantial weaknesses in faithful long-horizon procedural execution.
May 8, 2026cs.CL

Limits of Reliability and Scaling in Language Models

Large language models (LLMs) are trained and evaluated as though perfect reliability is achievable for any task given sufficient scale. We show that this assumption is information-theoretically unjustified. Every generative task has a reliability ceiling that no model can exceed, determined by how much output uncertainty is resolvable from observable context. The gap decomposes into a resolvable component closable with additional context and a subjective component inherent to task ambiguity. Autoregressive generation further degrades this ceiling at a rate governed by the task's dependency kernel, which quantifies inter-token correlations in the output. From these two primitives, we derive a first-principles scaling law where LLM performance is bottlenecked by the scarcer resource: training data or model capacity. This law recovers the Chinchilla scaling law as a special case and provides a structural account of when scaling improves reliability. Beyond scaling, our framework unifies diverse practical phenomena, such as the benefits of retrieval-augmentation and the spectral mechanics of catastrophic forgetting. Our work formalizes the resource-complexity tradeoffs that govern model performance across domains, offering a unified theory of performance limits in generative language models.