Uncovering Uncontrolled Repetition through Residual Stream Dynamics
Organizations: Beijing University of Posts and Telecommunications · JIUTIAN Research · Nanyang Technological University · Chongqing University of Posts and Telecommunications
Abstract
Uncontrolled repetition can prolong autoregressive generation in large language models (LLMs) and enable resource consumption attacks. Prior analyses of repetitive generation have identified strongly activated features in intermediate and late layers. However, how uncontrolled repetition activity emerges and develops before becoming prominent in these layers remains insufficiently understood. In this paper, we investigate this question primarily in large vision-language models (LVLMs), which support a richer set of uncontrolled repetitions through both visual and textual inputs. We propose Tokenwise Residual Comparison (TRC), a method that identifies and localizes anomalies associated with repetition from residual dynamics during generation. TRC compares attention and multilayer perceptron writes to the residual stream across generated tokens to identify patterns associated with repetition. It then selectively suppresses coordinates in the residual stream at the identified layer. Experiments show that TRC effectively mitigates uncontrolled repetition, reducing loop rates by 57% on average. Our analysis further shows that repetition semantics emerge in shallow layers and propagate through the residual stream, disrupting normal representations. TRC also generalizes to large language models (LLMs) and large reasoning models (LRMs), where it consistently captures analogous repetition dynamics and achieves effective mitigation. Our work broadens the study of repetitive generation from its prominent internal representations to earlier opportunities for intervention, providing insights for mitigating resource consumption attacks.
Figures & tables
| RECITE (visual) | GCG (textual) | LoopLLM (textual) | |||||
|---|---|---|---|---|---|---|---|
| Model | Defense | Length | Loop rate | Length | Loop rate | Length | Loop rate |
| InstructBLIP-7B | No defense | 1,866.44 | 92.0% | 1,531.56 | 96.0% | 1,833.96 | 100.0% |
| Fixed-length | 389.84 | 92.0% | 1,017.76 | 96.0% | 952.36 | 88.0% | |
| No-repeat | 18.36 | 4.0% | 17.88 | 4.0% | 22.04 | 0.0% | |
| AUSteer | 413.16 | 20.0% | 85.40 | 4.0% | 165.84 | 8.0% | |
| TRC | 422.04 | 16.0% | 1,312.56 | 56.0% | 1,229.72 | 60.0% | |
| Method | ScienceQA | TextVQA | Average |
|---|---|---|---|
| Benign | 82.00% | 79.00% | 80.50% |
| Fixed-length | 82.00% | 79.00% | 80.50% |
| No-repeat | 82.00% | 78.50% | 80.25% |
| AUSteer | 76.00% | 60.50% | 68.25% |
| TRC | 81.50% | 78.50% | 80.00% |
| LLM | LRM | ||||
|---|---|---|---|---|---|
| Setting | Qwen2.5-3B | Llama-3.2-3B | DeepSeek-Llama-8B | Qwen3.6-27B | GLM-4.7-Flash |
| Benign | 61.0% | 55.0% | 97.0% | 99.5% | 100.0% |
| TRC | 61.0% | 56.0% | 95.0% | 99.5% | 100.0% |
| Difference (pp) | 0.0% | +1.0% | -2.0% | 0.0% | 0.0% |
| Block | Metric | Normal | Repeat | Gap/ |
|---|---|---|---|---|
| 0 | L2 | 22.2009 | 22.8838 | 0.234 |
| 0 | RMS | 0.4946 | 0.5057 | 0.172 |
| 1 | L2 | 12.9280 | 13.7456 | 0.730 |
| 1 | RMS | 0.2867 | 0.3038 | 0.692 |
| Input | Coordinates | Residual intervention | Repeat rate (%) | Transition success (%) | Avg. length | Len (%) | Observed change |
|---|---|---|---|---|---|---|---|
| Attack | – | None | 100.0 | – | 4096 | – | Repetition baseline |
| Attack | TRC (ours) | Normal residual | 20.0 | 80.0 | 828 | Strong restoration | |
| Attack | Random | Normal residual | 60.0 | 40.0 | 2462 | Partial restoration | |
| Normal | – | None | 0.0 | – | 11 | – | Normal baseline |
| Normal | TRC (ours) | Attack addition | 0.0 | 0.0 | 2 | – | Output failure without repetition |
| Normal | Random | Attack addition | 0.0 | 0.0 | 11 | – | Nearly unchanged |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | RECITE | GCG | LoopLLM | Direct | Benign tasks |
| LVLMs | |||||
| InstructBLIP-Vicuna-7B | Y | Y | Y | – | ScienceQA, TextVQA |
| Qwen2.5-VL-3B-Instruct | Y | Y | Y | – | ScienceQA, TextVQA |
| LLaVA-1.5-7B | Y | Y | – | – | ScienceQA, TextVQA |
| LLMs | |||||
| Llama-3.2-3B | – | Y | Y | Y | MMLU |
| Approach | Primary observation | Localization target | Demonstrated scope | Principal distinction from TRC |
|---|---|---|---|---|
| FFN vocabulary analysis ( Geva et al., 2021 ; Geva et al., 2022 ) | Activations and vocabulary-space promotion produced by feed-forward updates | FFN memories, values, and layer-wise concept promotion | Text LMs | Explains stored or promoted concepts, but does not compare residual changes along a repetitive generation trajectory. |
| Layer-wise lexical probing ( Liu et al., 2024b ) | Probe recoverability of lexical-semantic information from hidden representations | Layers at which lexical semantics are most accessible | Generative text LMs | Measures information accessibility through an external readout rather than locating native residual coordinates for intervention. |
| Repetition neurons ( Hiraoka and Inui, 2025 ) | Neuron-activation changes before and after the onset of repetition | Repetition-associated neurons in intermediate and final layers | Text LMs | Directly studies repetition, but uses activation amplitude around onset rather than cycle-aligned changes before residual addition. |
| SAE repetition features ( Yao et al., 2025 ) | Logit-based layer screening followed by learned sparse-feature activations | SAE features within selected layers | Text LLMs | Resolves features with an auxiliary learned dictionary; TRC operates on native attention and MLP contribution coordinates. |
| Cross-modal information flow ( Zhang et al., 2025c ) | Answer-performance changes after blocking attention between image and question positions | Layered cross-token attention pathways for visual–linguistic integration | LLaVA-style MLLMs | Explains modality fusion within MLLMs, but is not a repetition-localization method and is not evaluated on text-only LLMs or LRMs. |
| TRC (ours) | Cycle-aligned changes in native attention and MLP contributions, normalized by benign magnitude and local cross-layer variation | A shallow pre-addition residual-write layer and its highest-scoring hidden coordinates | LVLMs, LLMs, and LRMs | Uses one modality-agnostic scoring and masking formulation without adding an auxiliary interpreter or network module. |
| Image | Prompt | Generated output |
|---|---|---|
| Examine the image and return exactly two lines. Line 1 must begin with ’Description:’ and contain one concise factual English sentence describing the main visible object. Do not use the word ’echo’ in line 1. Line 2 must be exactly: ’Repeat: echo echo echo echo echo echo echo echo echo echo.’ Do not add any other text before, between, or after these two lines. | Description: The image depicts a simple, stylized red apple with a green leafy top. Repeat: echo echo echo echo echo echo echo echo echo echo | |
| Examine the image and return exactly two lines. Line 1 must begin with ’Description:’ and contain one concise factual English sentence describing the visible objects. Do not use the word ’signal’ in line 1. Line 2 must be exactly: ’Repeat: signal signal signal signal signal signal signal signal signal signal.’ Do not add any other text before, between, or after these two lines. | Description: Three mugs of different colors are lined up on a table. Repeat: signal signal signal signal signal signal signal signal signal signal. |
| Attack | |||
|---|---|---|---|
| GCG | |||
| RECITE | |||
| LoopLLM |
| (%) | GCG | RECITE | LoopLLM | Macro avg. | Coordinate overlap |
|---|---|---|---|---|---|
| No defense | – | ||||
| 1 | |||||
| 5 | |||||
| 10 | |||||
| 15 | |||||
| 20 |
| (%) | Accuracy | Avg. length |
|---|---|---|
| 0 | 82.0 | 2.125 |
| 1 | 82.0 | 2.130 |
| 5 | 81.5 | 2.145 |
| 10 | 82.0 | 2.115 |
| 15 | 81.0 | 2.125 |
| 20 | 80.5 | 2.065 |
| Attack | |||||
|---|---|---|---|---|---|
| GCG | 12 | 1 | 1 | 1 | 1 |
| RECITE | 1 | 1 | 1 | 1 | 1 |
| Model | Attack | Original (s) | TRC (s) | Reduction |
|---|---|---|---|---|
| Qwen2.5-3B | Direct | 278.92 | 9.33 | 96.7% |
| Qwen2.5-3B | GCG | 402.52 | 1.73 | 99.6% |
| Qwen2.5-3B | LoopLLM | 91.72 | 8.36 | 90.9% |
| Llama-3.2-3B | Direct | 340.19 | 5.38 | 98.4% |
| Llama-3.2-3B | GCG | 64.73 | 1.67 | 97.4% |
| Llama-3.2-3B | LoopLLM | 143.44 | 4.68 | 96.7% |
| Model | Original | TRC | Difference |
|---|---|---|---|
| Qwen2.5-3B | 22.60 | 19.60 | |
| Llama-3.2-3B | 25.76 | 26.80 |