The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading
Organizations: King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia · Department of Computer Science, Edge Hill University, Ormskirk, England
Abstract
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.
Figures & tables
| Qwen3.5-397B-A17B | Kimi K2.5 | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Policy | Acc. | Tokens | Tok. save | Time | Time save | Acc. | Tokens | Tok. save | Time | Time save |
| LongBench-v2 | full reading | 0.535 | 382,971 | 0% | 86.0 h | 0% | 0.517 | 340,761 | 0% | 36.3 h | 0% |
| random stop | 0.474 | 185,632 | 52% | 44.2 h | 48.7% | 0.481 | 168,707 | 50% | 18.7 h | 48.7% | |
| verbalized gate | 0.533 | 267,179 | 30.2% | 54.9 h | 36.2% | 0.517 | 290,354 | 14.8% | 27.6 h | 24.1% | |
| END gate | 0.507 | 134,569 | 65% | 28.3 h | 67.1% | 0.525 | 183,604 | 46% | 18.1 h | 50.3% | |
| ACS , fixed | 0.541 | 190,966 | 50% | 40.4 h | 53.1% | 0.519 | 331,070 | 3% | 31.2 h | 14.1% | |
| S-NIAH ( per model) | ||||||
|---|---|---|---|---|---|---|
| Model | Policy | Premature | Over-read | Acc. | Regret | Capture |
| Qwen3-14B | fixed at 25% | 56.4% | 0.58 | 0.436 | +0.560 | 50.2% |
| random stop | 32.3% | 2.25 | 0.676 | +0.320 | 37.2% | |
| verbalized gate | 8.4% | 0.00 | 0.916 | +0.080 | 92.0% | |
| END gate | 5.2% | 0.00 | 0.948 | +0.048 | 95.2% | |
| ACS , fixed | 1.6% | 1.51 | 0.980 | +0.016 | 60.7% | |
| ACS | Confidence only | Verbalized gate | ||||
|---|---|---|---|---|---|---|
| Model | Acc. | Premature | Acc. | Premature | Acc. | Premature |
| Qwen2.5-7B | 0.948 | 3.2% | 0.892 | 9.2% | 0.700 | 29.2% |
| Qwen3-14B | 0.980 | 1.6% | 0.928 | 6.8% | 0.916 | 8.4% |
| Qwen3-32B | 1.000 | 0.0% | 1.000 | 0.0% | 0.592 | 40.8% |
| Gemma-3-12B | 0.980 | 1.6% | 0.872 | 12.8% | 0.544 | 45.6% |
| Gemma-3-27B | 0.876 | 12.0% | 0.520 | 23.2% | 0.876 | 12.0% |
| Model | Policy | Pre-proxy | Over-read | Acc. |
|---|---|---|---|---|
| Qwen3.5-397B | fixed at 25% | 37.8% | 0.66 | 0.420 |
| random stop | 24.9% | 2.71 | 0.522 | |
| verbalized gate | 3.3% | 2.03 | 0.708 | |
| END gate | 4.4% | 1.74 | 0.701 | |
| ACS , fixed | 2.4% | 3.20 | 0.697 | |
| full reading | 0.0% | 5.01 | 0.705 |
| Axis | Configuration | Acc. | Premature | Tokens |
|---|---|---|---|---|
| Stopping signal | verbalized gate | 0.916 | 8.4% | 27,019 |
| END gate | 0.948 | 5.2% | 28,403 | |
| ACS | 0.980 | 1.6% | 37,696 | |
| Stability test, five models | confidence only | 0.890 | 10.4% | 37,901 |
| both tests (default) | 0.957 | 3.7% | 44,768 | |
| Chunk size | 12K | 0.968 | 3.2% | 34,390 |
| Axis | Configuration | Acc. | Premature | Tokens |
|---|---|---|---|---|
| Stopping signal | verbalized gate | 0.916 | 8.4% | 27,019 |
| END gate | 0.948 | 5.2% | 28,403 | |
| ACS | 0.980 | 1.6% | 37,696 | |
| Stability test, five models | confidence only | 0.890 | 10.4% | 37,901 |
| both tests (default) | 0.957 | 3.7% | 44,768 | |
| Chunk size | 12K | 0.968 | 3.2% | 34,390 |
| Qwen3.5, LB-v2 | Kimi, LB-v2 | Needle premature | ||||
|---|---|---|---|---|---|---|
| Acc. | Saving | Acc. | Saving | Qwen3.5 | Kimi | |
| 0.80 | 0.479 | 79% | 0.521 | 58% | 30.4% | 0.4% |
| 0.90 | 0.513 | 73% | 0.531 | 51% | 28.0% | 0.4% |
| 0.92 | 0.517 | 72% | 0.535 | 49% | 24.8% | 0.0% |
| 0.95 | 0.527 | 68% | 0.533 | 44% | 10.4% | 0.0% |
| 0.98 | 0.533 | 61% | 0.523 | 29% | 1.2% | 0.0% |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Mean chunks | Saving, Qwen3.5 | Saving, Kimi | |
|---|---|---|---|---|
| Long-dialogue history | 39 | 13.9 | 31% | 31% |
| Single-document QA | 175 | 19.1 | 39% | 28% |
| Multi-document QA | 125 | 22.4 | 31% | 32% |
| Long in-context learning | 81 | 36.3 | 50% | 45% |
| Long structured data | 33 | 51.8 | 37% | 42% |
| Code repository | 50 | 150.2 | 62% | 67% |
| 0.5 | 0.8 | 0.9 | 0.95 | 0.97 | 0.98 | 0.99 | 0.995 | |
|---|---|---|---|---|---|---|---|---|
| Acc. | 0.764 | 0.788 | 0.856 | 0.872 | 0.900 | 0.912 | 0.956 | 0.980 |
| Tokens (K) | 26.1 | 27.1 | 31.1 | 31.8 | 33.1 | 33.6 | 36.5 | 37.7 |
| Model | Tuned | Tuned Acc. | Fixed Acc. | Fixed Premature |
|---|---|---|---|---|
| Qwen2.5-7B | 0.878 | 0.922 | 6.1% | |
| Qwen3-14B | 0.965 | 0.965 | 2.6% | |
| Qwen3-32B | 0.965 | 1.000 | 0.0% | |
| Gemma-3-12B | 0.983 | 0.983 | 1.7% | |
| Gemma-3-27B | 0.791 | 0.835 | 16.5% | |
| Mean | — | 0.917 | 0.941 | 5.4% |
| Model | Configuration | Acc. | With probes | Probe cost | Without probes | Saving w/o probes |
|---|---|---|---|---|---|---|
| Qwen3.5 | fixed/best | 0.541 | 190,966 | 24,079 | 166,887 | 56.4% |
| Kimi | fixed | 0.519 | 331,070 | 37,988 | 293,082 | 14.0% |
| Kimi | best | 0.535 | 172,648 | 19,968 | 152,680 | 55.2% |
| Qwen3.5 | Kimi | Qwen3-14B | Premature | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Acc. | Saving | Acc. | Saving | Acc. | Saving | Qwen3.5 | Kimi | Qwen3-14B | |
| 0.80 | 0.660 | 36% | 0.730 | 32% | 0.532 | 53% | 10.3% | 7.9% | 20.7% |
| 0.90 | 0.682 | 31% | 0.742 | 26% | 0.556 | 48% | 8.4% | 6.6% | 18.1% |
| 0.92 | 0.686 | 30% | 0.750 | 24% | 0.556 | 48% | 7.7% | 5.1% | 17.7% |
| 0.95 | 0.698 | 26% | 0.752 | 19% | 0.562 | 46% | 6.2% | 3.5% | 16.3% |
| 0.98 | 0.708 | 21% | 0.754 | 11% | 0.580 | 40% | 4.0% | 1.8% | 12.2% |
| Axis | Configuration | Acc. | Regret | Tokens |
|---|---|---|---|---|
| Chunk size | 12K | 0.425 | 3.8% | 209,470 |
| 24K (default) | 0.538 | 1.3% | 168,160 | |
| 48K | 0.500 | 1.3% | 151,960 | |
| Notes cap | 3K | 0.525 | 2.5% | 98,755 |
| 6K (default) | 0.538 | 1.3% | 168,160 | |
| 12K | 0.475 | 2.5% | 188,747 |
| Qwen3.5-397B-A17B | Kimi K2.5 | |||||
|---|---|---|---|---|---|---|
| Acc. | Tokens | Saving | Acc. | Tokens | Saving | |
| 2 | 0.533 | 185,722 | 51.5% | 0.529 | 168,644 | 50.5% |
| 3 (default) | 0.541 | 190,966 | 50.0% | 0.535 | 172,648 | 49.0% |
| 4 | 0.537 | 202,054 | 47.2% | 0.533 | 180,912 | 46.9% |
| 5 | 0.541 | 206,716 | 46.0% | 0.527 | 194,035 | 43.1% |
| full reading | 0.535 | 382,971 | 0% | 0.517 | 340,761 | 0% |
| Model | Metric | |||||||
|---|---|---|---|---|---|---|---|---|
| Qwen2.5-7B | Acc. | 0.568 | 0.596 | 0.596 | 0.596 | 0.700 | 0.700 | 0.700 |
| Premature | 43.2% | 40.4% | 40.4% | 40.4% | 29.2% | 29.2% | 29.2% | |
| Qwen3-14B | Acc. | 0.916 | 0.916 | 0.916 | 0.916 | 0.916 | 0.916 | 0.916 |
| Premature | 8.4% | 8.4% | 8.4% | 8.4% | 8.4% | 8.4% | 8.4% | |
| Qwen3-32B | Acc. | 0.588 | 0.592 | 0.592 | 0.592 | 0.592 | 0.592 | 0.592 |
| Premature | 41.2% | 40.8% | 40.8% | 40.8% | 40.8% | 40.8% | 40.8% |