ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models
Organizations: Department of Computer Science and Engineering, Korea University · Department of Computer Science and Engineering, Konkuk University · Human-inspired AI Research
Abstract
Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model's trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at https://github.com/js-lee-AI/ThinkFuse.
Figures & tables
| Method | Intervention unit | Selective trigger | Multi-model fusion | Trajectory-aware | Decision basis |
|---|---|---|---|---|---|
| Self-consistency ( Wang et al., 2023 ) | answer | post-hoc majority voting | |||
| Cool-Fusion ( Liu et al., 2025 ) | segment | fuses at every segment | |||
| AdaFuse ( Cui et al., 2026 ) | word / span | local token confidence | |||
| MUR ( Yan et al., 2026 ) | step | single-model momentum uncertainty | |||
| SpecReason ( Pan et al., 2025 ) | step | verifies every drafted step | |||
| ThinkFuse (ours) | segment | trajectory-calibrated segment uncertainty |
| Method | MATH-500 | GSM8K | AIME24 | GPQA | NQ-Open |
|---|---|---|---|---|---|
| Standalone models | |||||
| Qwen3-4B | 89.0 | 92.0 | 53.3 | 54.0 | 29.5 |
| Qwen3-1.7B ‡ | 83.0 | 90.0 | 33.3 | 39.4 | 28.0 |
| Ministral | 32.5 | 67.0 | 0.0 † | 25.8 | 24.0 |
| EXAONE-4.0 | 57.0 | 85.0 | 3.3 † | 38.9 | 22.0 |
| Test-time compute baseline: Qwen3-4B | |||||
| Method | Wall-clock (s) | TFLOPs | Acc. | Acc. / TFLOPs |
|---|---|---|---|---|
| Cool-Fusion | 39.4 | 86 | 21.8 | 0.25 |
| AdaFuse | 33.9 | 74 | 43.7 | 0.60 |
| ThinkFuse | 68.7 | 77 | 72.9 | 0.95 |
| Auxiliary | MATH-500 | GPQA |
|---|---|---|
| Llama-3.2-3B-IT | 89.5 (+0.5) | 50.5 ( 3.5) |
| Gemma-2-2B-it | 87.5 ( 1.5) | 51.0 ( 3.0) |
| Mistral-7B-Inst-v0.3 | 84.5 ( 4.5) | 44.9 ( 9.1) |
| Qwen2.5-3B-Inst | 88.0 ( 1.0) | 52.5 ( 1.5) |
| Qwen3-1.7B (no-think) | 89.5 (+0.5) | 50.0 ( 4.0) |
| Setting | MATH-500 | GSM8K |
|---|---|---|
| Fusion budget penalty | ||
| Soft budget (ours, ) | 90.0 (+1.0) | 96.9 (+4.9) |
| No budget penalty ( ) | 88.5 ( 0.5) | 90.5 ( 1.5) |
| Segment length | ||
| 87.0 ( 2.0) | 93.5 (+1.5) | |
| (ours) | 90.0 (+1.0) | 96.9 (+4.9) |
| Shared prefix. For AIME-2022 II #11, both traces begin with the same coordinate setup: , , , , and is the midpoint of . Both traces then try to use the two angle-bisector conditions at . | |
|---|---|
| Standalone | ThinkFuse |
| Generated trace near the divergence. I place the quadrilateral in coordinates and write the distance equations. Since is the midpoint of , it should lie on the two angle bisectors. the valid solution is b = -7/sqrt(2), a = 7/sqrt(2) So, D is at ((7sqrt(2))/2, -7sqrt(2)/2). I plug this coordinate for back into the remaining constraints and solve for . (2+sqrt(2))yˆ2 + 9(sqrt(2)+1)y + (22+7sqrt(2)) = 0 // Comment: this fixes D as if . The second angle-bisector condition has not been used, so the trace enters a dead-end quadratic. | Generated trace near the fusion point. I keep and symbolic, then express the two angle-bisector conditions as distance-to-line equations. [Two angle-bisector distance equations from M = (B+C)/2] -8 - 2p + 3r = 0 => r = (8 + 2p)/3 => s = 2q/3 Substituting these relations keeps both angle-bisector constraints active while the remaining equations are solved. // Comment: the auxiliary segment keeps symbolic. It uses both angle-bisector conditions before solving. |
| Answer: | Answer: |
| Scope | Setting | Implementation details |
| Shared | Engine / hardware | All generation and PPL computation use vLLM v0.8.5 with bfloat16 precision. Each model pair is served on two NVIDIA A100-80GB GPUs. |
| Generation budget / sampling | Standard decoding uses max_model_len =16384, max_new_tokens =16384, temperature=0.6, top- =0.95, top- =-1, and min- =0.0. The standalone Qwen3-1.7B alone uses top- =20 following the Qwen3 recommendation. Within this budget, explicit thinking-stage generation uses an additional 2K-token cap. Self-consistency uses separate sampling settings. | |
| ThinkFuse | Trigger signal | Every time the primary model generates tokens, we compute top-5 token entropy and aggregate token-level uncertainty into segment-level uncertainty using EWCA with . |
| Adaptive threshold | The EMA update rate is , the uncertainty tolerance is , and the initial statistics are . The soft-budget coefficient is ; the trigger uses both recent uncertainty history and the cumulative fusion ratio. | |
| Segment selection | When fusion is triggered, primary and auxiliary candidate segments are compared under the current context using primary-model PPL. Segment boundaries from different tokenizers are aligned by the shared prefix that decodes identically under both tokenizers, and each segment is restored to the reasoning-tag format of the model used for PPL scoring. | |
| Baselines | Cool-Fusion | Each participating LLM generates a candidate segment; all generated candidates are scored by the PPL of all participating LLMs, and the candidate with the lowest average PPL is selected. |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Auxiliary | MATH-500 | GSM8K | AIME24 | GPQA | NQ-Open |
|---|---|---|---|---|---|
| Ministral | +0.5 [ 1.5, +2.5] | +3.1 [+1.0, +5.6]* | +20.0 [+6.7, +36.7]* | +10.6 [+1.0, +20.2]* | +10.0 [+5.5, +15.0]* |
| EXAONE-Deep | +2.0 [+0.5, +4.0]* | +3.1 [+1.0, +5.8]* | +20.0 [+6.7, +33.3]* | +1.0 [ 8.1, +10.1] | +9.0 [+4.5, +14.0]* |
| DeepSeek-R1 | +1.5 [ 0.5, +4.0] | +3.0 [+0.0, +6.5] | +3.3 [ 10.0, +16.7] | +8.7 [ 0.5, +17.9] | +8.0 [+3.5, +13.1]* |
| Qwen3-1.7B | +2.0 [+0.5, +4.0]* | +2.0 [ 0.5, +5.0] | +16.7 [ 3.3, +36.7] | +1.0 [ 8.6, +10.1] | +7.5 [+3.5, +12.0]* |
| Method | MATH-500 | GSM8K | AIME24 | GPQA | NQ-Open |
|---|---|---|---|---|---|
| Standalone models | |||||
| Qwen3-4B | 89.0 | 92.0 | 53.3 | 54.0 | 29.5 |
| EXAONE-Deep | 83.5 | 90.5 | 43.3 | 57.1 | 17.5 |
| DeepSeek-R1 | 13.0 | 30.5 | 0.0 † | 21.2 | 14.5 |
| ThinkFuse cross-family pairs | |||||
| Qwen3-4B EXAONE-Deep | 91.0 (+2.0) | 99.0 (+7.0) | 73.3 (+20.0) | 55.1 (+1.1) | 38.5 (+9.0) |
| Benchmark | Qwen3-4B pass@5 |
|---|---|
| MATH-500 | 90.5 (+1.5) |
| GSM8K | 95.5 (+3.5) |
| AIME24 | 70.0 (+16.7) |
| GPQA | 71.2 (+17.2) |
| NQ-Open | 48.5 (+19.0) |
| Benchmark | ThinkFuse | Post-hoc critique |
|---|---|---|
| MATH-500 | 91.0 (+2.0) | 86.3 ( 2.7) |
| GSM8K | 99.0 (+7.0) | 92.0 (+0.0) |
| AIME24 | 73.3 (+20.0) | 30.0 ( 23.3) |
| GPQA | 55.1 (+1.1) | 27.8 ( 26.2) |
| NQ-Open | 38.5 (+9.0) | 40.0 (+10.5) |
| Method | AIME24 | GPQA |
|---|---|---|
| Qwen3-4B standalone | 53.3 | 54.0 |
| MUR ( -decoding) | 16.7 ( 36.6) | 37.9 ( 16.1) |
| MUR (Ministral critic) | 53.3 (+0.0) | 48.0 ( 6.0) |
| ThinkFuse | 73.3 (+20.0) | 64.6 (+10.6) |