CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment
Abstract
Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Many dual-mode models leave this choice to users. Automating it is challenging because routing targets evolve with the policy, initial mode preferences destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and modeconditioned responses in one shared policy directly from a native dual-mode checkpoint, without method-specific SFT warm-up. Paired current-policy counterfactual rollouts assign cross-mode credit only to the routing token, while within-mode GRPO trains response tokens. A paired-to-self-routed curriculum stabilizes early training with forced rollouts from both modes, then increases self-routed updates to improve autonomous routing. Across nine benchmarks, CounterRoute better balances accuracy and efficiency than heuristic and learned adaptive-routing methods. Relative to always-thinking checkpoints, it improves macro-average accuracy while reducing mean generated tokens by 51% for Qwen3-8B and 41% for Qwen3-14B. On instruction-following and commonsense benchmarks where direct answering is strong, think rates fall as low as 1% while response quality improves. Despite training only on math and instruction following, its routing behavior and response quality generalize to held-out coding, science, knowledge, and commonsense benchmarks.
Figures & tables
| Accuracy (%) | Efficiency | |||||||||||
| Method (Qwen3-8B) | GSM8K | MATH500 | GPQA | HumanEval | MMLU-Pro | MMLU | HellaSwag | IFEval | IFBench | Avg | Mean Tok. | Saving |
| Dual-mode post-trained checkpoints (no additional RL) | ||||||||||||
| Always Think | 95.8 | 92.0 | 57.6 | 90.2 | 65.8 | 82.7 | 84.0 | 90.2 | 34.7 | 77.0 | 2770 | 0% |
| Always No-Think | 89.2 | 81.8 | 46.7 | 86.0 | 52.4 | 77.3 | 80.4 | 89.7 | 28.2 | 70.2 | 488 | 82% |
| Fixed-mode GRPO (same data, no routing) | ||||||||||||
| GRPO Think-Only | 95.6 | 92.8 | 57.9 | 92.7 | 73.8 | 82.4 | 83.9 | 90.9 | 34.9 | 78.3 | 1571 | 43% |
| Accuracy (%) | Efficiency | |||||||||||
| Method (Qwen3-14B) | GSM8K | MATH500 | GPQA | HumanEval | MMLU-Pro | MMLU | HellaSwag | IFEval | IFBench | Avg | Mean Tok. | Saving |
| Dual-mode post-trained checkpoints (no additional RL) | ||||||||||||
| Always Think | 96.1 | 93.1 | 62.7 | 93.3 | 74.9 | 84.5 | 87.8 | 91.1 | 37.5 | 80.1 | 2211 | 0% |
| Always No-Think | 90.1 | 85.6 | 50.4 | 90.9 | 60.6 | 79.9 | 86.1 | 90.4 | 31.4 | 73.9 | 396 | 82% |
| Fixed-mode GRPO (same data, no routing) | ||||||||||||
| GRPO Think-Only | 96.3 | 94.4 | 61.2 | 96.3 | 76.6 | 84.9 | 87.5 | 91.4 | 37.8 | 80.7 | 1416 | 36% |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Shared Qwen3 setting |
|---|---|
| Initialization | Corresponding dual-mode Qwen3 post-trained checkpoint |
| Prompt template | Fixed system prompt; one direct and one reasoning demonstration |
| Training data | Math and synthetic instruction following, each weighted 0.5 |
| Batch and duration | Batch 32; 300 optimizer steps; 4 epochs; rollouts |
| Lengths | 4096 prompt tokens; 8192 response tokens |
| Sampling | temperature 1.0; top- 0.8; top- |
| Forced Think | Forced No-Think | CounterRoute Self-Routed | ||||||
|---|---|---|---|---|---|---|---|---|
| Benchmark | Acc (%) | Med Words | Acc (%) | Med Words | (pp) | Acc (%) | Med Words | Think (%) |
| GPQA | 57.07 | 1444 | 35.91 | 2 | 56.9 | 1383 | 90.0 | |
| MMLU-Pro | 73.24 | 501 | 53.86 | 2 | 73.3 | 467 | 93.0 | |
| MMLU | 81.87 | 299 | 71.95 | 2 | 81.2 | 268 | 81.0 | |
| MATH500 | 92.20 | 443 | 56.30 | 43 | 92.5 | 440 | 100.0 | |
| GSM8K | 96.13 | 210 | 91.58 | 61 | 95.8 | 212 | 100.0 | |
| Variant | Accuracy (%) | Mean Words | Think (%) |
|---|---|---|---|
| Control method : scheduled -decay (later ckpt) | 78.58 | 877.1 | 68.77 |
| Control method : fixed | 77.41 | 1015.4 | 69.02 |
| No counterfactual pairing | 77.42 | 1214.5 | 98.96 |
| Always paired | 77.25 | 1409.0 | 100.00 |
| Trajectory-wide routing broadcast | 51.78 | 2645.6 | 63.32 |
| Shared GRPO group | 72.16 | 595.8 | 42.19 |
| Fixed margin | Accuracy (%) | Med Words | Think (%) |
|---|---|---|---|
| 78.28 | 803 | 99.79 | |
| 77.41 | 603.8 | 69.02 | |
| 72.40 | 132 | 17.54 | |
| 71.36 | 120 | 19.79 |
| Model | Method | Avg Acc (%) | Think Rate | Mean Tok. | Saving |
|---|---|---|---|---|---|
| Qwen3-8B | Always Think | 77.0 | 100% | 2770 | 0% |
| Thinkless-style (DeGRPO) | 79.4 | 100% | 2499 | 10% | |
| CounterRoute | 79.1 | 65.6% | 1347 | 51% | |
| Qwen3-14B | Always Think | 80.1 | 100% | 2211 | 0% |
| Thinkless-style (DeGRPO) | 81.7 | 97% | 1848 | 16% | |
| CounterRoute | 82.4 | 68.3% | 1296 | 41% |