Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even with enhanced reasoning capabilities, cannot integrate such computations as an intrinsic and interpretable capability with existing architectures. To this end, we propose Physically-isolated Experts Routing Network (PiERN), an architecture that directs computation and reasoning at token level, thereby enabling iterative alternation within a single chain of thought. We systematically evaluate PiERN on representative computation-reasoning tasks, including PDEBench and battery management tasks. Results show that PiERN achieves not only higher accuracy than directly finetuning LLMs but also significant improvements in response latency, token usage, GPU energy consumption, and experts routing accuracy compared with mainstream multi-agent approaches, while exhibiting no significant degradation in performance on MMLU and GLUE benchmarks. PiERN offers an efficient, interpretable, and scalable paradigm for interfacing language models with scientific systems.
Figures & tables
Figure 1: PiERN achieves high precision with low GPU energy consumption. The horizontal axis represents GPU energy consumption, and the vertical axis represents normalized RMSE on Diffusion-Sorption.
Figure 2: (a): Training of Expert Model for specific high-precision numeric computation tasks. (b): Training the Text-to-Computation Module for alignment, we adpot LLM as text-to-computation Decoder. (c): Training the Token Router to determine LLM for next token prediction or experts for high-precision computational results. Middle: The overall architecture of PiERN .
Figure 3: Token-Level Routing for reasoning-computation inference paradigm in PiERN . Left: Token Router decides to send tokenized inputs into LLM for generating the next token. Middle: Token Router decides to send tokenized inputs into the expert model for high-precision computation. Right: Token Router decides to send tokenized inputs with computation results to LLM for subsequent reasoning and planning.
Component
Metric
Seen templates
Unseen templates
text2comp
Reconstruction RMSE ( ×10−3 ) ↓
6.607
6.569
Token Router
Routing accuracy ↑
100%
100%
Table 1: Generalization of core components to unseen instruction templates.
Figure 4: A five-expert example of the PiERN architecture . Left: Routing to the LLM expert for next token prediction and reasoning. Right: Routing to Burgers expert for high-precision computation.
Figure 5: Token usage and success rate across different tasks comparing PiERN with MAF-based multi-agent systems. The red dashed line denotes PiERN .
Figure 6: Execution-flow decomposition for PiERN and the MAF-based system .
Figure 7: PiERN-BMS Modeling Paradigm . Left: Routing to the LLM for reasoning. Middle: Routing to a scientific expert for capacity estimation. Right: Routing to another expert for profit computation using the computational result from the previous expert .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: The language templates and data synthesis pipeline. Top: Raw data is spatially downsampled and paired, then injected with controlled perturbations (e.g., Scaling ×k ) via natural language templates. Middle: An 80/20 split to evaluate text-to-computation module generalization. Bottom: Token router dynamically chose the LLM or experts.
Model (LLM)
Latency (s)
GPU Energy Consumption (kJ)
Diffusion Reaction
Diffusion Sorption
Burgers
Navier– Stokes
Acoustic Wave
Diffusion Reaction
Diffusion Sorption
Burgers
Navier– Stokes
Acoustic Wave
LangGraph
Qwen3.5-0.8B
54.96
41.73
20.42
65.68
108.94
5.511
4.187
2.076
7.240
13.660
Qwen3.5-2B
86.23
43.30
12.81
65.81
101.59
12.658
6.358
1.888
8.787
13.254
Qwen3.5-4B
15.23
72.24
12.35
85.08
135.01
2.354
11.208
1.911
12.889
20.738
Qwen3.5-9B
12.54
70.72
14.14
78.13
123.84
2.581
14.648
2.944
14.552
23.387
Appendix
Table 4: Latency and GPU energy consumption of LangGraph and Programmatic across five tasks.
Figure 9: Token usage and success rate across different tasks comparing PiERN with multi-agent systems. The red dashed line denotes PiERN .
Model
MMLU
MNLI
RTE
SST2
QQP
QNLI
MRPC
PiERN-2.4B
50.14
33.50
53.09
51.43
36.87
49.65
67.83
Qwen3.5-0.8B
50.24
33.25
52.35
51.03
36.81
49.55
68.38
PiERN-3.6B
57.67
46.57
64.87
80.80
78.59
62.08
74.11
Qwen3.5-2B
57.62
46.58
64.26
80.16
79.90
62.02
72.79
PiERN-5.6B
70.23
47.04
73.30
89.00
64.41
76.45
47.54
Qwen3.5-4B
70.12
46.77
73.29
88.88
64.49
76.37
47.30
Appendix
Table 5: Performance Comparison of PiERN and Qwen Models on MMLU and GLUE Benchmark
Figure 10: Pipeline of language-template-based data synthesis for PiERN-BMS . Top : Extraction of raw physical and numerical information, including battery capacity data and profit data. Middle : Application of Language templates to wrap structured data into natural language contexts. Bottom : The Final data stream, where reasoning tokens and computational tokens are interleaved to form a unified training sequence.
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives -- via final candidate selection, semantic synthesis, or neuro-symbolic verification -- to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines.
Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently sequential decoding incurs substantial latency, motivating parallelization of the generation process. However, existing parallel reasoning approaches suffer from performance degradation compared to their sequential counterparts, and often rely on specialized inference engines. We introduce ThreadWeaver, a framework for adaptive parallel reasoning that matches the accuracy of comparably sized sequential reasoning models while significantly reducing inference latency via three key innovations: 1) a two-stage parallel trajectory generator that produces high-quality parallel chain-of-thought data for supervised fine-tuning; 2) a trie-based rollout design that enables parallel reasoning on any off-the-shelf autoregressive inference engine; and 3) a parallelization-aware reinforcement learning framework that trains the model to balance reasoning accuracy with effective parallelization. Across six challenging math reasoning benchmarks, ThreadWeaver trained on top of Qwen3-8B achieves performance on par with cutting-edge sequential reasoning models (79.9% on AIME24 and 71.9% on average) while delivering up to 1.53x speedup in token latency, establishing a new Pareto frontier between accuracy and efficiency.
Long Lian, Sida Wang, Felix Juefei-Xu +7
Meta Superintelligence Labs (MSL), Menlo Park, CA, United States · UC Berkeley, Berkeley, CA, United States · UCSF, San Francisco, CA, United States
Inference-time computation has greatly enhanced the performance of large language models (LLMs) on challenging reasoning tasks, but this strategy can incur high inference costs. One solution is to route intermediate chain-of-thought (CoT) states to language models of different sizes; however, existing approaches rely on handcrafted routing strategies that limit performance, or on training large process reward models that may be infeasible in many applications. We formulate stepwise model routing as a constrained decision-making problem, which we solve by training a small control policy using reinforcement learning in conjunction with threshold calibration to tune the performance-efficiency tradeoff. We validate our method on three math benchmarks (GSM8K, MATH500, and OmniMath) on both open and closed models. Our method consistently improves the accuracy-cost tradeoff compared to handcrafted approaches, while achieving a comparable tradeoff to methods that require training large process reward models.
Wenwen Si, Insup Lee, Osbert Bastani
Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA 19104, USA