Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even with enhanced reasoning capabilities, cannot integrate such computations as an intrinsic and interpretable capability with existing architectures. To this end, we propose Physically-isolated Experts Routing Network (PiERN), an architecture that directs computation and reasoning at token level, thereby enabling iterative alternation within a single chain of thought. We systematically evaluate PiERN on representative computation-reasoning tasks, including PDEBench and battery management tasks. Results show that PiERN achieves not only higher accuracy than directly finetuning LLMs but also significant improvements in response latency, token usage, GPU energy consumption, and experts routing accuracy compared with mainstream multi-agent approaches, while exhibiting no significant degradation in performance on MMLU and GLUE benchmarks. PiERN offers an efficient, interpretable, and scalable paradigm for interfacing language models with scientific systems.
Figures & tables
Figure 1: PiERN achieves high precision with low GPU energy consumption. The horizontal axis represents GPU energy consumption, and the vertical axis represents normalized RMSE on Diffusion-Sorption.
Figure 2: (a): Training of Expert Model for specific high-precision numeric computation tasks. (b): Training the Text-to-Computation Module for alignment, we adpot LLM as text-to-computation Decoder. (c): Training the Token Router to determine LLM for next token prediction or experts for high-precision computational results. Middle: The overall architecture of PiERN .
Figure 3: Token-Level Routing for reasoning-computation inference paradigm in PiERN . Left: Token Router decides to send tokenized inputs into LLM for generating the next token. Middle: Token Router decides to send tokenized inputs into the expert model for high-precision computation. Right: Token Router decides to send tokenized inputs with computation results to LLM for subsequent reasoning and planning.
Component
Metric
Seen templates
Unseen templates
text2comp
Reconstruction RMSE ( ×10−3 ) ↓
6.607
6.569
Token Router
Routing accuracy ↑
100%
100%
Table 1: Generalization of core components to unseen instruction templates.
Figure 4: A five-expert example of the PiERN architecture . Left: Routing to the LLM expert for next token prediction and reasoning. Right: Routing to Burgers expert for high-precision computation.
Figure 5: Token usage and success rate across different tasks comparing PiERN with MAF-based multi-agent systems. The red dashed line denotes PiERN .
Figure 6: Execution-flow decomposition for PiERN and the MAF-based system .
Figure 7: PiERN-BMS Modeling Paradigm . Left: Routing to the LLM for reasoning. Middle: Routing to a scientific expert for capacity estimation. Right: Routing to another expert for profit computation using the computational result from the previous expert .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: The language templates and data synthesis pipeline. Top: Raw data is spatially downsampled and paired, then injected with controlled perturbations (e.g., Scaling ×k ) via natural language templates. Middle: An 80/20 split to evaluate text-to-computation module generalization. Bottom: Token router dynamically chose the LLM or experts.
Model (LLM)
Latency (s)
GPU Energy Consumption (kJ)
Diffusion Reaction
Diffusion Sorption
Burgers
Navier– Stokes
Acoustic Wave
Diffusion Reaction
Diffusion Sorption
Burgers
Navier– Stokes
Acoustic Wave
LangGraph
Qwen3.5-0.8B
54.96
41.73
20.42
65.68
108.94
5.511
4.187
2.076
7.240
13.660
Qwen3.5-2B
86.23
43.30
12.81
65.81
101.59
12.658
6.358
1.888
8.787
13.254
Qwen3.5-4B
15.23
72.24
12.35
85.08
135.01
2.354
11.208
1.911
12.889
20.738
Qwen3.5-9B
12.54
70.72
14.14
78.13
123.84
2.581
14.648
2.944
14.552
23.387
Appendix
Table 4: Latency and GPU energy consumption of LangGraph and Programmatic across five tasks.
Figure 9: Token usage and success rate across different tasks comparing PiERN with multi-agent systems. The red dashed line denotes PiERN .
Model
MMLU
MNLI
RTE
SST2
QQP
QNLI
MRPC
PiERN-2.4B
50.14
33.50
53.09
51.43
36.87
49.65
67.83
Qwen3.5-0.8B
50.24
33.25
52.35
51.03
36.81
49.55
68.38
PiERN-3.6B
57.67
46.57
64.87
80.80
78.59
62.08
74.11
Qwen3.5-2B
57.62
46.58
64.26
80.16
79.90
62.02
72.79
PiERN-5.6B
70.23
47.04
73.30
89.00
64.41
76.45
47.54
Qwen3.5-4B
70.12
46.77
73.29
88.88
64.49
76.37
47.30
Appendix
Table 5: Performance Comparison of PiERN and Qwen Models on MMLU and GLUE Benchmark
Figure 10: Pipeline of language-template-based data synthesis for PiERN-BMS . Top : Extraction of raw physical and numerical information, including battery capacity data and profit data. Middle : Application of Language templates to wrap structured data into natural language contexts. Bottom : The Final data stream, where reasoning tokens and computational tokens are interleaved to form a unified training sequence.