cs.LGSep 17, 2025

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning

Authors: Jingyuan Fan, Purui Liu, Hengbo Xiao, Yuxuan Zheng, Jingzhao Zhang, Chao Lu, Guannan He

Organizations: Peking University · Tsinghua University

Abstract

Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even with enhanced reasoning capabilities, cannot integrate such computations as an intrinsic and interpretable capability with existing architectures. To this end, we propose Physically-isolated Experts Routing Network (PiERN), an architecture that directs computation and reasoning at token level, thereby enabling iterative alternation within a single chain of thought. We systematically evaluate PiERN on representative computation-reasoning tasks, including PDEBench and battery management tasks. Results show that PiERN achieves not only higher accuracy than directly finetuning LLMs but also significant improvements in response latency, token usage, GPU energy consumption, and experts routing accuracy compared with mainstream multi-agent approaches, while exhibiting no significant degradation in performance on MMLU and GLUE benchmarks. PiERN offers an efficient, interpretable, and scalable paradigm for interfacing language models with scientific systems.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models

    Nov 24, 2025Long Lian, Sida Wang, Felix Juefei-Xu +7LLM Reasoning StrategiesInference Latency

  2. Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning

    May 7, 2026Wenwen Si, Insup Lee, Osbert BastaniInference CostChain-of-Thought Reasoning