Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix methods reduce this cost by reusing encoded context and scoring multiple decisions in parallel, but do not model decision dependencies or verify execution. We propose SharedKV-BT, where each active node of a behavior tree (BT) exposes stage-local fields and candidates, and Shared-KV scores the candidates in parallel and passes the selected decision to a separate execution system. We tested SharedKV-BT on robot manipulation, mobile navigation, and computer-use tasks. Across three tasks, SharedKV-BT made typed decisions 2.36-4.15 times faster than prompt-matched autoregressive decoding. On the manipulation task, node-local Shared-KV improved joint decision accuracy from 75% to 94% and closed-loop success from 0% to 60%. Fixed-score policy replay showed that stage gating prevented out-of-order actions and external postconditions prevented premature completion.
Figures & tables
Fig. 1: SharedKV-BT schematic. (a) An example pick-and-place BT sequences Acquire , Transport , and Deposit ; the active Acquire branch expands into a bounded retry loop over typed decision and execution followed by an external postcondition check. (b) The instruction and observed state are encoded once, Shared-KV scores the illustrative node-local skill, object, and completion candidates in parallel, and fieldwise argmax forms the decision. Completion is diagnostic; Check Result determines whether the active stage succeeded and controls BT progression.
Fig. 2: Evaluated tasks. The top row shows Stack, full PickPlace, and NutAssemblySquare; the bottom row shows Door, NavigateKitchen, and a composite of Clock, Notepad, and File Explorer from WindowsAgentArena runs that passed the platform evaluator. Images illustrate the task interfaces rather than additional evaluation trials.
Context
Matched batch
End-to-end comparison
Full prefix (ms)
Shared-KV (ms)
Speedup
Shared-KV (ms)
Constrained AR (ms)
Speedup
Stack
207.1±0.6
98.9±0.5
2.09×
99.8
413.8
4.15×
NavigateKitchen
208.6±0.6
119.5±0.4
1.75×
121.4
286.0
2.36×
Clock
170.4±0.9
131.9±10.0
1.29×
134.2
318.8
2.38×
TABLE I: Qwen2.5-3B latency on model inputs saved at decision points. Matched-batch values are mean ± 95% confidence intervals over ten runs. End-to-end values are mean latencies from input preparation through decision output in the prompt-matched constrained AR comparison.
Fig. 3: Qwen2.5-3B matched-batch cache-reuse comparison. (a) Latency is compared using identical candidate batches with and without prefix reuse; error bars show 95% confidence intervals over ten runs. The sweeps vary (b) prefix length, (c) the number of field-candidate pairs, and (d) candidate length. The dashed line indicates equal latency. Shared-KV is slower with only two candidates because cache-copy overhead outweighs the computation saved by prefix reuse, but becomes faster as more prefix computation is reused.
Fig. 4: Adaptive PickPlace isolates node-local interface routing. (a) Correct joint typed decisions on the predefined schedule of saved pre-execution states; counts are descriptive because the same state descriptions repeat across seeds. Closed-loop outcomes report (b) successful four-object episodes and (c) mean completed object–bin relations over 30 episodes per condition. NL-KV denotes node-local Shared-KV, Global denotes global-interface Shared-KV, and AR denotes node-local constrained AR. Global-interface selections were rejected rather than corrected.
Fig. 5: Representative archived failure modes, not additional trials. (a) Door perception gate: incomplete evidence that the handle was visible caused the safety rule to withhold action. (b) Door controller failure: repeated permitted OPEN_DOOR calls failed to open the door physically. (c) NutAssemblySquare insertion-recovery failure: recovery exhausted its attempt budget before insertion succeeded.