General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional Single Prompt (SP) approaches suffer from context bloat and the "Lost in the Middle" phenomenon, leading to unreliable task planning. We propose an LLM chaining architecture that separates instruction classification and action generation into two specialized stages, reducing per-inference prompt length by approximately 45% while improving planning consistency. We evaluate our method using 100 randomly generated GPSR commands across three language models spanning local open-source and frontier cloud deployment contexts. Results show consistent planning improvements over SP across all models, with gains of up to +37 percentage points on local models. Further, real-robot execution experiments on the Toyota Human Support Robot (HSR) reveal that planning success alone does not guarantee task completion, with 6 of 10 tasks completing successfully and execution-layer failures identified as the primary remaining bottleneck.
Figures & tables
Fig. 1 : Overview of the proposed LLM chaining framework deployed on the HSR platform. Left: the two-stage pipeline (Instruction Classifier → Action Generator). Right: real-robot execution with semantic map overlay and generated navigation action.
Fig. 4 : GPSR arena layout used in the real-robot experiments: four rooms connected by open doorways, 14 furniture locations, and a fixed instruction point. Layout follows the RoboCup@Home Japan GPSR definition [ 10 ] .
Model
Type
SP
Chaining
Δ (pp)
Qwen2.5-14B
Local
7
33
+26
Cogito-14B
Local
17
54
+37
Claude Sonnet 4.6
Cloud
81
88
+7
TABLE I : Planning Success Rate (%): SP vs. Chaining
Figure 4
Command: “Take the cola on the table and pass it to me.”
Chaining (success)
Stage 1
FetchObjectAndDeliverToMe
Stage 2
Navigation(dining_table)
DetectObject(cola)
Grasp(cola)
Navigation(instruction_point)
TABLE II : Planning Example: Chaining vs. SP on the Same Command
Fig. 7 : Three of the four execution-layer failures observed on the HSR. Planning succeeded in all three cases; the failures arose in the perception and localization components the plans depended on. See Table III .
Task
Plan
Exec
Layer
Cause
Fetch green tea
✓
×
Recognition
SAM3 unstable
Guide by shirt color
✓
×
Recognition
Timeout
Follow waving person
✓
△
Navigation
Localization drift
TABLE III : Execution-Layer Failures on the HSR (three of four cases)