OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous
Authors: Yuji Takubo, Daniele Gammelli, Marco Pavone, Simone D'Amico
Organizations: Stanford University, Stanford, CA 94305, USA · Department of Aeronautics and Astronautics, Stanford University, Stanford, CA 94305, USA · Italian Institute of Artificial Intelligence (AI4I), Turin, Italy · NVIDIA Research, Santa Clara, CA 95051, USA
Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-level operational intent into safe, dynamically feasible trajectories, creating a bottleneck to scalable operations. Large language model (LLM)-based agents could offer an intuitive interface for this process, although their outputs are not inherently grounded in orbital dynamics, operational constraints, or the structure of admissible spacecraft maneuvers. To exploit their semantic reasoning while ensuring the generated plan's physical validity, this paper presents a hierarchical framework for spacecraft task-and-motion planning (TAMP) that grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules. Within this framework, a pretrained LLM maps a natural-language command to a partial mission specification. The associated planners then resolve unspecified decisions within the admissible operational space. Finally, trajectory optimization converts the completed mission specification into a dynamically feasible trajectory. Numerical experiments demonstrate that this architecture substantially improves intent recovery over direct LLM generation, achieving 98% exact recovery of partial mission specifications across all evaluated splits when backed by frontier LLMs. Additional test-time-compute experiments show that, for a compact 9B model, verifier-guided revision increases exact recovery from 75% to 88%, while broader behavior-plan search independently improves selection among admissible trajectory realizations. Overall, these results establish a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO.
Figures & tables
Fig. 1 : Overview of the proposed framework. The intent parsing initializes the partial mission specification M0 ; behavior sequencing produces M1 ; waypoint generation produces M2 ; and trajectory optimization generates the final state and control trajectories. All three specifications are successive states of the same MissionConfig .
Fig. 2 : Example behavior–waypoint graph.
Domain
Description
aδλ [m]
aδey [m]
aδiy [m]
E0
E/I-separated, Central
[−5,5]
±[30,70]
±[30,70]
E+
E/I-separated, +V -bar
[100,250]
±[30,70]
±[30,70]
E−
E/I-separated, −V -bar
[−250,−100]
±[30,70]
±[30,70]
F+
Flat, +V -bar
[100,250]
[−5,5]
[−5,5]
F−
Flat, −V -bar
[−250,−100]
[−5,5]
[−5,5]
TABLE I : Waypoint domains used in the case study.
ID
Behavior primitive
Directed transitions
Duration (orbits)
1
Station-keeping
Self-loop at all domains
[1,3]
2
Drift, slow short, +V direction
E0→E+ , E−→E0
[2,4]
3
Drift, slow short, −V direction
E0→E− , E+→E0
[2,4]
4
Expand R/N separation
F+→E+ , F−→E−
[1,3]
5
Shrink R/N separation
E+→F+ , E−→F−
[1,3]
6
Approach from −V -bar
F−→E0
[1.5,4]
TABLE II : Behavior primitives used as directed edges in the behavior-waypoint graph.
Random
Heuristic
Learned pϕ
Fixed
Variable
Fixed
Variable
Fixed
Variable
SCP success [%]
84.0
95.0
95.7
95.1
77.1
95.6
Fuel [ cm/s ]
15.31±8.83
11.05±5.37
9.37±5.42
6.73±4.19
7.00±4.08
6.65±4.13
SCP iterations
6.86±8.01
9.98±7.43
5.85±7.25
8.05±6.13
6.72±7.27
7.21±7.12
TABLE III : Waypoint-generation performance on 1,000 held-out test cases; fuel and SCP iterations are the mean ± standard deviation over the cases in which the SCP converges.
Fig. 3 : Preference win rate of the structured behavior–duration planning against the best of K uniformly sampled graph-valid candidates over 100 scenarios.
Behavior selector
Waypoint proposal
Waypoint contract in SCP
SCP success
Fuel [cm/s]
SCP iterations
Aux. SCP calls/request
Percentile [%]
Online time [s]
Strict
Baseline
Uniform random (best of 64)
Representative waypoint
Fixed
100/100
7.65±7.04
2.90±3.80
0
–
–
–
Leg library
100/100
5.12±8.06
3.37±7.37
60
98.9
99.6
74.77
Uniform random (best of 64)
Min.-displacement heuristic
Variable
100/100
3.82±3.94
4.75±4.30
0
–
–
–
Learned Qψ
97/100
2.07±2.79
3.57±3.93
0
96.2
96.5
3.02
TABLE IV : Behavior-duration planning and corresponding trajectory performance after trajectory optimization. Fuel and SCP iterations are the mean ± standard deviation over successful trajectories.
Entry
Representative realization
Command
Drift towards the target, traverse the near-target region for 1 orbit, and then withdraw toward +V -bar. Prioritize minimizing fuel over maximizing observation quality, and complete the entire operation within 5 orbits.
Scenario
w0=[0,−156.24,0,−38.97,0,−60.99]
rKOZ=30m (default), M0=0.0∘
M0
P= [fuel, observation], H=∅ (default)
T∈[⊥,5.0]orbits
b=[2,1,7] , d=[⊥,1.0,⊥]orbits , w=[⊥,E0,F+]
TABLE V : Representative progressive completion of a natural-language command into M0 , M1 , and M2 , followed by the SCP-optimized waypoint sequence.
Fig. 4 : Representative trajectory optimized via SCP based on the completed MissionConfig (cf. Table V ).
Fig. 5 : Architecture comparison against the proposed hierarchical architecture, tested with three dataset splits.
Dataset split
Backend
Exact M0 [%]
Atomic match [%]
Precision
Recall
Development
Qwen3.5-9B
75
90.6
85.1
Ministral3-8B-Instruct
70
88.1
89.0
GPT-5.6 Terra (high)
98
99.0
99.5
Held-out instance
Qwen3.5-9B
78
91.1
89.8
Ministral3-8B-Instruct
70
89.1
90.3
TABLE VI : Intent-parsing performance on the prompt-development split and two held-out splits.
Fig. 6 : End-to-end mission satisfaction and pipeline completion across semantic parsers on 100 commands per split. All parsed specifications are screened by the same deterministic boundary/path verifier at Nrev=0 .
Fig. 7 : Effect of semantic test-time compute on intent parsing with Qwen3.5-9B and the development-split commands.
Fig. 8 : Effect of planning-search budget over the top- Nplan behavior-duration candidates ranked by Qψ .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Style
Command realization
Formal
Circumnavigate the target through the near-target operating region. Prioritize maximizing observation quality over minimizing completion time.
Natural
Move the chaser around the target along a circumnavigation path. Give observation quality higher priority than completion time.
Trajectory optimization is a critical component for enabling safe and reliable autonomous operations in space exploration. As space missions increase in frequency, complexity, and scope, there is a growing need to rapidly formulate mathematically sound trajectory optimization problems that accurately reflect mission objectives and operational constraints. However, translating mission intent into tractable analytical formulations for trajectory optimization requires substantial domain expertise. This paper presents a framework that leverages large language models (LLMs) to translate natural language descriptions of mission requirements and constraints into executable trajectory optimization code and corresponding mathematical formulations. Experiments in spacecraft rendezvous scenarios demonstrate a high success rate in reconditioning a convex trajectory optimization problem from semantic mission requirements. Ultimately, this work highlights the potential of LLMs to bridge high-level intent and formal optimization models, enabling more flexible and efficient trajectory design of spacecraft.
Eleanor Brosius, Yuji Takubo, Daniele Gammelli +2
Stanford University 496 Lomita Mall, Stanford, CA, 94305
Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We introduce AstroAgentBench, a seven-family benchmark for executable space mission planning in the domains of scheduling, observation planning, constellation design, and relay support. For each case, an agent submits a planning artifact that is checked by an external verifier for schema, timing, geometry, resources, and mission value. Results report validity and normalized scores, with comparisons to task-specific solver references. Across five LLM agent systems and 35 held-out cases, the strongest systems approach or exceed solver-reference scores on several families, while weaker systems often fail to produce high-value valid plans and even strong systems lose quality on geometric, product-level, or design-heavy tasks. Trace analyses separate two failure points: task-contract misformulation and weak solution construction. Successful runs instead calibrate agent-written implementations against verifier feedback and adapt search to case-specific structure. Ablations show that procedure injection and memory accumulation help selectively, when they supply the missing formulation, calibration, or search support.
Weiyi Wang, Xinchi Chen, Jingjing Gong +2
Fudan University · OpenMOSS Team · Shanghai Innovation Institute
Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing GRASP, a strategy-aware, multi-stage planning framework. GRASP decouples the planning pipeline across specialized, context-isolated modules: it pre-compiles global macro-guidelines (GenPlan), explores alternative localized strategies within isolated context windows (RevPlan), and independently evaluates trajectories using a multi-criteria discriminator (VerPlan). Empirical evaluations show that GRASP consistently establishes a new state-of-the-art frontier across diverse datasets, yielding substantial accuracy gains over direct LLM planners on Natural Plan Calendar Scheduling (∼12.4\%$$\uparrow), ZebraLogic (∼30.8\%$$\uparrow), and SciBench Math. Crucially, under multi-task scaling-where standard planners suffer immediate performance collapse-GRASP completely flattens the multi-task degradation penalty. In interleaved dual-task environments, GRASP achieves an absolute accuracy gain of up to 16.7% over direct LLM planners. Furthermore, by isolating context and enforcing strict macro-regularization, GRASP outperforms frontier reasoning models (such as GPT-5-mini) by a margin of 14.5%.
Arunabh Srivastava, Mohammad A., Khojastepour +2
Amir · University of Maryland, College Park, MD · NEC Laboratories America, Inc.