cs.AIJan 16, 2026

AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks

Authors: Weiyi Wang, Xinchi Chen, Jingjing Gong, Xuanjing Huang, Xipeng Qiu

Organizations: Fudan University · OpenMOSS Team · Shanghai Innovation Institute

Abstract

Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We introduce AstroAgentBench, a seven-family benchmark for executable space mission planning in the domains of scheduling, observation planning, constellation design, and relay support. For each case, an agent submits a planning artifact that is checked by an external verifier for schema, timing, geometry, resources, and mission value. Results report validity and normalized scores, with comparisons to task-specific solver references. Across five LLM agent systems and 35 held-out cases, the strongest systems approach or exceed solver-reference scores on several families, while weaker systems often fail to produce high-value valid plans and even strong systems lose quality on geometric, product-level, or design-heavy tasks. Trace analyses separate two failure points: task-contract misformulation and weak solution construction. Successful runs instead calibrate agent-written implementations against verifier feedback and adapt search to case-specific structure. Ablations show that procedure injection and memory accumulation help selectively, when they supply the missing formulation, calibration, or search support.

Figures & tables

Explore similar work

CardsList
  1. Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

    Jun 3, 2026Haoyu Sun, Wenxuan Wang, Mingyang Song +5Large Language Model PlanningAgentic Benchmarks

  2. AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

    May 23, 2026Hao Liu, Siyuan Yang, Qinglei Hu +1Human-Annotated BenchmarkHigh-Fidelity

  3. OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous

    Oct 1, 2026Yuji Takubo, Daniele Gammelli, Marco Pavone +1Large Language Model PlanningSpacecraft