AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks
Organizations: Fudan University · OpenMOSS Team · Shanghai Innovation Institute
Abstract
Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We introduce AstroAgentBench, a seven-family benchmark for executable space mission planning in the domains of scheduling, observation planning, constellation design, and relay support. For each case, an agent submits a planning artifact that is checked by an external verifier for schema, timing, geometry, resources, and mission value. Results report validity and normalized scores, with comparisons to task-specific solver references. Across five LLM agent systems and 35 held-out cases, the strongest systems approach or exceed solver-reference scores on several families, while weaker systems often fail to produce high-value valid plans and even strong systems lose quality on geometric, product-level, or design-heavy tasks. Trace analyses separate two failure points: task-contract misformulation and weak solution construction. Successful runs instead calibrate agent-written implementations against verifier feedback and adapt search to case-specific structure. Ablations show that procedure injection and memory accumulation help selectively, when they supply the missing formulation, calibration, or search support.
Figures & tables
| Task family | Horizon | Primary input entries per case |
|---|---|---|
| AEOSSP | 12 h | 20–28 satellites; 1,614–1,970 tasks |
| SatNet | 7 d | 257–333 requests; 2,513–3,370 view periods |
| SPOT-5 | / | 8–1,057 candidate photographs |
| Stereo | 48 h | 10–12 satellites; 121–144 targets |
| Regional | 72 h | 6–12 satellites; 10,492–18,325 grid cells |
| Revisit | 48 h | 24–32 targets |
| Component | Description |
| Validity | |
| Submission | The agent submits the required planning artifact before timeout. |
| Schema | The artifact parses into the family-specific decision format. |
| Time | Actions lie inside horizons, access or service windows, and required temporal gaps. |
| Geometry | The verifier recomputes visibility, pointing, range, footprint, illumination, or line of sight. |
| Resources | The plan respects concurrency limits, stateful consumables, and capacity or demand budgets. |
| Method | Valid | AEOSSP | Regional | Relay | Revisit | SatNet | SPOT-5 | Stereo |
|---|---|---|---|---|---|---|---|---|
| Claude Code + Claude Opus 4.6 | 26/35 | 58.45 | 21.22 | 12.25 | 76.50 | 60.36 | 43.96 | 19.16 |
| Codex CLI + GPT-5.4 | 35/35 | 72.91 | 72.08 | 64.91 | 69.58 | 69.45 | 61.92 | 68.91 |
| Kimi CLI + Kimi K2.6 | 35/35 | 72.03 | 73.15 | 57.77 | 73.36 | 66.14 | 61.92 | 0.25 |
| OpenCode + MiniMax M2.7 | 23/35 | 36.44 | 17.73 | 0.00 | 0.00 | 37.54 | 37.47 | 0.00 |
| OpenCode + DeepSeek V4 Pro | 34/35 | 71.44 | 35.79 | 35.67 | 70.95 | 60.23 | 60.91 | 13.00 |
| Best solver baseline | / | 75.95 | 86.44 | 61.79 | 70.98 | 61.59 | 60.92 | 96.05 |
| System | Revisit | Stereo | SatNet | Reg. | Relay |
|---|---|---|---|---|---|
| Claude | 13.6(0) | 20.2(2) | 19.0(1) | 0.0(5) | 1.0(4) |
| Codex | 5.8(0) | 8.8(0) | 6.0(0) | 8.2(0) | 6.4(0) |
| Kimi | 9.2(0) | 10.4(0) | 16.0(0) | 11.6(0) | 8.8(0) |
| OC+MM | 3.0(4) | 18.0(3) | 19.4(0) | 25.8(1) | 19.4(0) |
| OC+DS | 6.0(0) | 47.2(0) | 10.0(0) | 18.2(0) | 12.4(0) |
| Aspect | AEOSSP | SatNet | SPOT-5 | Stereo | Regional | Revisit | Relay |
|---|---|---|---|---|---|---|---|
| State and reference geometry | |||||||
| Propagation | SGP4 (TEME) | precomputed | abstracted | SGP4 (TEME) | SGP4 (TEME) | J2 | J2 |
| Frame conversion | GCRF–ITRF | / | / | GCRF–ITRF | GCRF–ITRF | GCRF–ITRF | GCRF–ITRF |
| Surface model | WGS84 | / | / | WGS84 | WGS84 | spherical | spherical |
| Geometry and action feasibility | |||||||
| Pointing variable | off-nadir | / | / | along/across | roll | off-nadir | / |
| Family | Submitted solution | Hard-validity layers | Native metrics |
|---|---|---|---|
| AEOSSP | point-observation actions | task window; required duration; sensor type; visibility; off-nadir; same-satellite overlap; slew; battery | completion; turnaround; energy |
| SatNet | antenna-track rows | view-period containment; setup; teardown; antenna occupancy; maintenance; request/resource match; minimum duration | unsatisfied demand; satisfied requests; tracking hours |
| SPOT-5 | one camera-mode assignment per photograph | domain membership; binary conflicts; ternary conflicts; multi-orbit memory cap | profit; memory use |
| Stereo Imaging | raw observation actions | timing; access interval; solar elevation; off-nadir; overlap; slew; stereo-product geometry | target coverage; product quality |
| Regional Coverage | roll-only strip actions | time grid; duration; sensor band; strip intersection; overlap; roll slew; battery; duty limit; minimum regional coverage | weighted coverage; actions; battery |
| Revisit Constellation | initial satellite states; observation actions | satellite cap; orbit bounds; visibility; range; off-nadir; timing; overlap; slew; battery | revisit gap; satellite count |
| Family | Total | Train | Test |
|---|---|---|---|
| AEOSSP | 30 | 10 | 5 |
| SatNet | 5 | 0 | 5 |
| SPOT-5 | 21 | 10 | 5 |
| Stereo | 15 | 10 | 5 |
| Regional | 15 | 10 | 5 |
| Revisit | 15 | 10 | 5 |
| Family | Reference | Core abstraction |
|---|---|---|
| AEOSSP | MWIS conflict graph [ 9 ] | Enumerates feasible observation candidates, links incompatible candidates in a conflict graph, and selects a weighted independent set before verifier-facing repair. |
| AEOSSP | Greedy LNS [ 2 ] | Builds a satellite-local schedule by greedy candidate insertion, then reinserts bounded neighborhoods to improve completion and turnaround. |
| SatNet | Delta-MILP [ 7 ] | Uses a mixed-integer contact-assignment model to allocate antenna time while controlling request-level unsatisfied demand. |
| SatNet | PPO [ 11 ] | Uses a trained reinforcement-learning policy to choose contact assignments under the SatNet request-service metric. |
| SPOT-5 | Reference lookup [ 22 ] | Matches known held-out instances to archived challenge solutions and recomputes their profit under the benchmark verifier. |
| Stereo Imaging | CP/local search [ 22 ] | Constructs a library of pair and tri-stereo products, then inserts and repairs products as coupled scheduling objects. |
| Layer | Contents exposed to the run |
|---|---|
| Task package | Rendered family brief, rendered task prompt, case files, required output contract, and the runnable opaque verifier helper. |
| Runtime tools | Shared Linux container with Python 3.13, Node.js 24.x, OpenJDK 17, shell tools, Git, JSON utilities, file-search utilities, and the evaluated agent CLI packages. |
| Python libraries | Pinned astrodynamics, geometry, optimization, and data libraries, including Brahe, Basilisk, Orekit bindings, Skyfield, OR-Tools, PuLP, NetworkX, NumPy, SciPy, pandas, Shapely, pyproj, and plotting utilities. |
| Agent-local guidance | The Brahe skill document is installed under .agents/skills/ for each harness where supported. Procedure-injection and memory-accumulation experiments add only the configured procedures or prior-run notes for the corresponding condition. |
| Execution controls | A fixed two-hour wall-clock limit, 8 CPU allocation, 32 GB memory limit, 16 GB shared-memory allocation, no human repair during the run, and post-run official scoring of the final submitted artifact. |
| Collected evidence | Final submitted artifact, official verifier output, aggregate metrics, and the harness-specific session logs needed for later trace analysis. |
| Evaluated system | Harness package | Reasoning configuration |
|---|---|---|
| Claude Code + Claude Opus 4.6 | @anthropic-ai/claude-code 2.1.123 | high |
| Codex CLI + GPT-5.4 | @openai/codex 0.125.0 | high |
| Kimi CLI + Kimi K2.6 | kimi-cli 1.40.0 | thinking |
| OpenCode + MiniMax M2.7 | opencode-ai 1.14.30 | thinking |
| OpenCode + DeepSeek V4 Pro | opencode-ai 1.14.30 | max |
| Family | Claude | Codex | Kimi | OC+MM | OC+DS |
|---|---|---|---|---|---|
| AEOSSP | 58.45 [29.06, 73.95] | 72.91 [69.87, 76.15] | 72.03 [69.43, 74.14] | 36.44 [16.60, 52.71] | 71.44 [69.15, 73.73] |
| Regional | 21.22 [20.71, 22.22] | 72.08 [60.04, 80.37] | 73.15 [67.40, 77.38] | 17.73 [9.98, 24.82] | 35.79 [24.15, 44.92] |
| Relay | 12.25 [0.00, 36.75] | 64.91 [55.81, 77.71] | 57.77 [42.85, 72.63] | 0.00 [0.00, 0.00] | 35.67 [11.55, 59.78] |
| Revisit | 76.50 [73.72, 78.85] | 69.58 [61.90, 76.01] | 73.36 [69.78, 78.13] | 0.00 [0.00, 0.00] | 70.95 [69.98, 71.93] |
| SatNet | 60.36 [29.25, 79.41] | 69.45 [61.78, 77.13] | 66.14 [58.88, 74.96] | 37.54 [16.55, 55.85] | 60.23 [50.07, 70.39] |
| SPOT-5 | 43.96 [11.09, 76.83] | 61.92 [44.58, 82.18] | 61.92 [44.58, 82.18] | 37.47 [7.15, 71.09] | 60.91 [43.99, 82.17] |
| Repetition | Case 8 | Case 28 | Case 1021 | Case 1403 | Case 1506 | Valid | Mean |
|---|---|---|---|---|---|---|---|
| 1 | 100.00 | 0.00 | 54.80 | 64.36 | 3/5 | 43.83 | |
| 2 | 100.00 | 34.37 | 0.00 | 0.00 | 0.00 | 2/5 | 26.87 |
| 3 | 100.00 | 34.37 | 0.00 | 35.60 | 0.00 | 3/5 | 33.99 |
| Agent systems | Reference | ||||||
| Family / metric | Claude | Codex | Kimi | OC+MM | OC+DS | Ref1 | Ref2 |
| AEOSSP (weighted) | |||||||
| valid (/5) | 5 | 5 | 5 | 5 | 5 | 5 | 5 |
| WCR (.45) | 0.570 | 0.708 | 0.692 | 0.201 | 0.682 | 0.758 | 0.682 |
| CR (.20) | 0.594 | 0.736 | 0.720 | 0.176 | 0.710 | 0.790 | 0.721 |
| TAT (s) (.20) | 9477 | 1025 | 973 | 9369 | 987 | 1018 | 1128 |
| System | No procedure | Compact | Procedure pack |
|---|---|---|---|
| Regional Coverage (weighted coverage ratio ) | |||
| OpenCode + DeepSeek V4 Pro | 0.397 | 0.415 | 0.644 |
| OpenCode + MiniMax M2.7 | 0.044 | 0.127 | 0.225 |
| Relay Constellation (service fraction ) | |||
| OpenCode + DeepSeek V4 Pro | 0.559 | 0.658 | 0.601 |
| OpenCode + MiniMax M2.7 | 0.000 | 0.124 | 0.174 |
| System | No memory | Codex-derived | DeepSeek-derived |
|---|---|---|---|
| Regional Coverage (weighted coverage ratio ) | |||
| OpenCode + DeepSeek V4 Pro | 0.397 | 0.725 | 0.646 |
| OpenCode + MiniMax M2.7 | 0.044 | 0.097 | 0.175 |
| Relay Constellation (service fraction ) | |||
| OpenCode + DeepSeek V4 Pro | 0.559 | 0.681 | 0.833 |
| OpenCode + MiniMax M2.7 | 0.000 | 0.043 | 0.110 † |
| System | Invalid | Valid-but-zero | Ungrounded |
|---|---|---|---|
| Claude | 9 | 1 | 17 |
| Codex | 0 | 0 | 1 |
| Kimi | 0 | 4 | 6 |
| OC+MM | 12 | 6 | 15 |
| OC+DS | 1 | 5 | 0 |
| All | 22 | 16 | 39 |