cs.NIApr 28, 2026

EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling

Authors: Qian YinJiaxing LiJiaqi ChengQizhang LuoAnnalisa RiccardiAbhijit ChatterjeeRafael VazquezCarlo Novara+18 more

Organizations: School of Traffic and Transportation Engineering, Central South University, Changsha 410083, China · School of Engineering and Materials Science, Queen Mary University of London, London E1 4NS, UK · College of Automation, Central South University, Changsha, 410083, China · Mechanical and Aerospace Engineering, University of Strathclyde, Glasgow G1 1XQ, UK · Department of Computer Science, University of Exeter, Exeter EX4 4QJ, UK · Department of Aerospace Engineering, Universidad de Sevilla, Camino de los Descubrimientos s.n., Sevilla, 41092, Spain · Department of Electronics and Telecommunications, Politecnico di Torino, Corso Duca degli Abruzzi, 24, Turin, 10129, Italy · ERATOSTHENES Centre of Excellence, Limassol, 3012, Cyprus · Department of Civil Engineering and Geomatics, Cyprus University of Technology, Limassol, 3036, Cyprus · Department of Computer Science and Engineering, College of Engineering, Qatar University, Doha, 2713, Qatar · Department of Aerospace Engineering, Korea Advanced Institute of Science and Technology, Daejeon, 34141, South Korea · School of Management, Hefei University of Technology, Hefei, 230009, China · Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University, Xian 710071, China · School of Astronautics, Beihang University, 102206 Beijing, China · Advanced Space Technology Laboratory, College of Astronautics, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, China · National Key Laboratory of Aerospace Flight Dynamics, Northwestern Polytechnical University, Xian, 710072, China · State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan, 430079, China · School of Computer Science, China University of Geosciences, Wuhan, 430074, China · School of Information Science and Technology, Dalian Maritime University, Dalian, 116026, China · Department of Electrical & Computer Engineering, University of Alberta, Edmonton, AB T6R 2V4, Canada · Division of Geological and Planetary Science, California Institute of Technology, CA, USA · Department of Automation and Systems Engineering, Federal University of Santa Catarina, Florianopolis, SC, Brazil · School of Geography, Archaeology and Environmental Studies, University of the Witwatersrand, Braamfontein, Johannesburg, South Africa

Abstract

Earth observation satellite imaging scheduling is a challenging NP-hard combinatorial optimisation problem central to space mission operations. While next-generation agile Earth observation satellites (EOS) increase operational flexibility, they also significantly raise scheduling complexity. The lack of a unified, open-source benchmark makes it difficult to compare algorithms across studies. This paper introduces EOS-Bench, a comprehensive framework for systematic and reproducible evaluation of scheduling methods. By integrating high-fidelity orbital dynamics and platform constraints, EOS-Bench generates 1,390 scenarios and 13,900 benchmark instances, spanning from small-scale validation cases to large coordination problems with up to 1,000 satellites and 10,000 requests. We further propose a scenario characterisation scheme to quantify structural difficulty based on factors such as opportunity density, task flexibility, conflict intensity, and satellite congestion. A multidimensional evaluation protocol is introduced, assessing performance across five metrics: task profit, completion rate, workload balance, timeliness, and runtime. The framework is evaluated using mixed-integer programming, heuristics, meta-heuristics, and deep reinforcement learning across both agile and non-agile settings. Results show that EOS-Bench effectively distinguishes solver performance across scales and conditions, revealing trade-offs between solution quality and computational efficiency, and providing deeper insight into scenario complexity. EOS-Bench offers a unified and extensible open testbed for advancing research in Earth observation satellite scheduling. The code and data are available at https://github.com/Ethan19YQ/EOS-Bench.

Explore similar work

May 26, 2026cs.AI

DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

Progress in neural combinatorial optimization for Dynamic Flexible Job Shop Scheduling Problem (DFJSP) is currently hindered by a methodological tension: static benchmarks encourage benchmark overfitting, while uncalibrated generators obscure algorithmic capability with stochastic noise. To resolve this, we introduce \textbf{DynaSchedBench}, a diagnostic framework for DFJSP that rigorously controls the instance-generation process. Instead of relying on parameter sampling, our approach utilizes Sequential Event-Space Calibrator (SESC) that computes a novel Schedule Stress Index (SSI) to stratify instances by difficulty. We demonstrate that SESC is substantially more computationally efficient than evolutionary baselines while converging reliably to the target metrics. The framework integrates modular components for instance generation, snapshot-based simulation, agents, evaluation, and visualization, thereby enabling rigorous testing of reactive and lookahead-based policies. Leveraging this calibrated environment, we identify key limitations of LLM-based scheduling agents. Specifically, in step-wise online decision-making for dynamic scheduling, we identify an ``Observability Paradox'': providing agents with oracle access to full structural information can degrade policy performance, underperforming concise information. Furthermore, despite substantial token overhead, tool-augmented and refinement strategies fail to reliably improve performance, and most LLM agents fail to consistently surpass strong dispatching baselines-behaving more like robust heuristic approximators than superior optimizers.
Shijie Cao, Yuan Yuan, Jing Liu
Jun 20, 2026astro-ph.IM

Deep RL for Fast Long-Horizon Operations Scheduling on NASA's Carruthers Geocorona Observatory Mission

Spacecraft operations scheduling is a highly constrained, long-horizon combinatorial optimization problem that traditionally relies on heuristics, constraint programming, or manual planning. We present a scalable deep reinforcement learning framework developed and deployed for NASA's Carruthers Geocorona Observatory mission. Our framework introduces a macro-action abstraction known as activity blocks coupled with dynamic action-masking to navigate the intractably large search space and strictly enforce complex power, thermal, and instrument constraints. The resulting architecture generates globally feasible schedules with overwhelming probability, establishes operational trust, and executes a full training cycle in under six hours, circumventing the need for policy robustness by enabling rapid, on-demand retraining. Further, resulting schedules outperform baseline heuristics in scheduled science quality. The deep reinforcement learning framework was deployed as the default operational scheduler for the Carruthers Geocorona Observatory mission from the outset of the mission, demonstrating that deep reinforcement learning can be trusted for real spacecraft operations under complex, evolving constraints.
Alex Zhang, Jackson Craig, Lara Waldrop
Jun 18, 2026cs.AI

ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear. Existing OR evaluations often decouple modeling from solving, rely on pre-formalized or text-only instances, and rarely test the full workflow from operational artifacts to validated decisions. In this work, we introduce ORAgentBench, an execution-grounded benchmark for evaluating autonomous agents on challenging end-to-end operations research tasks. It contains 107 human-reviewed tasks across diverse operational scenarios, each packaged in an isolated environment with a natural-language brief, multi-file data, configuration artifacts, and a required submission schema. Agents must write and run solution code, and their submissions are evaluated by hidden validators for schema validity, hard-constraint feasibility, and normalized objective quality. Experiments with fourteen frontier agent-model configurations show that current agents remain far from reliable OR practice. The best agent passes only 35.51% of all tasks and 20.59% of hard tasks, and many feasible submissions still fall below the required quality threshold. Failure analysis further shows that errors are dominated by strategic weaknesses, including missed operational rules, brittle formulations, weak feasible-solution construction, and insufficient solution improvement. OR-specific procedural skills increase hard-task feasibility, but do not reliably improve solution quality or pass rate. These results suggest that progress in OR agents requires moving beyond plausible optimization code toward dependable, high-quality operational decision-making.
Jiajun Li, Mingshu Cai, Yixuan Li +5