cs.AIAug 25, 2026

Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites

Authors: He Wang, Junyu Wu, Yeye Liu, Yifan Zhou, Jie Zhang, Hui Li, Yanjie Song, Liang Li

Organizations: College of Intelligent Science and Engineering, Harbin Engineering University, Harbin 150001, China · School of Electrical and Control Engineering, North University of China, Taiyuan 030051, China · School of Information Science and Technology, Dalian Maritime University, Dalian 116026, China

Abstract

Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target motion and satellite orbital geometry. The scheduler must jointly determine task selection, satellite assignment, observation-window selection, and observation ordering under time-window, attitude-maneuvering, and onboard-resource constraints. This paper proposes an implicit Q-learning-bootstrapped ant colony optimization method, termed IQACO, for multi-satellite maritime moving-target observation scheduling. Rather than directly learning a task-selection policy, IQACO embeds an offline implicit Q-learning module into constructive ant colony optimization to adaptively adjust the pheromone factor, heuristic factor, and evaporation rate. A compact search-state representation captures pheromone distribution, current and historical-best solution quality, and iteration progress. During online scheduling, ant colony optimization constructs feasible observation sequences, while the learned policy adjusts the search behavior according to the current search state. Experiments on 14 scenarios with different scales and satellite configurations show that IQACO consistently outperforms the compared algorithms, improving the mean objective value over conventional ant colony optimization by 2.86%-9.41%. Further comparative and supplementary experiments demonstrate its effectiveness and robustness across different scheduling conditions and problem settings. These results indicate that offline value learning provides an effective adaptive search-control mechanism for constrained maritime moving-target observation scheduling.

Figures & tables

Explore similar work

CardsList
  1. EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling

    Apr 28, 2026Qian Yin, Jiaxing Li, Jiaqi Cheng +23Mixed-Integer Programming

  2. Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling

    Jul 28, 2026Itai Zilberstein, Pranav Rajbhandari, Steve Chien +1Decentralized OptimizationDynamic Pricing

  3. Dueling DDQN-Based Adaptive Multi-Objective Handover Optimization for LEO Satellite Networks

    May 4, 2026Po-Heng Chou, Chiapin Wang, Chung-Chi Huang +1Low Earth Orbit Satellite NetworksDeep Q-Networks