cs.LGAug 10, 2026

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

Authors: Logan LunaJuan Ortiz CouderRaul Alejandro Vargas-Acosta

Organizations: Georgia Institute of Technology · Department of Electrical Engineering and Computer Science, Embry-Riddle Aeronautical University, Daytona Beach, FL 32114 USA · Present address: School of Computer Science, College of Computing, Georgia Institute of Technology, Atlanta, GA 30332 USA · Embry-Riddle Aeronautical University

Abstract

Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO). However, these events have been growing in frequency as orbital congestion worsens with the launch of megaconstellations. Consequently, conjunction alerts and collision risks are becoming increasingly common. Current practices, which are commonly manual or rule-based, have difficulty scaling to these worsening dynamic environments. To address this intensifying situation, we propose a reinforcement-learning policy for autonomous collision avoidance, trained via Proximal Policy Optimization (PPO) along with an open-source, high-fidelity astrodynamics simulator for training and evaluation. In 1,000 deterministic GEO episodes, our agent achieves a 97.5% collision avoidance success rate, outperforming traditional controllers such as a rule-based baseline (20.7% success) and an impulsive delta-v planner baseline (27.5% success). To achieve these results, we designed a simulator to train and evaluate our agent, using real-world and simulated debris. We simulate Newtonian two-body dynamics using Sun/Moon third-body perturbations, fuel-dependent thrust, and configurable debris fields. The agent is trained with curriculum learning and shaped rewards oriented toward encouraging survival, adequate projected miss distance, and delta-v conservation. Finally, our evaluation consisted of a fully deterministic pipeline, including shared seeds, per-episode logs, and telemetry exports. Our work is a publicly available framework at https://purl.org/sat-trajectory-avoidance

Explore similar work

CardsList
  1. Time-Optimal Collision Avoidance Via a Greedy Polynomial Backward Sweep

    May 31, 2026Zeno Pavanello, Frank De Veld, Roberto ArmellinCollision Avoidance