q-fin.TRSep 30, 2026

Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games

Authors: Christos Spyridon Koulouris, Carlo Campajola

Organizations: Institute of Finance and Technology University College London Gower Street, London WC1E 6BT, United Kingdom · UZH Blockchain Center Andreasstrasse 15, 8050 Zürich, Switzerland

Abstract

In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

Figures & tables

Explore similar work

CardsList
  1. Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution

    May 19, 2026Christos Spyridon Koulouris, Carlo CampajolaMulti-Agent Reinforcement LearningSingle Trajectory

  2. Mitigating Retaliatory Algorithmic Collusion in Repeated Games

    Sep 17, 2026Karthik Sivachandran, Rohan PalejaCollusionQ-Learning