Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games
Authors: Christos Spyridon Koulouris, Carlo Campajola
Organizations: Institute of Finance and Technology University College London Gower Street, London WC1E 6BT, United Kingdom · UZH Blockchain Center Andreasstrasse 15, 8050 Zürich, Switzerland
In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.
Figures & tables
Figure 1. Testing implementation shortfall for the two players. Each circle averages 500 test episodes from one trained pair; stars mark the closed-loop Nash and joint TWAP benchmarks. Lower values on both axes indicate a joint improvement in execution costs. Ten run centroids lie below the Nash benchmark in both players' implementation shortfall, but above the TWAP benchmark.
Figure 2. Mean testing inventory paths (left) and trade quantities (right), compared with closed-loop Nash. The shaded region is the full pointwise range across both players, ten trained runs and 500 test episodes per run. It measures realised trajectory dispersion. Mean inventory and trading paths with a minimum-to-maximum envelope covering every deterministic testing trajectory, compared with the closed-loop Nash schedule. Learned liquidation is slower initially, and the envelope is narrow.
Figure 3. Punisher’s trade under the learned response to the deviation minus its trade along the learned supra-competitive path without the deviation. The line averages these differences over both role assignments, 500 paired test episodes and ten trained runs; shading spans the full range of run means. Grey and hatched regions mark the deviation and compulsory final liquidation. The punisher sells more in timesteps two to five than in the unforced episode, then sells less at later timesteps. Its first trade is unchanged.
Figure 4. Execution-cost differences given the same forced deviation: learned punishment minus no punishment (orange), and sampled alternative minus no punishment (green).Positive values indicate a loss; negative values indicate a gain. The learned response has a near-zero cost effect on the punisher and a positive cost effect on the deviator. The alternative schedules lower both players' costs relative to learned punishment, while the deviator's cost remains above the replay control.
Figure 5. Occurrence of the tested first-step deviation during training. Orange counts matching episodes; blue counts the subset followed by additional selling from the other player. Counts pool a 200-episode trailing window from each of ten runs. Faint curves show rolling counts; bold curves are centred averages over 2,000 successive window endpoints. Matching deviations and the subset followed by additional selling rise during training, peak near episode eighteen thousand, and then decline.
Figure 6. Deviator gain without punishment, G (blue), and loss from punishment, H (orange), in IS/N . Connected points pair values from the same trained run, averaging both player identities and all test episodes. Boxes summarise the ten run means. The punishment loss exceeds the gain in every run. For every trained run, the orange loss point lies to the right of its connected blue gain point.
Figure 7. Measured exposure-weighted TV and its required boundary from ( 18 ). Connected points compare the two quantities within each trained run; boxes summarise their distributions and black diamonds show pooled values. The measured change exceeds the boundary in every run. The measured timing change lies above its required boundary in all ten trained runs and in the pooled comparison.
Institute of Finance and Technology, University College London, Gower Street WC1E 6BT London, United Kingdom · UZH Blockchain Center, Andreasstrasse 15, 8050 Zürich, Switzerland