cs.LGOct 4, 2026

How Execution Assumptions Change Short-Horizon Sharpe Rankings: Evidence from a Synthetic Trading Benchmark

Authors: Weicheng Xue

Organizations: Virginia Tech

Abstract

Backtests of LLM trading agents often assume that every order fills at the closing price. We ask whether this choice changes only reported returns or also the order of the agents. Five prompted LLM signal policies and seven classical baselines trade the same synthetic price paths under six execution settings, from near-ideal fills to latency, spread, participation, and impact stresses. The main experiment contains 2,4622{,}462 runs with matched decision frequencies and paired market paths. On the compressed two-asset board, agreement between the near-ideal and default-stress rankings falls to Kendall τb=0.21τ_b=0.21 in the high-volatility regime, compared with 0.820.82 in the calm regime. The seed-bootstrap intervals, [0.00,0.52][0.00,0.52] and [0.48,0.94][0.48,0.94], are wide and overlap. On a fixed 11-policy board, agreement rises from 0.24 with two assets to 0.85 with ten; the two-asset point estimate differs substantially from the wider settings we tested. Rank changes are related to turnover, and comparisons with buy-and-hold also depend on how that anchor is initialized. The experiment does not compare LLM trading skill. It shows that, on a short horizon, an execution convention can become part of the benchmark's headline. Execution assumptions and rank stability should be reported alongside returns.

Figures & tables

Explore similar work

CardsList
  1. Execution Realism and Reproducibility in LLM-Based Trading Systems: A Systematic Scoping Review and Evidence Audit

    Jun 6, 2026Junyi Yao, Zihao Zheng, Baichuan Li +1TradingReproducibility

  2. Agentic Trading: When LLM Agents Meet Financial Markets

    May 19, 2026Yihan Xia, Panpan You, Taotao Wang +4Large Language Model AgentsEvaluation Protocols