Organizations: The Chinese University of Hong Kong · University of Science and Technology of China · Jilin University · RMIT University · Infplane Computing Lab · Singapore Management University
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at https://github.com/Starlien95/Awesome-TradingAI.
Figures & tables
Method
Cum. Ret. (%) ↑
Alpha (%) ↑
Sharpe ↑
Max DD (%) ↓
Vol. (%) ↓
Buy&Hold
-7.63
0.00
-0.18
33.01
41.74
LightGBM
8.57
13.70
0.31
52.05
42.14
CatBoost
21.46
25.61
0.57
61.93
58.07
Linear
21.08
25.87
0.63
44.59
51.30
XGBoost
41.93
53.60
1.08
42.42
59.50
MLP
30.33
39.15
1.08
35.24
43.06
Table 1: Backtest performance of the evaluated trading methods from 2025-01-01 to 2025-12-31.
Figure 2
Method
Cum. Ret. (%) ↑
Alpha (%) ↑
Sharpe ↑
Max DD (%) ↓
Vol. (%) ↓
BT
PT
BT
PT
BT
PT
BT
PT
BT
PT
Buy&Hold
30.35
0.00
1.88
13.56
41.78
XGBoost
33.95
27.53
64.23
45.32
2.90
2.59
12.01
9.56
37.88
32.15
MLP
-45.56
-46.79
-256.24
-251.80
-4.69
-4.77
49.63
48.97
45.52
44.03
TabNet
-54.37
-50.80
-257.37
-237.74
-10.30
-8.34
56.27
52.72
24.91
27.67
TCN
-35.24
-33.20
-191.51
-188.46
-3.42
-3.24
37.76
37.45
42.30
42.81
Table 2: Backtest (BT) and paper trading (PT) performance.
Figure 3: Cumulative returns of backtesting and paper trading.
Method
Latency
Price Diff.
Order Value
(s)
(‰)
Dev. (‰)
XGBoost
14.57
0.31
0.41
MLP
13.31
0.35
0.22
TabNet
18.17
0.50
0.63
TCN
16.89
0.33
0.22
LSTM
27.98
0.40
0.30
Table 3: Paper trading execution discrepancies.
Method
Cum. Ret. (%) ↑
Alpha (%) ↑
Sharpe ↑
Max DD (%) ↓
Vol. (%) ↓
BT
PT
LT
BT
PT
LT
BT
PT
LT
BT
PT
LT
BT
PT
LT
Buy&Hold
11.63
0.00
0.00
0.00
1.59
1.59
1.59
7.48
40.66
LSTM
-2.25
-0.48
-0.33
-54.89
-107.36
-90.88
-1.96
-5.53
-4.88
8.12
9.29
7.66
24.86
18.75
17.25
TRA
-9.21
-10.82
-14.74
-160.46
-235.95
-284.03
-2.18
-3.11
-4.41
18.38
19.14
21.49
58.71
58.26
51.85
XGBoost
-3.60
0.34
-0.79
-9.51
-52.32
-57.90
-0.11
-1.31
-1.59
10.24
8.73
9.00
28.54
27.74
25.20
Qwen
6.18
10.10
3.72
-0.63
15.80
-28.10
1.34
1.85
0.34
4.85
7.27
6.06
29.91
39.28
32.02
Table 4: Backtest (BT), paper trading (PT) and live trading (LT) performance.
Figure 4: Cumulative returns of backtesting, paper trading and live trading.
Method
Latency
Price Diff.
Order Value
(s)
(‰)
Dev. (‰)
LSTM
10.63
2.63
1.23
TRA
14.23
2.28
0.50
XGBoost
13.57
1.39
0.41
Qwen
151.00
1.13
0.27
Table 5: Live trading execution discrepancies.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Method
Category
Information
Trading Decision
Market Data
Factors
News
Asset Sel.
Direction
Risk Ctrl.
Buy&Hold
Passive
×
×
×
×
Long-only
×
LightGBM
ML
×
✓
×
✓
Long-only
×
CatBoost
ML
×
✓
×
✓
Long-only
×
Linear
ML
×
✓
×
✓
Long-only
×
XGBoost
ML
×
✓
×
✓
Long-only
×
Appendix
Table 6: Overview of the evaluated trading methods.
Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling. This article presents a targeted topical review and reproducibility audit of execution realism in LLM-based trading research. A coded evidence matrix covering 30 trade-relevant primary studies is used to assess point-in-time controls, split transparency, held-out evaluation, cost and turnover treatment, execution semantics, universe definition, and artifact release. Across the audited sample, architecture reporting is generally clearer than the evaluation assumptions needed to judge whether a trading result is economically interpretable or reproducible. A 10-equity worked example is included only as a methodological scaffold to illustrate how explicit friction and timing choices can materially compress active-strategy results. The main conclusion is that the next useful step for LLM trading research is not only better agent design, but also clearer reporting standards for execution realism, reproducibility, and evaluation comparability.
The intersection of crypto x AI is spawning papers, products, online posts, and companies. All the surrounding buzz, though, obscures what exactly has been done, what the opportunities and challenges are, and what open questions deserve attention. This survey paper asks what AI can do for blockchain-based technologies (broadly construed as "crypto") (crypto x AI), and vice versa (AI x crypto). We systematize existing work, summarize key takeaways, highlight open research questions, and offer a perspective on pervasive industry misconceptions, concluding that AI and crypto are still in the very early stages of meaningful integration.
Sarah Allen, Pranay Anchuri, James Austgen +22
1Initiative for CryptoCurrencies and Contracts (IC3) · 5Flashbots · 6Offchain Labs +11
Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A deterministic multiple-choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three difficulty tiers, with an independent checker that re-derives every answer. A generator-solver filtering pipeline autonomously mines harder questions: a generator writes questions verified by executable code, converts them to MCQs, and discards any that a no-tool solver can answer without code execution. We evaluate 11 models without tools (10 runs each) and four with-tools configurations on a 30-question curated set. Tool-augmented agents reach 90.0% accuracy in a single pass (GPT-5.5 and Opus 4.7), outperforming the best no-tools baselines (73.0%, averaged over 10 runs) by 17 percentage points. On 38 separately mined questions, no-tools accuracy drops further, with half the models falling to roughly random-chance level (25%). Beyond evaluation, the scalable MCQ infrastructure is designed to produce a training corpus for reinforcement learning, with the ultimate goal of building a specialized agent for quantitative trading workflows.