Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challenges, including nonlinear payoffs and multi-leg strategy construction, requiring structured decisions rather than simple directional bets. We introduce LiveOption, an evaluation framework for LLM-based agents in option trading. LiveOption formulates the problem as structured sequential decision-making under realistic execution and capital constraints, and provides a reproducible environment with standardized interaction protocols. The framework includes three task suites covering portfolio overlays, event-driven earnings trading, and 0DTE intraday trading. We further propose a hierarchical metric suite that evaluates action validity, decision quality, risk characteristics, and outcome-level performance. Experiments show that current agents often fail to achieve competitive returns in most scenarios. LiveOption offers a principled testbed for evaluating structured decision-making beyond outcome-based metrics.
Figures & tables
Model
Hedging
Covered Call
Δ Ret
Δ MDD
Δ Ret
Δ MDD
DeepSeek-V4-Flash
− 6.15
3.80
0.34
3.85
GPT-OSS-120B
− 3.05
1.45
− 7.38
− 0.42
Qwen3-235B-A22B
− 7.25
− 0.52
− 2.79
0.09
Llama-3.3-70B-Instruct
− 7.86
6.87
− 1.37
− 0.90
MiniMax-M3
− 6.34
2.76
− 9.36
0.86
Table 1 : Overlay performance against the matched buy-and-hold portfolio, five assets per model and mandate. Δ Ret is mean active return and Δ MDD the mean drawdown reduction (%).
Hedge
Income
Model
Fills
Moneyness
DTE
Fills
Moneyness
DTE
DeepSeek-V4-Flash
511
0.970 ± 0.079
23.2 ± 13.4
222
1.057 ± 0.075
25.9 ± 13.9
GPT-OSS-120B
807
0.988 ± 0.057
14.1 ± 11.6
852
1.037 ± 0.055
12.5 ± 9.3
Qwen3-235B-A22B
703
0.951 ± 0.068
16.7 ± 15.2
397
1.033 ± 0.051
14.2 ± 15.6
Llama-3.3-70B-Instruct
588
0.979 ± 0.053
6.4 ± 5.4
164
1.012 ± 0.037
3.3 ± 4.3
MiniMax-M3
789
0.955 ± 0.078
27.1 ± 12.9
405
1.063 ± 0.076
23.2 ± 14.0
Table 3 : Overlay instrument choice by mandate ( μ±σ over option fills). Fills counts executed option contracts over the five assets, Moneyness is strike over underlying at fill ( K/S ), and DTE is days to expiry at entry against prescribed bands of 7–45 days (Income) and 7–60 days (Hedge). Every hedge leg in the record is a put and every income leg a call, so a call-share column is omitted.
Model
Total fills
Per episode
Hedge
Income
Earnings
0DTE
Earnings
0DTE
DeepSeek-V4-Flash
511
222
616
3,274
6.6
32.1
GPT-OSS-120B
807
852
682
3,427
7.2
33.6
Qwen3-235B-A22B
703
397
803
2,519
8.5
24.7
Llama-3.3-70B-Instruct
588
164
678
2,397
7.5
23.5
MiniMax-M3
789
405
572
1,499
6.0
14.7
Table 4 : Trading activity by model and scenario, as executed option fills in the selected main runs. The two right-hand columns normalise by episode, per earnings case and per 0DTE session; 0DTE totals are the per-session mean over the 102 selected sessions.
Model
0DTE Intraday
Earnings Bet
Shape
BSM attribution
Shape
BSM attribution
↑
↓
q90
δ
γ
θ
ν
Resid.
↑
↓
q90
δ
γ
θ
ν
Resid.
DeepSeek-V4-Flash
2
0
+3.6
− 19
+123
− 163
+17
− 32
2
0
+13.5
+553
+9659
− 1642
+11
− 8737
GPT-OSS-120B
10
2
+9.8
+73
+211
− 271
+32
− 32
2
0
+10.6
+659
+20157
− 1503
+260
− 19642
Qwen3-235B-A22B
0
0
+2.3
− 9
+86
− 133
+21
− 18
3
0
+9.3
+239
+32689
− 3108
+658
− 31340
Llama-3.3-70B-Instruct
3
0
+5.1
− 36
+120
− 186
+14
− 16
2
0
+11.7
+750
+271
− 1681
+202
+279
Table 5 : Episode-level return shape and BSM PnL attribution for 0DTE Intraday (left) and Earnings Bet (right). ↑/↓ count positive/negative tail episodes, q90 is the 90th-percentile return, and attribution reports mean PnL per episode by Greek plus residual. Detail metrics are reported in Tables 21 and 27 .
Ablation
Task
Δ Ret (pp)
Fills
NT%
MDD (%)
min p
Fixed chain
0DTE
− 0.90
29.5 → 26.0
0.0
4.4 → 4.2
0.085
No portfolio state
0DTE
+2.36
29.5 → 16.1
0.0
4.4 → 12.4
0.698
No task rules
0DTE
+0.31
29.5 → 30.7
0.0
4.4 → 4.5
0.211
No news
Earnings
− 0.31
7.1 → 7.3
4.2
7.0 → 8.0
1.000
No IV
Earnings
− 2.82
7.1 → 7.3
5.6
7.0 → 8.3
0.059
No Greeks / multi-TF
0DTE
+0.25
29.5 → 4.4
47.2
4.4 → 1.0
0.835
Table 6 : Targeted context ablations averaged over three models (24 paired cases each). Δ Ret is the paired return change; Fills and MDD show baseline → ablation; NT is the no-trade rate; and p is the minimum Holm-adjusted sign-flip p -value. Under No Greeks , lower drawdown reflects reduced trading rather than better risk control. Per-model results are in Table 29 .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Greek
Definition
Risk Dimension
Key Scenario(s)
Delta ( δ )
∂V/∂S
Directional exposure
Overlay
Gamma ( γ )
∂2V/∂S2
Price convexity
0DTE
Theta ( θ )
−∂V/∂T
Time decay
0DTE, Overlay
Vega ( ν )
∂V/∂σ
Volatility level
Earnings Bet
Rho ( ρ )
∂V/∂r
Interest rate
(negligible in short-dated)
Appendix
Table 7 : Summary of option Greeks and their relevance to LiveOption evaluation scenarios.
Name
Construction
Payoff / Profit Intuition (at T )
Long Call
Buy 1 call (K,T) , pay c
Profit =(ST−K)+−c ; bullish with limited loss and uncapped upside.
Long Put
Buy 1 put (K,T) , pay p
Profit =(K−ST)+−p ; bearish protection with limited loss and downside convexity.
Covered Call
Long stock + sell 1 call (K,T)
Collects premium; upside capped above K , downside partially cushioned by premium.
Protective Put
Long stock + buy 1 put (K,T)
Downside bounded (insurance-like); sets a floor near K on portfolio value.
Cash-Sec. Put
Sell 1 put (K,T) (cash-backed)
Receives premium; profits if ST>K ; risk of buying stock at K on drops.
Bull Call Spread
Buy call K1 , sell call K2 ( K1<K2 )
Moderately bullish; bounded profit/loss; reduces cost but caps upside.
Appendix
Table 8 : Common option strategies and their constructions. Payoff intuition is described at expiration T ; profit equals payoff minus net premium, and short legs flip the corresponding payoff.
View
Source fields
Market path and chain
Underlying bars and option-chain snapshots from the query layer, at the decision timestamp
News and event triggers
Time-aligned news and calendar entries from the event layer
Agent context and rationale
The observation packet and thesis text recorded per decision step
Orders and fills
Submitted intents, verification outcome, and the resulting fills with price, quantity and fees
Portfolio and risk
Cash, equity, open positions, margin usage and per-position Greeks from the portfolio record
Performance panels
Episode return, drawdown, win rate and exposure, recomputed from the portfolio series
Appendix
Table 9: Web interface views and the recorded fields each one reads.
Setting
Overlay
Earnings Bet
0DTE Intraday
Universe
QQQ, NVDA, AMD, GOOG, TSLA
24 single names across mega-cap tech, semis, software, consumer, fintech (Tab. 11 )
Highly volatile around earnings; retail-flow driven
High Beta / Growth
RKLB
Elevated volatility and strong reaction to earnings surprises
Appendix
Table 11 : Tradable universe in the Earnings Bet scenario.
Ticker
Tot. Ret.
Ann. Ret.
Ann. Vol.
Sharpe
Skew.
Kurt.
Best
Worst
MDD
(%)
(%)
(%)
(%)
(%)
(%)
QQQ
20.40
20.67
23.62
0.88
1.31
17.81
12.00
−6.21
−22.88
NVDA
34.84
35.33
49.64
0.71
−0.08
8.00
18.72
−16.97
−36.89
AMD
77.54
78.77
60.73
1.30
1.79
10.68
23.82
−8.90
−39.63
GOOG
64.61
65.60
31.99
2.05
0.33
3.99
9.88
−7.51
−29.43
TSLA
18.57
18.82
63.31
0.30
0.41
4.79
22.69
−15.43
−48.19
Appendix
Table 12 : Performance metrics and risk characteristics of selected technology tickers. Returns and volatility are annualized; Sharpe ratio is calculated assuming a risk-free rate of zero ( rf=0 ). Metrics are derived from daily closing prices for the 2025 fiscal year.
Q1
Q2
Q3
Q4
Overall
Ticker
Open
Close
Open
Close
Open
Close
Open
Close
Open
Close
AAPL
−3.20
−4.03
−4.79
−6.76
−1.48
−2.03
−0.36
−0.87
−2.46
−3.42
AMD
−7.17
−7.82
3.35
3.12
−4.29
−1.10
1.37
−4.94
−1.69
−2.68
AMZN
−3.47
−2.38
−1.94
−2.02
−7.14
−9.59
14.58
13.97
0.51
−0.01
AVGO
5.66
2.79
−5.73
−6.02
11.97
12.92
−10.95
−16.38
0.24
−1.67
CELH
27.09
22.88
3.72
5.61
19.68
21.55
−23.72
−30.71
6.69
4.83
Appendix
Table 13 : Earnings-announcement price reactions by ticker in 2025. Values represent percentage changes ( Δ% ) aligned by decimal point. Open denotes the reaction from T−1 close to T+1 open; Close denotes T−1 close to T+1 close. Information is curated from evaluated technical reports.
Return (%)
GK Volatility (%)
Group
N
Mean
Std
P25
P75
Mean
Std
P25
P75
Overall
250
0.0911
1.2558
−0.3760
0.4846
0.8214
0.8630
0.4289
0.8557
Wednesday
52
0.2181
1.4654
−0.2520
0.4612
0.8762
1.1162
0.4289
0.8599
Friday
50
−0.0409
0.9903
−0.3955
0.5506
0.7643
0.4816
0.4389
0.8477
Appendix
Table 14 : Summary statistics of daily returns and GK volatility by day-of-week group over 250 trading days. All values are in percentage (%).
Metric
Wed.
Fri.
Overall
Mean intraday range (%)
1.2697
1.1902
1.1791
Mean absolute return (%)
0.6677
0.6916
0.6115
Skewness
5.3916
− 1.2568
2.6255
Kurtosis
30.239
− 2.636
32.623
Up >2σ
1 (1.9%)
1 (2.0%)
3 (1.2%)
Up >3σ
1 (1.9%)
0 (0.0%)
2 (0.8%)
Appendix
Table 15: Intraday and tail-risk statistics by day-of-week group. Tail counts report observations exceeding the stated multiple of the full-sample standard deviation, with the group percentage in parentheses.
Model
Ret%
Δ Ret
Sharpe
MDD%
Δ MDD
IR
Fills
Fees ($)
(a) QQQ buy-and-hold 17.00%, maximum drawdown 19.24%
Hedge
DeepSeek-V4-Flash
10.38
− 6.62
0.669
16.62
2.62
− 0.93
95
3,255
GPT-OSS-120B
15.95
− 1.05
1.011
14.06
5.18
− 0.28
131
3,788
Qwen3-235B-A22B
11.66
− 5.33
0.690
19.95
− 0.71
− 1.74
71
2,280
Llama-3.3-70B-Instruct
12.30
− 4.70
0.943
14.44
4.81
− 0.43
132
1,766
Appendix
Table 17 : Overlay backtest performance per asset, full-year 2025. Δ Ret is active return against the matched buy-and-hold portfolio, Δ MDD the reduction in maximum drawdown (positive means shallower than the benchmark), IR the information ratio, and Fees the total commission and exchange fees paid. Each sub-heading gives that asset’s buy-and-hold return and maximum drawdown, which are the benchmark the Δ columns are measured against.
QQQ
NVDA
AMD
GOOG
TSLA
Δ Ret
Δ MDD
Hedging
Protective put rule
-9.1
-11.6
+3.2
-7.1
-15.2
-7.97
+8.20
LLM panel mean
-4.9
-12.1
-0.8
-6.1
-12.5
-7.28
+3.78
panel − rule
+4.2
-0.5
-4.0
+1.0
+2.7
+0.69
3/5
Covered Call
Covered call rule
-7.7
-16.8
-47.0
-47.5
+8.8
-22.04
+3.57
Appendix
Table 18: Mechanical overlay baselines against the LLM panel. Each baseline runs the paper’s own overlay configuration with only the decision layer replaced by a fixed 0.30 -delta rule, so universe, capital, dates, rebalance cadence, risk rules and fill model are identical to the agent runs. Cells are active return against the matched buy-and-hold in percentage points; Δ Ret and Δ MDD average over the five assets, and the last row of each block is the panel mean minus the rule on the same asset, with the count of assets on which the panel is ahead.
Month
Model
Sess.
Win%
PnL%
Mean%
P/L
PF
Best%
Worst%
2025-01
DeepSeek-V4-Flash
9
33.3
− 10.97
− 1.22
0.45
0.23
+2.07
− 4.59
GPT-OSS-120B
9
33.3
+13.55
+1.50
3.14
1.57
+21.87
− 7.93
Qwen3-235B-A22B
9
33.3
− 5.29
− 0.59
0.94
0.47
+3.03
− 2.75
Llama-3.3-70B-Instruct
9
11.1
− 19.69
− 2.19
1.01
0.13
+2.86
− 5.50
MiniMax-M3
9
44.4
− 1.87
− 0.21
0.91
0.73
+2.49
− 3.72
GLM-5.3-Flash
9
55.6
+3.17
+0.35
1.10
1.37
+5.06
− 3.82
Appendix
Table 19 : Per-month, per-model intraday 0-DTE SPY performance. Each session is an independent fresh-$10k episode; PnL% sums the daily realized return percentages inside the month and Overall sums every covered session in 2025. Mean%, Win%, P/L, PF, Best% and Worst% are computed over the sessions inside each month, and Fills is the per-session average.
Strategy
Model
Trades
Win%
PnL (k$)
P/L
PF
Best%
Worst%
Buy (k$)
Sell (k$)
Long Call
DeepSeek-V4-Flash
384
23.4
− 5.9
1.98
0.61
+100.2
− 78.4
83.3
78.7
GPT-OSS-120B
296
30.1
− 5.1
1.80
0.78
+130.8
− 101.0
99.9
96.0
Qwen3-235B-A22B
211
36.0
− 2.1
1.34
0.75
+203.7
− 223.9
47.3
45.9
Llama-3.3-70B-Instruct
210
20.9
− 5.8
1.89
0.50
+162.6
− 107.5
43.7
38.4
MiniMax-M3
200
37.0
+1.3
2.04
1.20
+499.3
− 100.3
31.2
32.8
GLM-5.3-Flash
282
29.8
− 0.8
2.23
0.94
+272.2
− 69.8
66.3
66.4
Appendix
Table 20 : Per-strategy-bucket, per-model intraday 0-DTE SPY performance. Each contract traded on each session is one trade, and the bucket is set by the first fill’s side and the option right (Long Call = first BUY of a call, Short Put = first SELL of a put, and so on). PnL is FIFO-matched realized PnL net of fees, with any residual at end of day closed at the contract’s last mark. Per-trade return uses the entry-side notional as denominator, so Worst% can fall well below −100 % for short trades whose closing buyback dwarfs the premium received. Buy and Sell are total notionals (qty × price × 100) in thousands of USD.
Explained
Share of attribution
Model
Pairs
Moved
Step
Session
Time
Spot
Vol
Resid
DeepSeek-V4-Flash
40,159
92%
0.92
0.87
0.04
0.76
0.13
0.06
GPT-OSS-120B
53,180
91%
0.92
0.88
0.04
0.76
0.14
0.07
Qwen3-235B-A22B
56,443
84%
0.80
0.76
0.03
0.71
0.12
0.15
Llama-3.3-70B-Instruct
63,000
88%
0.85
0.86
0.03
0.72
0.14
0.12
MiniMax-M3
39,182
90%
0.89
0.84
0.04
0.74
0.13
0.09
Appendix
Table 22: Full-repricing P&L-explain for the 0DTE track. Each consecutive step pair of each held position is revalued exactly, sequentially in time, spot and volatility, and the remainder is the residual; because this reprices rather than expands, it stays valid across jumps. Moved is the share of step pairs whose recorded mark changed, and the explained fractions are computed over those pairs, since a pair with no mark change contributes nothing to episode PnL. Explained is 1−∑∣residual∣/∑∣Δmark∣ at step level, and Session the same after signed aggregation within a session. The last four columns are each term’s share of total absolute attributed magnitude.
Ticker
Model
Cases
Mean%
Med%
Win%
PF
MDD%
Fills
AAPL
DeepSeek-V4-Flash
4
− 6.03
− 7.24
50.0
0.38
14.10
28
GPT-OSS-120B
4
− 3.08
− 2.60
25.0
0.12
4.39
32
Qwen3-235B-A22B
4
− 3.19
− 3.07
0.0
0.00
3.67
28
Llama-3.3-70B-Instruct
4
− 1.56
− 1.45
50.0
0.54
5.21
27
MiniMax-M3
4
− 4.57
− 7.49
25.0
0.33
9.11
32
GLM-5.3-Flash
4
− 1.04
− 1.53
50.0
0.72
6.29
38
Appendix
Table 25: Per-ticker earnings metrics across all models, aggregated over the quarterly earnings cases per ticker. Mean and median are per-case returns, MDD is the mean maximum drawdown within a case, Fills counts executed option contracts and NT the cases in which the model chose not to trade.
Structure
Model
Cases
Share%
Mean%
Med%
Win%
PF
bear call spread
DeepSeek-V4-Flash
3
3.2
− 4.05
− 4.37
0.0
0.00
GPT-OSS-120B
5
5.3
3.49
1.68
60.0
2.55
Qwen3-235B-A22B
1
1.1
52.37
52.37
100.0
–
Llama-3.3-70B-Instruct
3
3.3
− 10.19
− 9.91
0.0
0.00
MiniMax-M3
3
3.2
− 1.45
− 6.34
33.3
0.71
GLM-5.3-Flash
4
4.2
− 2.13
− 4.66
25.0
0.56
Appendix
Table 26 : Earnings performance by the structure opened in each case. Cases counts the cases in which the model opened that structure and Share its percentage of that model’s classified cases. Buckets a model used only a handful of times are reported for completeness but do not support a ranking.
Mean return
Worst episode
Model
n
Eps
Off
Guard
Off
Guard
Earnings
DeepSeek-V4-Flash
94
1
−0.6
−0.5
−23.8
−23.8
GPT-OSS-120B
95
0
−0.2
−0.2
−22.0
−22.0
Qwen3-235B-A22B
94
5
−1.2
−1.2
−22.8
−22.8
Llama-3.3-70B-Instruct
90
1
−0.4
−0.4
−23.2
−23.2
Appendix
Table 28: Execution-layer ablation of the inventory check on the closing path. Every recorded episode is replayed with the decision stream held fixed and inversions clamped on any contract that received a duplicate closing intent, which is the one defect this guard addresses. Eps counts episodes in which the guard binds at all; Mean and Worst are the mean and minimum episode return, with and without it, in percent.
Ablation
Task
Model
n
Δ Ret [95% CI]
Med
Worse%
NT%
MDD%
p
Fixed chain
0DTE
DeepSeek-V4-Flash
24
− 0.09 [ − 0.86, 0.59]
0.01
50.0
0.0 → 0.0
4.4 → 4.3
0.899
0DTE
GPT-OSS-120B
24
− 2.16 [ − 4.36, − 0.43]
− 0.73
62.5
0.0 → 0.0
6.3 → 5.7
0.085
0DTE
Qwen3-235B-A22B
24
− 0.44 [ − 1.53, 0.64]
− 0.22
54.2
0.0 → 0.0
2.6 → 2.7
0.899
Covered Call
DeepSeek-V4-Flash
2
− 3.68 [ − 9.83, 2.46]
− 3.68
50.0
0.0 → 0.0
22.1 → 24.6
1.000
Covered Call
GPT-OSS-120B
2
− 4.98 [ − 10.29, 0.33]
− 4.98
50.0
0.0 → 0.0
25.0 → 23.5
1.000
Covered Call
Qwen3-235B-A22B
2
− 2.15 [ − 6.48, 2.18]
− 2.15
50.0
0.0 → 0.0
24.1 → 23.7
1.000
Appendix
Table 29 : Every context ablation cell, by profile, task and model. Three profiles (Fixed chain, No portfolio state, No task rules) were run on all four tasks; the other two were defined for a single task. NT and MDD are reported as baseline → ablation. The Hedging and Covered Call cells rest on two paired cases each and are shown for completeness rather than as evidence. p is Holm-adjusted within each profile–task family. The main-text Table 6 reports one task per profile.
Task
Model A
Model B
n
Δ [95% CI]
Med
p
pHolm
Earnings
DeepSeek-V4-Flash
GPT-OSS-120B
94
− 0.28 [ − 2.27, 1.64]
0.05
0.781
1.000
Qwen3-235B-A22B
94
0.64 [ − 2.31, 3.57]
0.30
0.672
1.000
Llama-3.3-70B-Instruct
90
− 0.05 [ − 2.45, 2.31]
0.51
0.968
1.000
MiniMax-M3
94
1.28 [ − 0.79, 3.35]
0.45
0.225
1.000
GLM-5.3-Flash
94
− 0.08 [ − 2.09, 2.05]
− 0.13
0.939
1.000
GPT-OSS-120B
Qwen3-235B-A22B
94
0.92 [ − 1.97, 3.75]
0.46
0.536
1.000
Appendix
Table 30 : Pairwise model comparisons on the two episodic tasks, paired on shared episodes. Δ is the paired mean return difference of model A minus model B in percentage points with a 95% paired bootstrap interval and Med is the paired median. Model A is printed once per block. p is a two-sided Monte Carlo sign-flip p -value from 50,000 draws and pHolm is its Holm-adjusted value within the 15 pairs of that task. No pair is significant at the 5% level after correction.
A growing body of work explores how Large Language Models (LLMs) can be embedded in trading systems as agents that perceive market information, retrieve context, reason about decisions, emit tradable actions, and adapt under market feedback. This paper reframes LLM-based trading agents as expert-system decision pipelines and presents an audit-oriented evidence map of 77 included studies in a protocol-coded snapshot screened through 2026-03-09. A primary empirical subset (n=19) satisfies the minimum boundary of Action Output plus Closed-Loop Evaluation; the remaining 58 included studies are retained as background and design context. The central empirical finding is protocol incomparability: within the primary subset, only 2/19 studies report extractable time-consistent split protocols, 1/19 reports an explicit transaction-cost model, 1/19 documents universe or survivorship handling, 11/19 report execution timing or semantics, 15/19 are coded as R0, and no study reaches R3 reproducibility. We therefore use Architecture-Capability-Adaptation as a working analytical lens rather than a validated taxonomy, and we foreground the evidence ledger, reproducibility audit, and reporting checklist as the main contributions. The resulting survey shows that architectural experimentation is expanding rapidly, while comparable evaluation protocols, execution semantics, and reproducible artifacts remain the field's immediate bottlenecks.
Yihan Xia, Panpan You, Taotao Wang +4
College of Electronic and Information Engineering, Shenzhen University, Shenzhen, China
Recent deployments of large language models (LLMs) as autonomous trading agents raise questions about whether financial decision-making competence generalizes beyond specific market patterns and how it should be trained and evaluated in noisy markets lacking ground truth. We propose a structured framework for training and evaluating such models. Central to our approach is a curated, multiple-choice question (MCQ) dataset derived from classic textbooks and historical markets, verified by an AI committee, enriched with structured reasoning traces, and augmented to reduce shortcut learning. To evaluate whether performance on isolated MCQs generalizes to real-world trading, we introduce a two-stage protocol combining test-set evaluation with an MCQ-based chronological trading simulation. Extensive evaluations across market regimes provide statistically robust evidence that open models trained with our framework exhibit competitive, risk-aware behavior over time, outperform open-source baselines, and approach frontier-model performance at smaller scale. We release the dataset and evaluation framework to support further research.
Yuchen Pan, Soung Chang Liew
Department of Information Engineering, The Chinese University of Hong Kong
Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment reasoning. Second, raw returns are a noisy proxy for stock-selection ability, since positive performance may come from market beta, style exposure, or favorable regimes rather than genuine alpha. We introduce KTD-Fin (Knowing-To-Doing Financial Benchmark), an end-to-end stock-market trading benchmark that addresses both issues. KTD-Fin uses a data-side masking protocol to anonymize key identifiers and calendar information consistently across prompts and tools, separating historical market memory from investment decision-making. It also incorporates a Barra-style performance attribution framework that decomposes portfolio returns into market, style, and stock-selection alpha components. Across ten frontier LLM agents evaluated on the Chinese CSI300 over a 2024--2026 window, masking substantially changes agent rationales, pushing them towards anonymized factor-based reasoning. Attribution analysis further shows that LLM agents' cumulative returns under leakage-controlled evaluation are largely explained by passive market and style exposure, with limited evidence of persistent stock-selection alpha. These findings suggest that financial LLM benchmarks should evaluate not only whether an agent makes money, but also whether the source of returns reflects transferable investment skill. We release KTD-Fin as a reproducible template for leakage-controlled and attribution-aware evaluation of LLM trading agents.