SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
Organizations: Carnegie Mellon University · Amazon
Abstract
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.
Figures & tables
| Phase | Dates | Role |
|---|---|---|
| Supervised fine-tuning | 2024-09-03 to 2024-11-29 | Frontier-model generation and supervised fine-tuning |
| Reinforcement learning | 2024-12-02 to 2025-02-28 | Reinforcement learning to optimize the policy |
| Test | 2025-03-03 to 2025-08-29 | Out-of-sample evaluation of the policy and baselines |
| Profit | Risk-adjusted profit | Risk | Trade | |||||
| Policy | TR(%) | ASR | ACR | ASoR | AVOL(%) | MDD(%) | WR(%) | PLR |
| Rule-based policies | ||||||||
| Equal-weighted rule | ||||||||
| GARCH rule | ||||||||
| Threshold rule | ||||||||
| Machine-learning policy | ||||||||
| Factor | Profit | Risk-adjusted profit | Risk | Trade | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Variant | RL | News | TR(%) | ASR | ACR | ASoR | AVOL(%) | MDD(%) | WR(%) | PLR |
| Full method | ||||||||||
| SOTA | ✓ | |||||||||
| Ablations | ||||||||||
| w/ news during RL | ✓ | ✓ | ||||||||
| w/o RL (SFT only) | ✓ | |||||||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Assumption | Value |
|---|---|
| Traded price | Midpoint of the best bid and best ask |
| Option fee | $0.65/contract/leg |
| Assignment fee | $5/event |
| Share commission | $0.005/share |
| Stock borrow | 0.5% annually |
| Stop-loss | 60% of entry premium |
| Profit | Risk-adjusted | Risk | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Code | Family | Trades | TR(%) | ASR | ACR | ASoR | AVOL(%) | MDD(%) | WR(%) |
| Directional | |||||||||
| ol | Outright | 4.29 | 1.43 | 5.08 | 2.93 | 6.91 | 2.33 | 46.3 | |
| dv | Debit vertical | 2.84 | 2.10 | 8.24 | 4.60 | 5.51 | 1.68 | 75.3 | |
| cv | Credit vertical | 1.46 | 1.97 | 14.88 | 7.28 | 2.54 | 0.70 | 82.4 | |
| dg | Defined risk reversal | 0.28 | 1.31 | 4.48 | 1.99 | 4.87 | 1.51 | 60.9 | |
| Parameter | Symbol | Value | Selection method |
|---|---|---|---|
| ARCH coefficient | 0.060 | QML on estimation sample | |
| GARCH coefficient | 0.762 | QML on estimation sample | |
| Persistence | 0.822 | implied (half-life 3.5 d) | |
| Winsorisation | — | fixed a priori | |
| Forecast horizon | 10 trading days | matches the 14-day tenor anchor | |
| Estimation sample | — | 3,710 obs, 2023-09-01 to 2025-02-28 | strictly pre-test |
| Family | Code | Feature | Entry | Edge | Friction |
| Participation gate, applied to every family | |||||
| all | — | — | — | ||
| Volatility: , fair value | |||||
| Long straddle | ls | 10% | 3.2% | ||
| Iron butterfly | ib | 20% | 9.1% | ||
| Iron condor | ic | 25% | 10.6% | ||
| Parameter | Qwen3.8-27B |
|---|---|
| Model and supervised fine-tuning | |
| Hugging Face revision | 1d4bf0f2 |
| Precision | bf16 |
| SFT episodes | |
| Maximum sequence length | |
| Train batch size | |
| Parameter | Symbol | Value |
| Universe and action space | ||
| Underlyings | 10 (nine single names, one index slot) | |
| Initial net asset value | $1,000,000 | |
| Cash accrual | 0 | |
| Registered families | 9 | |
| Tenor buckets (days to expiry) | 0–7, 8–30, 31–90, 91–180 | |
| Family | Default coordinates | Admitted range |
|---|---|---|
| Outright | long delta 0.55 | |
| Debit vertical | long delta 0.55, width 0.10 | long , width |
| Credit vertical | short delta 0.35, width 0.10 | short , width |
| Butterfly | centre 0.45, widths 0.10/0.10 | centre , widths |
| Defined-risk reversal | directional 0.35, tail wing 0.20 | directional , wing |
| Long straddle | call 0.50, put 0.50 |