Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.
Figures & tables
Figure 1: Post-training brings a 4B model to performance comparable to frontier models on BETA. (a) Scores on the 240 scored test tasks improve from 20.94 to 37.94 to 43.31 through Base, SFT, and PPO. (b) Median calls across all 398 test tasks, including submission, change from 5 to 18 to 16. (c) AURA-4B ranks third among fifteen systems; the intermediate SFT checkpoint is also shown.
Evidence
Available observations
Analysis supported
Individual stock
OHLCV candles, quotes, technical indicators
Price trends, volume, volatility, and multiple resolutions
Table 1: Market information available in the sandbox. Full tool descriptions and task prompts appear in appendix B .
Property
Training
Test
Tasks
3,807
398
Distinct stocks
489
224
Decision dates
Feb. 2011–Dec. 2023
Feb. 2024–Mar. 2026
Scored evaluation tasks
—
240
Table 2: Benchmark split within a 499-equity market environment. The latest training label is realized on December 29, 2023; the earliest test decision is February 1, 2024.
System
Score S
Direction D
Magnitude M
Later score
Starting model
Qwen3-4B base
20.94
62.9
33.3
17.41
Frontier language models
glm-5.3
49.84
79.2
63.0
23.43
claude-fable-5
44.48
85.4
52.1
20.02
kimi-k3
36.28
80.4
45.1
21.84
Table 3: Existing models exhibit different strengths in direction and magnitude. Results for fourteen systems before adding AURA-4B. S is the joint forecast score, D is directional accuracy, and M is magnitude agreement conditional on a correct sign. Overall results use 240 tasks; the later-period column uses 71.
Checkpoint
Score
Relative gain vs. previous
Relative gain vs. Base
Base Qwen3-4B
20.94
—
—
SFT
37.94
+81.2%
+81.2%
AURA-4B (SFT + PPO)
43.31
+14.2%
+106.8%
Table 4: Training-stage results on the 240 scored test tasks. Gains are relative improvements.
Figure 2: AURA approaches GLM on easy and medium tasks and exceeds momentum on medium and hard tasks. (a) Scores for the three systems with recorded difficulty aggregates. (b) AURA scores higher than GLM on short horizons (1–10 days), while GLM scores higher on longer horizons (21–126 days). Sample sizes are shown below each group.
Figure 3: Post-training nearly doubles conditional magnitude agreement. The joint score factors as S=DM/100 . Direction changes from 62.9 to 65.4, while conditional magnitude changes from 33.3 to 66.2.
Recorded behavior
Base
SFT
AURA-4B
Median tool calls
5
18
16
Median distinct retrieval tools
4
7
6
Ranking and market/sector query share
16.2%
31.4%
38.0%
Median absolute forecast
0.0120
0.0312
0.0420
Positive forecasts on scored tasks
80.4%
47.9%
58.8%
Table 5: Behavior across Base, SFT, and AURA-4B (SFT + PPO) on the same test tasks. The first four rows use all 398 tasks; positive-forecast shares use the 240 scored tasks. Call counts include the terminal submission; distinct retrieval tools and query shares exclude it. Query shares pool calls across tasks and count get_ranking and get_market_kline .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Numerical information returned
get_kline
OHLCV candles at resolutions from one minute to one month.
Market or sector candles and market-breadth statistics.
get_indicator
A caller-parameterized indicator time series.
list_indicators
Indicator names and default parameters.
get_ranking
Cross-sectional metric rankings, including the target stock’s rank.
Appendix
Table 6: The ten available actions: nine retrieval tools and one terminal prediction action.
System
Earlier ( n=169 )
Later ( n=71 )
Change (%)
AURA-4B
42.17
46.04
+9.2
Momentum extrapolation
41.99
42.01
+0.0
Momentum-20
29.77
36.37
+22.2
GBM classifier
32.32
33.13
+2.5
gpt-5.6-sol
37.06
25.08
-32.3
glm-5.3
60.94
23.43
-61.6
Appendix
Table 7: Calendar-period results for every system. Earlier: February 2024–July 2025. Later: August 2025–March 2026. Relative changes are computed from the rounded endpoints.
Figure 4: Absolute forecast errors for the recorded DOW task. The realized log return is −0.0796 ; errors are multiplied by 100.
Financial markets are characterized by extreme non-stationarity, low signal-to-noise ratios, and strong dependence on external information such as news, company fundamentals, and macroeconomic signals. Yet, existing approaches either abstract time-series into text or decouple forecasting from language-based reasoning, leading to a fundamental mismatch between qualitative reasoning and quantitative outcomes. To address this, we introduce StockR1, a time-series-enhanced LLM that unifies stock forecasting and financial reasoning through a verifiable forecast action. Based on a tool-call design, the model first emits a forecast action, which is a structured and interpretable representation of its qualitative market outlook. It then invokes a time-series decoder conditioned on this action to generate distributional future trajectories, leading to more informed question answering and financial reasoning. We optimize the full pipeline with reinforcement learning, where rewards jointly reflect answer validity, forecast accuracy, and consistency between generated actions and observed time-series dynamics. In addition, rewards are reweighted by a sample-level uncertainty scalar, encouraging the model to accommodate varying uncertainty in market dynamics. We evaluate StockR1 on financial question answering and stock forecasting over a large-scale 10-year benchmark. Our method consistently outperforms time-series baselines and general-purpose LLMs, improving reasoning accuracy by 17.7% (4B) and 25.9% (8B). These findings demonstrate that structuring the forecast actions establishes a powerful synergy between language reasoning and temporal prediction, enabling LLMs to reason through verifiable, interpretable, and numerically grounded decisions.
Jialin Chen, Aosong Feng, Harshit Verma +7
Yale University · Arizona State University · University of Texas Rio Grande Valley
Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ultimately evaluated. We investigate whether a causal language model can instead represent forecasts and decisions directly through constrained token generation. FinATOM introduces a unified, head-free interface for three-step stock-return forecasting and dynamic five-ETF allocation. The forecasting model autoregressively emits volatility-standardized return tokens and is trained with ordinal and ranking supervision followed by a one-epoch token-level policy stage. The allocation model generates normalized long-only weights; supervised fine-tuning imitates a causal mean--variance anchor, and DAPO-augmented GRPO optimizes realized 21-day Sharpe subject to anchor consistency. In 2023--2025 ETF tests, the allocation policy improves pooled gross Sharpe from 1.428 to 1.529 and net Sharpe under a 5-bp transaction-cost model from 1.394 to 1.494. The multimodal allocation input attains the highest three-period mean Sharpe of 1.540, with its clearest advantage in 2025. On FinTexTS, the SFT and policy strategies achieve 73.52%/2.68 and 73.72%/2.69 cumulative-return/Sharpe, respectively. These results support the feasibility of direct language-model token generation for financial numerical prediction and decision-making, while motivating broader tests across assets, regimes, and random seeds.
Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.