Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, information availability, and transition reliability are modeled. Existing LLM-based financial agents incorporate historical evidence, yet they provide limited support for preserving issuer-specific chronology under point-in-time constraints and for identifying when historical transitions contribute information beyond the current forecast. We present RICE-Alpha (Reliability-Informed Correction with Event Graphs), a point-in-time stock-scoring framework that separates a history-aware multi-view Base Alpha from a reliability-calibrated residual correction derived from historical event continuation. A Multi-Tier Memory Layer grounds news interpretation in temporally eligible issuer-specific history, while a Typed Event Agent constructs event states whose successor relations are formed within issuers and pooled across firms only after valid local pairing. Matured transitions are calibrated by their empirical reliability, and the resulting graph signal is residualized against the Base Alpha and technical view to obtain the RICE Delta. On daily Nasdaq-100 and Hang Seng Index panels from 2024 to 2026, RICE-Alpha achieves the strongest results among the evaluated LLM-based agents and momentum across four predictive and four portfolio-level metrics. Its ICIR more than doubles that of the strongest baseline, while net Sharpe ratios reach 1.656 and 1.725 in the U.S. and Hong Kong, respectively. U.S. ablations further show significant reductions in IC and RankIC after Holm adjustment when major components are removed. These results indicate that historical event continuation adds incremental information when it is temporally grounded, reliability-calibrated, and introduced as a residual correction to a multi-view forecast.
Figures & tables
Figure 1: Overview of the RICE-Alpha framework. A point-in-time multi-view Base Alpha is refined by a reliability-calibrated residual derived from historical event propagation.
U.S. (Nasdaq-100)
Hong Kong (HSI)
Panel A: Prediction quality
Method
IC ↑
ICIR ↑
RankIC ↑
RankICIR ↑
IC ↑
ICIR ↑
RankIC ↑
RankICIR ↑
MEME
0.0216
0.1135
0.0254
0.1228
0.0013
0.0067
0.0156
0.0794
R&D-Agent-Quant
0.0188
0.0796
0.0177
0.0706
0.0011
0.0048
0.0125
0.0630
AI Hedge Fund
0.0077
0.0595
0.0094
0.0845
−0.0018
−0.0121
0.0033
0.0264
Momentum (12–1)
0.0316
0.1274
0.0322
0.1324
0.0160
0.0645
0.0327
0.1312
Table 1: Main results on common data and dated universes. ICIR and RankICIR are not annualized; ARR and MDD (a positive loss) are in %. Bold: market best. All strategies use the long-only, equal-weighted top-10% weekly rule and costs of Section 4.1 ; NDX/HSI exclude costs. † / ‡ : lag-5 Newey–West t>1.96 / 2.58 .
Figure 4: Net asset value of Rice-Alpha , the baselines, and the index in (a) the U.S. and (b) Hong Kong from 2024-01-03; models are net of costs, indices exclude costs.
Configuration
IC
ICIR
RankIC
RankICIR
Rice-Alpha (complete)
0.0350
0.2928
0.0360
0.2927
w/o C3 (direct)
0.0260
0.2203
0.0290
0.2437
Reliability only
0.0326
0.2706
0.0330
0.2665
Residualization only
0.0208
0.1777
0.0242
0.2040
w/o C2–C3 (memory)
0.0191
0.1469
0.0205
0.1531
w/o C1–C2–C3
0.0064
0.0472
0.0090
0.0641
Table 2: U.S. ablations and two separate references. C1 is retrieved memory, C2 numerical graph evidence, and C3 reliability weighting plus residualization. All component rows use multi-agent stock scoring; the single-agent row reports that run’s stored final score. Component paired tests are in Table 3 ; the single-agent test is in the text.
Panel A: Rice-Alpha means
Market
Dates
IC ( t )
RankIC ( t )
U.S.
562
0.0350 (4.31)
0.0360 (4.24)
Hong Kong
551
0.0283 (2.48)
0.0395 (3.10)
Pooled
534
0.0317 (4.17)
0.0377 (4.41)
Panel B: paired increments of Rice-Alpha over each comparator
Market
Comparator
Metric
Increment
tdiff
pboot ( pHolm )
Table 3: Daily signal significance and mechanism contrasts. Panel A gives mean IC and RankIC with Newey–West t -statistics (lag 5; 20-lag IC values 4.72 and 2.28). Pooled means weight the two markets equally on 534 shared valid dates. Panel B compares the complete system with six U.S. rows or Hong Kong AI Hedge Fund on common valid dates; p -values use one-sided 20-session paired block bootstrap (5,000 replications), with Holm adjustment by market and metric. Panel C reports the separate unadjusted comparison of reliability-only against residualization-only; it is outside the Panel B Holm family.
U.S. (Nasdaq-100)
Hong Kong (HSI)
Method
Capture
Crash
Down mo.
Capture
Crash
Down mo.
Index (NDX / HSI)
1.00/1.00
−24.4
−4.5
1.00/1.00
−21.2
−3.3
MEME
0.28/0.33
−20.8
−4.1
0.02/0.11
+3.5
+0.2
R&D-Agent-Quant
0.29/0.35
−24.9
−4.9
0.25/0.33
−11.8
−1.8
AI Hedge Fund
0.73/0.71
−17.0
−3.2
0.93/0.97
−20.1
−2.8
Momentum (12–1)
0.83/0.85
−16.8
−4.1
0.58/0.64
−9.3
−1.4
Table 4: Behavior in falling and rising markets. Capture: mean return on index-down/up sessions relative to the index. Crash: return (%) in the index’s largest drawdown. Down mo.: mean return (%) in the 7 and 13 months the index fell. Bold: best Crash and Down mo.
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.
Liqin Ye, Haorui Wang, Fardin Ahmed +8
Georgia Institute of Technology · Amazon · University of Florida
Financial retrieval-augmented generation (RAG) systems typically rank evidence by textual relevance, but in financial markets evidence utility depends on event type, forecast horizon, and market context. We study news-triggered event-impact prediction as a point-in-time financial RAG problem. For each company-news anchor, the system retrieves financial news and SEC filing passages, appends a pre-decision market-context card, and predicts multi-horizon residual-return signals. Our method keeps the LLM frozen and adapts retrieval through an external Bayesian source memory updated from matured residual-return feedback. On a fixed 89-stock Nasdaq-oriented universe derived from the FinRL-DeepSeek/FNSPID task, using original FNSPID news and point-in-time EDGAR filing passages, Frozen Reader with Source Memory improves held-out macro-F1 from 0.438 to 0.471 and downstream portfolio Sharpe from 0.52 to 0.84 relative to Frozen Reader with No Memory. Supervised LoRA gives modest gains under static retrieval, but after source-memory adaptation, the LoRA reader does not improve over the frozen reader. These results suggest that, for financial RAG systems, learning where to retrieve can be as important as learning how to read, offering a modular route to market-feedback adaptation.
Zijie Zhao, Roy E. Welsch
Massachusetts Institute of Technology · Cambridge, MA, USA
Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, several living benchmarks have been proposed. However, existing benchmarks either lack the multidimensional events essential for accurate forecasting due to data scarcity, or focus on relatively closed environments. To assess the predictive capabilities of LLMs in complex, real-world scenarios, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including future event probabilities, trend and time series forecasting. LEAF utilizes a recursive retrieval agent system paired with dual-agent cross-validation to provide comprehensive and relevant auxiliary text for forecasting. Evaluating state-of-the-art proprietary and open-weight LLMs, we find that these models can leverage signals extracted from complex events to enhance predictive performance. In the stock domain, we find that LLMs achieve better performance on equities they confidently identify as more predictable. Furthermore, the events demonstrate a strong correlation with the target equities. To this end, LEAF provides a necessary, dynamically updating testbed to continuously track and drive progress in event-driven forecasting tasks.