Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions
Authors: Sujung Kim, Seung Hwan Cho, Sangjin Park, Young-Min Kim
Organizations: Department of Industrial Data Engineering, Hanyang University, Republic of Korea · School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea
In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.
Figures & tables
Figure 1: Overview of the proposed framework for generating evidence-grounded LLM narratives of stock return predictions.
Condition
Evidence provided
C0-Evidence Control
Prediction, local SHAP, and current global attribution
C1-Raw
C0 + raw 6-month global attribution sequence
C2-Numerical
C1 + Δ6P , slope, and trend direction
C3-Relational
C2 + deterministic OVERTAKES relations
C3-Analog
C3 + historical analog evidence
Table 1: Experimental Conditions and Evidence Provided
Model
RankIC ↑
RankICIR ↑
RMSE ↓
MAE ↓
Linear Regression
0.0170
0.2380
0.1631
0.0845
Ridge Regression
0.0209
0.2810
0.1631
0.0848
Random Forest
0.0357
0.4900
0.1627
0.0841
XGBoost
0.0361
0.5380
0.1660
0.0859
Table 2: Out-of-Sample Prediction Performance
Method
K=1
K=3
K=5
L=3
L=6
L=9
L=3
L=6
L=9
L=3
L=6
L=9
SHAP-Trajectory
0.05146
0.05135
0.04845
0.04180
0.04260
0.04091
0.04012
0.03996
0.04121
Market-State
0.04752
0.04822
0.04899
0.04225
0.04076
0.04057
0.04098
0.03997
0.04147
Recent-Return
0.04544
0.04128
0.03972
0.04544
0.04128
0.03972
0.04544
0.04128
0.04042
Random
0.04902
0.04858
0.04874
0.04236
0.04205
0.04186
0.04047
0.04040
0.04114
Table 3: Mean Absolute Error of Subsequent Universe Returns across Retrieval Settings
Model
Parse
EF C1
EF C2
EF C3
Temp. C1
Temp. C2
Temp. C3
Rel. C1
Rel. C2
Rel. C3
Qwen3-8B
0.9363
0.7219
0.7662
0.9195
0.6318
0.8784
0.8458
0.2155
0.3585
0.9295
Qwen3-8B-Instruct
1.0000
0.8259
0.9704
0.9686
0.3604
0.6125
0.5750
0.8625
0.8875
1.0000
Qwen3-14B
0.9779
0.7570
0.8061
0.9984
0.6979
0.9916
0.9916
0.3375
0.2059
0.9979
Qwen3-14B-Instruct
1.0000
0.7194
0.8844
0.9995
0.4854
0.9479
0.9521
0.4271
0.4979
1.0000
Qwen3-32B-Instruct
0.9988
0.6960
0.9032
0.9964
0.3396
0.9896
0.9313
0.5250
0.5729
1.0000
Ministral-8B-Instruct
0.5433
0.8733
0.8536
0.9204
0.5506
0.6273
0.5246
0.7341
0.8636
0.9836
Table 4: Cross-Model Consistency of Reasoning Externalization
Condition
EF ↑
Temporal Acc. ↑
Relational Acc. ↑
C0 Abstention
C0-Evidence Control
0.9104
–
–
0.8386
C1-Raw
0.6960
0.3396
0.5250
–
C2-Numerical
0.9032
0.9896
0.5729
–
C3-Relational
0.9964
0.9313
1.0000
–
C3-Analog
0.9891
0.9000
1.0000
–
Table 5: Structured Claim Performance by Reasoning Condition (Qwen3-32B-Instruct)
Financial markets are characterized by extreme non-stationarity, low signal-to-noise ratios, and strong dependence on external information such as news, company fundamentals, and macroeconomic signals. Yet, existing approaches either abstract time-series into text or decouple forecasting from language-based reasoning, leading to a fundamental mismatch between qualitative reasoning and quantitative outcomes. To address this, we introduce StockR1, a time-series-enhanced LLM that unifies stock forecasting and financial reasoning through a verifiable forecast action. Based on a tool-call design, the model first emits a forecast action, which is a structured and interpretable representation of its qualitative market outlook. It then invokes a time-series decoder conditioned on this action to generate distributional future trajectories, leading to more informed question answering and financial reasoning. We optimize the full pipeline with reinforcement learning, where rewards jointly reflect answer validity, forecast accuracy, and consistency between generated actions and observed time-series dynamics. In addition, rewards are reweighted by a sample-level uncertainty scalar, encouraging the model to accommodate varying uncertainty in market dynamics. We evaluate StockR1 on financial question answering and stock forecasting over a large-scale 10-year benchmark. Our method consistently outperforms time-series baselines and general-purpose LLMs, improving reasoning accuracy by 17.7% (4B) and 25.9% (8B). These findings demonstrate that structuring the forecast actions establishes a powerful synergy between language reasoning and temporal prediction, enabling LLMs to reason through verifiable, interpretable, and numerically grounded decisions.
Jialin Chen, Aosong Feng, Harshit Verma +7
Yale University · Arizona State University · University of Texas Rio Grande Valley
Large language models (LLMs) have the potential to aid and improve human decision-making in classification tasks, not only by providing fairly accurate predictions, but also in their ability to generate cogent narrative explanations of those predictions. Prior work has demonstrated that people generally find AI narrative explanations to be understandable, trustworthy, and convincing for changing beliefs and opinions; however, less is known about the impact of narrative explanations on objective human decision-making performance. Here we conduct a large-scale human behavioral experiment to evaluate decision-making performance with LLM-generated narrative explanations of varying persuasiveness. We found the degree of persuasiveness, or lack thereof, for LLM-based explanations did not meaningfully impact decision accuracy over a simple AI prediction alone, in agreement with typical results with explainable AI based on feature importance. We found evidence that narratives increased reliance on AI, but both when the AI prediction was correct and incorrect. Exploratory analyses also indicated that the more persuasive narratives may have had a detrimental effect on decision response times and the ability to discriminate between a correct and incorrect AI prediction. Overall, this work indicates that including narrative explanations with AI predictions may involve tradeoffs for decision-making performance, and more work is needed to determine how and when narrative explanations impact human decision-making.
Laura R. Marusich, Mary Grace Kozuch Dhooghe, Jonathan Z. Bakdash +1
DEVCOM Army Research Laboratory Adelphi, MD 20783 · University of Texas at Dallas Richardson, TX 75080 · Virginia Polytechnic Institute and State University Blacksburg, VA 24060
Can financial news reliably predict short-term stock movements? Despite advances in large language models, this question remains unresolved. We revisit this problem using a zero-shot natural language processing framework, investigating whether models can extract actionable signals from financial news without domain-specific training. We design a structured pipeline that combines zero-shot natural language inference with temporal aggregation, explicitly modelling recency and event-dependent impact horizons when integrating information across articles. To address the need for transparency in high-stakes settings, we introduce a multi-layered explainability framework that links predictions to token-level, article-level, and aggregate evidence, and produces grounded natural language rationales. Across multiple models and prediction horizons, we find that zero-shot approaches consistently fail to outperform simple baselines, with particularly weak performance on negative movements, suggesting deeper structural limitations in mapping news sentiment to short-term price dynamics. However, explainability signals reliably distinguish between trustworthy and unreliable predictions, offering practical value even when accuracy is limited. These findings highlight the limits of zero-shot financial NLP and motivate a shift toward decision-support systems that prioritise transparency and uncertainty awareness. Code: https://github.com/alimert05/zero-shot-stock-xai
Ali M Karaoglu, Shreyank N Gowda
School of Computer Science, University of Nottingham, Nottingham, United Kingdom