Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions
Organizations: Department of Industrial Data Engineering, Hanyang University, Republic of Korea · School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea
Abstract
In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.
Figures & tables
| Condition | Evidence provided |
| C0-Evidence Control | Prediction, local SHAP, and current global attribution |
| C1-Raw | C0 + raw 6-month global attribution sequence |
| C2-Numerical | C1 + , slope, and trend direction |
| C3-Relational | C2 + deterministic OVERTAKES relations |
| C3-Analog | C3 + historical analog evidence |
| Model | RankIC | RankICIR | RMSE | MAE |
| Linear Regression | 0.0170 | 0.2380 | 0.1631 | 0.0845 |
| Ridge Regression | 0.0209 | 0.2810 | 0.1631 | 0.0848 |
| Random Forest | 0.0357 | 0.4900 | 0.1627 | 0.0841 |
| XGBoost | 0.0361 | 0.5380 | 0.1660 | 0.0859 |
| Method | |||||||||
| SHAP-Trajectory | 0.05146 | 0.05135 | 0.04845 | 0.04180 | 0.04260 | 0.04091 | 0.04012 | 0.03996 | 0.04121 |
| Market-State | 0.04752 | 0.04822 | 0.04899 | 0.04225 | 0.04076 | 0.04057 | 0.04098 | 0.03997 | 0.04147 |
| Recent-Return | 0.04544 | 0.04128 | 0.03972 | 0.04544 | 0.04128 | 0.03972 | 0.04544 | 0.04128 | 0.04042 |
| Random | 0.04902 | 0.04858 | 0.04874 | 0.04236 | 0.04205 | 0.04186 | 0.04047 | 0.04040 | 0.04114 |
| Model | Parse | EF C1 | EF C2 | EF C3 | Temp. C1 | Temp. C2 | Temp. C3 | Rel. C1 | Rel. C2 | Rel. C3 |
| Qwen3-8B | 0.9363 | 0.7219 | 0.7662 | 0.9195 | 0.6318 | 0.8784 | 0.8458 | 0.2155 | 0.3585 | 0.9295 |
| Qwen3-8B-Instruct | 1.0000 | 0.8259 | 0.9704 | 0.9686 | 0.3604 | 0.6125 | 0.5750 | 0.8625 | 0.8875 | 1.0000 |
| Qwen3-14B | 0.9779 | 0.7570 | 0.8061 | 0.9984 | 0.6979 | 0.9916 | 0.9916 | 0.3375 | 0.2059 | 0.9979 |
| Qwen3-14B-Instruct | 1.0000 | 0.7194 | 0.8844 | 0.9995 | 0.4854 | 0.9479 | 0.9521 | 0.4271 | 0.4979 | 1.0000 |
| Qwen3-32B-Instruct | 0.9988 | 0.6960 | 0.9032 | 0.9964 | 0.3396 | 0.9896 | 0.9313 | 0.5250 | 0.5729 | 1.0000 |
| Ministral-8B-Instruct | 0.5433 | 0.8733 | 0.8536 | 0.9204 | 0.5506 | 0.6273 | 0.5246 | 0.7341 | 0.8636 | 0.9836 |
| Condition | EF | Temporal Acc. | Relational Acc. | C0 Abstention |
| C0-Evidence Control | 0.9104 | – | – | 0.8386 |
| C1-Raw | 0.6960 | 0.3396 | 0.5250 | – |
| C2-Numerical | 0.9032 | 0.9896 | 0.5729 | – |
| C3-Relational | 0.9964 | 0.9313 | 1.0000 | – |
| C3-Analog | 0.9891 | 0.9000 | 1.0000 | – |
| Metric | C3-Relational | C3-Analog |
| Textual Temporal Transfer Accuracy | 0.9975 | 0.9975 |
| Deterministic Temporal Accuracy | 0.8625 | 0.7425 |
| Evidence Scope Adherence | ||
| Narrative Synthesis |
| Item | C3-Relational | C3-Analog | |
| Q1 Understandability | |||
| Q2 Decision Usefulness | |||
| Q3 Contextual Usefulness | |||
| Q4 Reliance Calibration | |||
| Q5 Adoption Intention | |||
| Overall |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Feature | Definition / Formula | Measurement & Source |
|---|---|---|
| Price Level | ||
| lowest_bid | Month CRSP: BIDLO | |
| highest_ask | Month CRSP: ASKHI | |
| close_price | Month CRSP: PRC | |
| ma3 | , minimum 2 observations | Trailing 3 months CRSP: PRC, CFACPR |
| ma12 | , minimum 6 observations | Trailing 12 months CRSP: PRC, CFACPR |