Applying LLMs to predictive tasks in finance is challenging due to look-ahead bias resulting from their training on long time-series data. This precludes the backtests typically employed in finance since retraining frontier models from scratch with a specific knowledge cutoff is prohibitive. In this paper, we introduce a fast, effective, and low-cost alternative. Our method guides generation at inference time by adjusting the logits of a large base model using a pair of smaller, specialized models -- one fine-tuned on information to be forgotten and another on information to be retained. We demonstrate that our method effectively removes both verbatim and semantic knowledge, corrects biases, and outperforms prior methods.
Figures & tables
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Initial LR
Best Verbatim
Best Q&A
Stupid Backoff Trigram
TopK=1
Alpha=10
princeton-nlp/Sheared-LLaMA-1.3B
5e-5
TopK=100
Alpha=0.8
princeton-nlp/Sheared-LLaMA-2.7B
4e-5
TopK=200
Alpha=1.0
Appendix
Table 1 : Configuration MUSE
Method
Epochs
Method-Specific Hyperparameters
GradDiff
1 ∗
α=1.0,γ=1.0
NPO
10
β=0.1,α=1.0,γ=1.0
SimNPO
10
δ=0,β=4.5,α=1.0,γ=0.125
Appendix
Table 2 : MUSE Configurations
Name
Score
google/gemma-3-27b-it
77.94
Unlearning Split B α =2
77.58
Unlearning Split B TopK=250
75.09
Appendix
Table 3 : MMLU CoT 0-Shot
Set A
Set B
Bucyrus International <> Caterpillar
Allegheny Energy <> FirstEnergy
El Paso <> Kinder Morgan
Goodrich <> United Technologies
Hillshire Brands <> Tyson Foods
Biomet <> Zimmer Holdings
Rockwood Holdings <> Albemarle
Family Dollar Stores <> Dollar Tree
Rock-Tenn <> MeadWestva
Bally Technologies <> Scientific Games
Bright House Networks LLC <> Charter Communications
Pharmacyclics <> AbbVie
Appendix
Table 4 : Final Deal Sets
Tell me about the acquisition of {target name} by {acquirer name}
When did {acquirer name} announce the acquisition of {target name}?
What firm bought {target name} in {year}?
Why did {acquirer name} acquire {target name}?
What strategic benefits did {acquirer name} gain from acquiring {target name}?
How did the {target name} acquisition help {acquirer name}’s business strategy?
What synergies were expected from the {acquirer name}-{target name} deal?
Appendix
Table 5 : Prompts used to distill data for the M&A (Temperature=0.4)
Ran for 2014-2016 and 2022-2024 for both Southwest Airlines and United Airlines.
Summarize the financial performance of {airline} in {year}
How well did {airline} do in {year}?
Summarize the operational performance of {airline} in {year}
What was the outlook for {airline} going into {year}
Was {year} a good year for {airline}?
Appendix
Table 6 : Prompts used to distill data for Airlines (Temperature=0, Gemma API)
Backtesting large language models (LLMs) on historical financial data is unreliable because pre-training cuts off after the events happened. An LLM trained in 2024 already "knows" which way 2018-2020 stocks moved. We name this failure parametric look-ahead bias and propose FinCAD, an inference-time adaptation of Context-Aware Decoding that suppresses an LLM's memory of historical outcomes without retraining. FinCAD pairs an adversarial bias-discovery pipeline that learns a model-specific memory-activating prior prompt with an entity- and date-adaptive rule that scales the CAD strength to per-(entity, date) memorisation, so the penalty fires on memorised in-sample dates and decays to zero out-of-sample. Across five 7-14B LLMs and five mega-cap equities, FinCAD cuts in-sample backtest returns by up to -67.1% on memorised dates while leaving 2025 out-of-sample returns within $8K and Sharpe within 0.10 of baseline, and preserves general-purpose reasoning within 1.7 pts. On an eleven-model leaderboard, it raises the in-sample / out-of-sample Spearman correlation from +0.779 to +0.846, recovering rankings that genuinely predict out-of-sample performance.
Large language models (LLMs) are increasingly deployed in financial contexts, raising critical concerns about reliability, alignment, and susceptibility to adversarial manipulation. While prior finance-related benchmarks assess LLMs' capabilities in stock trading, they are often restricted to small sample and fail to demonstrate LLM susceptibility to context with potential human bias. We introduce Fin-Bias (financial herding under long and uncertain financial context), a benchmark for evaluating LLM investment decision-making when faced with uncertainty and possible human-biased opinions. Fin-Bias includes 8868 long firm-specific analyst reports, including firm aspects summarized and analyzed by sophisticated analysts with investment ratings (Bullish/Neutral/Bearish) spanning from various industries. We present large language models with firm analyst reports with/without analyst investment ratings and even with 'fake' rating, to get investment ratings generated by LLMs. Our results reveal that LLMs tend to herd the explicit bias in context. We also develop a method to detect potential human opinions, which can encourage LLMs to think independently, some models even exceed human performance in predicting future stock return.
Xiaoyu Hu, Jinman Zhao
Rutgers University · Department of Computer Science, University of Toronto
Successful forecasting involves identifying patterns between historical and future states of the world which generalize to future observations. We apply LLMs to a variety of forecasting tasks and inspect their internal states using sparse autoencoders to understand whether they appear to rely on time-specific pieces of knowledge versus generalizable patterns. Our analyses identify features associated with both time-aware reasoning and look-ahead-biased reasoning. We then apply the LLMs to an entirely different domain and intervene on these features. We find that amplifying time-awareness features substantially reduces look-ahead bias on forecasting prompts while preserving general reasoning performance. In contrast, steering the candidate look-ahead-bias features does not produce an effect. These results suggest that interpretable temporal features can be used to causally shift LLMs toward more historically grounded reasoning.