cs.AIOct 8, 2026

MemTrial: Learning When to Trust Memory in LLM Portfolio Agents

Authors: Guanghao Wu, Zhuo Cai, Shoujin Wang

Organizations: University of Technology Sydney

Abstract

Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/NN) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/NN. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2% below 1/NN on PortBench and InvestorBench, against 15--38% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.

Figures & tables

Appendix figures & tables33 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

    May 27, 2026Taojie Zhu, Wentao Zhao, Rui Sun +7Benchmark ContaminationLLM Agent Evaluation

  2. FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

    Aug 4, 2026Ben Wang, Kang Zhou, Lifan Guo +2Financial ServicesLLM Agent Memory

  3. OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

    Aug 6, 2026Xinying Cai, Minghao Guo, Jiahe Liu +7LLM AuditingLLM Agent Evaluation