MemTrial: Learning When to Trust Memory in LLM Portfolio Agents
Organizations: University of Technology Sydney
Abstract
Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2% below 1/ on PortBench and InvestorBench, against 15--38% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.
Figures & tables
| Spearman correlation ( ) | ||||
| Benchmark | Outcome credit vs. 1/ | Contribution vs. 1/ | Outcome credit vs. contribution | Dates with a detectable effect |
| PortBench-Full | 0.00 | 0.03 | 0.06 | 3–4 / 56 |
| PortBench-Raw | 0.00 | 0.04 | 0.03 | 2–4 / 56 |
| InvestorBench | 0.00 | 0.02 | 0.02 | 126–127 / 149 |
| ClassAlloc | 0.00 | 0.07 | 0.07 | 53–57 / 71 |
| PortBench | InvestorBench | ClassAlloc | PlantedMem | |
| Assets | 99 (6 classes) | 4 stocks, cash | 5 classes | 4 classes |
| Decisions | monthly | daily | monthly | monthly |
| Development | 2019–22 | none | none | 15 worlds |
| Test dates | 20 (2023–24) | 149 (10/20–05/21) | 71 (05/20–03/26) | 120 / world |
| Experiences | 338 decisions | 63 lessons | 58 lessons | 12/48 planted |
| Seeds | 3 | 3 | 3 | 100 worlds |
| Method | PortBench-Full | PortBench-Raw | InvestorBench | ClassAlloc | PlantedMem | |
| %/month | %/month | bp/day | %/month | %/month | ||
| Rule-based | 1/ ( DeMiguel et al., 2009 ) | 2.02 0.00 | 2.02 0.00 | 10.45 0.00 | 0.623 0.000 | 0.57 0.00 |
| Minimum variance ( Markowitz, 1955 ) | 0.25 0.00 | 0.25 0.00 | 5.53 0.00 | 0.230 0.000 | 0.04 0.00 | |
| LLM w/o experience learning | Memory-free ( Brown et al., 2020 ) | 1.60 0.05 | 1.07 0.02 | 8.43 0.29 | 0.657 0.001 | 0.76 0.08 |
| Self-consistency ( Wang et al., 2022 ) | 1.56 0.11 | 1.17 0.03 | 7.89 0.21 | 0.686 0.007 | 0.79 0.03 | |
| Similarity retrieval (top-2) ( Lewis et al., 2020 ) | 1.43 0.19 | 1.24 0.29 | 8.67 0.19 | 0.649 0.033 | 0.72 0.07 |
| PortBench-Full | InvestorBench | ClassAlloc | ||||||||
| Method | CR % | SR | MDD % | CR % | SR | MDD % | CR % | SR | MDD % | |
| Rule-based | 1/ ( DeMiguel et al., 2009 ) | 61.2 0.0 | 1.77 0.00 | 5.6 0.0 | 19.1 0.0 | 2.87 0.00 | 5.7 0.0 | 82.0 0.0 | 1.22 0.00 | 12.0 0.0 |
| Minimum variance ( Markowitz, 1955 ) | 5.4 0.0 | 1.19 0.00 | 1.5 0.0 | 9.9 0.0 | 1.47 0.00 | 6.4 0.0 | 19.1 0.0 | 4.19 0.00 | 0.2 0.0 | |
| LLM w/o exp. learning | Memory-free ( Brown et al., 2020 ) | 48.5 0.8 | 1.47 0.06 | 7.3 1.0 | 14.6 0.5 | 2.24 0.08 | 6.3 0.1 | 88.3 1.5 | 1.11 0.00 | 13.9 0.3 |
| Experience- learning agents | FinMem ( Yu et al., 2025 ) | 49.6 9.2 | 1.53 0.21 | 6.2 1.4 | 14.7 0.4 | 2.27 0.03 | 6.1 0.1 | 75.4 2.6 | 1.07 0.03 | 15.4 0.5 |
| MemRL ( Zhang et al., 2026 ) | 47.7 11.1 | 1.45 0.19 | 6.9 1.1 | 15.4 0.7 | 2.33 0.09 | 6.1 0.3 | 84.0 1.2 | 1.06 0.03 | 15.4 1.1 | |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
| , | decision date, number of dates |
| market state at date | |
| number of assets (plus cash) | |
| investor profile: maximum share of risky assets, minimum share of defensive assets, risk aversion | |
| projection of a portfolio onto the investor’s mandate | |
| ; , | investor utility (certainty-equivalent return) of portfolio , Eq. ( 1 ); net return and realized variance over the holding period |
| PortBench (Full / Raw) | InvestorBench | ClassAlloc | PlantedMem | |
| Decision dates | 59 monthly, Aug 2019 – Dec 2024 (a month whose 20-day window would overlap the previous one is skipped); 56 usable per configuration | 63 warm-up and 149 test trading days, Jul 2020 – May 2021 | 129 monthly, Jul 2015 – Mar 2026: 58 warm-up and 71 test | 132 monthly, Apr 2015 – Mar 2026, per episode; the first 12 are a warm-up |
| Development / test | 36 / 20 months | none (fully held out) / 149 days | none (fully held out) / 71 months | seeds 0–14 / 15–114 |
| Assets | 99 in 6 classes, including cash | HON, JNJ, MSFT, UVV and cash | equities, bonds, commodities, real estate and cash: equal-weighted indices of 89 assets | equities, bonds, cash, commodities |
| Experiences | 338 records of earlier decisions (Apr 2015 – Jul 2019); 29 retrieved | 63 lessons written in the warm-up | 58 lessons written in the warm-up | 12 or 48 planted per episode |
| Drafts per date and seed | 19: the 16 subsets of the 4 retrieved experiences, FinMem, MemRL and the benchmark’s native memory; Reflexion runs separately | 20: the 16 subsets, ExpeL, FinMem, MemRL and Reflexion | 20: the 16 subsets, ExpeL, FinMem, MemRL and Reflexion | resampled from 2,828 logged LLM allocations |
| LLM executions | 6,612 (6,518 valid) and 354 Reflexion decisions | 9,003 | 4,318 (58 in the warm-up), and 210 redrawn (Appendix D.3 ) | none |
| Benchmark | Method | Utility | Gross return | Trading cost | Risk penalty | L1 |
| PortBench-Full | 1/ | 2.02 0.00 | 2.25 0.00 | 0.03 0.00 | 0.21 0.00 | 0.17 0.00 |
| Memory-free | 1.60 0.05 | 2.12 0.05 | 0.17 0.02 | 0.34 0.03 | 1.13 0.11 | |
| Sim. retrieval (top-4) | 1.70 0.06 | 2.19 0.04 | 0.18 0.02 | 0.31 0.02 | 1.17 0.14 | |
| Draft averaging | 1.62 0.05 | 2.09 0.06 | 0.16 0.00 | 0.32 0.02 | 1.02 0.01 | |
| PortBench-Raw | 1/ | 2.02 0.00 | 2.25 0.00 | 0.03 0.00 | 0.21 0.00 | 0.17 0.00 |
| Memory-free | 1.07 0.02 | 1.43 0.03 | 0.18 0.01 | 0.17 0.00 | 1.19 0.08 |
| Variant | PortBench | InvestorBench | ClassAlloc | PlantedMem (core) | PlantedMem (informative) |
| Ablation (Figure 4 ) | |||||
| MemTrial (ours) | 0.3 0.5 | 0.3 0.5 | 14.6 21.7 | 33.7 9.2 | 44.9 22.5 |
| per-experience only (no content) | 0.0 0.0 | 0.3 0.5 | 14.6 21.7 | 31.8 8.9 | 22.0 20.7 |
| content-aware only (no learner selection) | 1.1 1.3 | 0.3 0.5 | 14.6 21.7 | 32.6 9.0 | 44.5 23.1 |
| in-sample F-test | 0.0 0.0 | 0.0 0.0 | 2.3 4.1 | 29.3 8.4 | 15.1 17.8 |
| no trust test | 24.7 21.4 | 34.5 18.8 | 53.8 24.7 | 54.3 12.6 | 64.9 18.6 |
| PlantedMem | |||||||||
| Step changed | Variant | PortBench-Full | PortBench-Raw | InvestorBench | ClassAlloc | no signal | signal | informative content | core |
| – | MemTrial (ours) | 1.98 0.00 | 2.01 0.00 | 10.40 0.04 | 0.613 0.017 | 0.784 0.031 | 0.932 0.070 | 0.862 0.068 | 0.873 0.045 |
| Value estimation | per-experience learner only | 1.98 0.00 | 2.01 0.00 | 10.40 0.04 | 0.613 0.017 | 0.786 0.030 | 0.929 0.069 | 0.817 0.048 | 0.872 0.045 |
| content-aware learner only | 1.98 0.00 | 1.94 0.12 | 10.40 0.04 | 0.613 0.017 | 0.785 0.031 | 0.929 0.072 | 0.866 0.072 | 0.871 0.046 | |
| Trust test | in-sample F-test | 1.98 0.00 | 2.01 0.00 | 10.37 0.00 | 0.628 0.001 | 0.786 0.029 | 0.924 0.064 | 0.814 0.045 | 0.869 0.043 |
| none (always act on values) | 1.89 0.11 | 1.77 0.41 | 9.91 0.45 | 0.612 0.005 | 0.767 0.038 | 0.940 0.080 | 0.882 0.068 | 0.871 0.052 | |
| PlantedMem | |||||||
| Method | PortBench-Full | PortBench-Raw | InvestorBench | ClassAlloc | no signal | signal | informative content |
| 1/ | 2.021 0.000 | 2.021 0.000 | 10.45 0.00 | 0.623 0.000 | 0.575 0.000 | 0.575 0.000 | 0.575 0.000 |
| MemTrial (ours) | 1.976 0.005 | 2.007 0.002 | 10.40 0.04 | 0.613 0.017 | 0.784 0.031 | 0.932 0.070 | 0.862 0.068 |
| Hedge, prior 0.9 on the reference | 1.992 0.009 | 2.003 0.003 | 10.41 0.01 | 0.627 0.001 | 0.786 0.029 | 0.787 0.029 | 0.787 0.030 |
| With the anchored action of MemTrial | |||||||
| Memory-free | 1.985 0.008 | 1.998 0.004 | 10.34 0.02 | 0.630 0.001 | 0.789 0.030 | 0.789 0.032 | 0.788 0.030 |
| Method | PortBench | InvestorBench | ClassAlloc | ||||||
| End | Falling | Rising | End | Falling | Rising | End | Falling | Rising | |
| Memory-free | 0.46 | 0.58 | 0.61 | 0.36 | 0.06 | 0.39 | 0.97 | 0.25 | 0.89 |
| FinMem | 1.93 | 1.48 | 0.96 | 0.20 | 0.03 | 0.22 | 0.25 | 0.10 | 0.35 |
| MemRL | 2.82 | 0.38 | 2.64 | 0.10 | 0.11 | 0.11 | 0.22 | 0.20 | 0.07 |
| Reflexion | 1.16 | 1.31 | 1.78 | 0.10 | 0.06 | 0.11 | 0.87 | 0.67 | 0.71 |
| ExpeL | 0.88 | 1.03 | 0.83 | 0.16 | 0.03 | 0.18 | 0.39 | 0.26 | 0.33 |
| Method | Shortfall | Shortfall | Shortfall | Median largest shortfall (%) |
| Memory-free | 27.8 3.7 | 77.8 3.7 | 94.4 0.0 | 3.15 0.19 |
| FinMem | 19.1 3.9 | 67.9 3.9 | 90.7 3.2 | 3.59 0.12 |
| MemRL | 31.5 4.9 | 74.7 1.1 | 93.2 1.1 | 3.04 0.26 |
| Reflexion | 14.2 4.3 | 55.6 1.9 | 88.9 3.7 | 4.60 0.24 |
| ExpeL | 26.5 2.1 | 69.8 3.9 | 83.3 3.2 | 3.01 0.36 |
| MemTrial (ours) | 93.2 1.1 | 100.0 0.0 | 100.0 0.0 | 0.20 0.01 |
| Method | PortBench | InvestorBench | ClassAlloc | ||||||
| CR | SR | MDD | CR | SR | MDD | CR | SR | MDD | |
| 1/ | 53.1 0.0 | 1.72 0.00 | 5.7 0.0 | 18.1 0.0 | 2.87 0.00 | 5.4 0.0 | 82.0 0.0 | 1.22 0.00 | 12.0 0.0 |
| Memory-free | 36.0 0.8 | 1.43 0.04 | 6.8 0.8 | 14.4 0.4 | 2.26 0.06 | 6.0 0.1 | 82.5 1.8 | 1.14 0.02 | 12.7 0.3 |
| FinMem | 39.1 3.0 | 1.53 0.08 | 6.2 0.6 | 15.1 0.2 | 2.42 0.03 | 5.8 0.0 | 73.5 0.5 | 1.12 0.00 | 13.4 0.0 |
| MemRL | 35.4 4.4 | 1.39 0.05 | 6.5 0.4 | 15.2 0.1 | 2.39 0.03 | 5.9 0.1 | 77.4 0.4 | 1.11 0.00 | 13.5 0.3 |
| Reflexion | 35.4 1.8 | 1.43 0.09 | 6.8 0.8 | 12.6 0.1 | 2.02 0.01 | 5.9 0.0 | 77.4 1.6 | 1.14 0.01 | 13.0 0.4 |
| Hyperparameter | Value | PortBench-Full | PortBench-Raw | InvestorBench | PlantedMem (core) | PlantedMem (informative) |
| Test level | 0.01 | 1.98 0.00 | 2.01 0.00 | 10.37 0.00 | 0.87 0.04 | 0.84 0.07 |
| 0.02 | 1.98 0.00 | 2.01 0.00 | 10.37 0.00 | 0.87 0.04 | 0.85 0.07 | |
| 0.05 (default) | 1.98 0.00 | 2.01 0.00 | 10.40 0.04 | 0.87 0.04 | 0.86 0.07 | |
| 0.1 | 2.01 0.05 | 1.98 0.04 | 10.42 0.09 | 0.87 0.05 | 0.87 0.07 | |
| 0.2 | 1.94 0.05 | 1.93 0.13 | 10.39 0.07 | 0.87 0.05 | 0.88 0.07 | |
| Minimum forward scores | 5 | 1.98 0.00 | 2.01 0.00 | 10.40 0.04 | 0.87 0.05 | 0.86 0.07 |
| Design | Drafts | PortBench-Full | PortBench-Raw | InvestorBench | PlantedMem (core) | PlantedMem (informative) | |
| 2 | full factorial | 4 | 1.98 0.01 | 2.01 0.00 | 10.27 0.11 | 0.812 0.043 | 0.798 0.042 |
| 3 | half fraction | 4 | 1.97 0.02 | 2.01 0.00 | 10.14 0.21 | 0.829 0.041 | 0.799 0.047 |
| 3 | full factorial | 8 | 1.98 0.01 | 2.01 0.00 | 9.96 0.36 | 0.865 0.051 | 0.850 0.070 |
| 4 | half fraction (default) | 8 | 1.98 0.00 | 2.01 0.00 | 10.40 0.04 | 0.868 0.045 | 0.862 0.072 |
| 4 | full factorial | 16 | 1.97 0.00 | 1.95 0.11 | 9.99 0.34 | 0.889 0.047 | 0.903 0.085 |
| 5 | half fraction | 16 | – | – | – | 0.896 0.044 | 0.914 0.078 |
| PortBench- Full | PortBench- Raw | InvestorBench | PlantedMem (core) | ||
| 0.5 | 1 | 1.80 0.02 | 2.02 0.00 | 9.99 0.03 | 0.87 0.04 |
| 0.5 | 4 | 1.83 0.02 | 1.92 0.01 | 9.71 0.02 | 0.87 0.04 |
| 0.5 | 16 | 1.85 0.02 | 1.72 0.01 | 9.65 0.04 | 0.86 0.04 |
| 0.75 | 1 | 1.86 0.03 | 2.02 0.00 | 10.29 0.05 | 0.87 0.04 |
| 0.75 | 4 | 1.92 0.01 | 1.98 0.00 | 10.16 0.02 | 0.87 0.04 |
| 0.75 | 16 | 1.94 0.01 | 1.89 0.00 | 10.12 0.01 | 0.87 0.04 |
| Design | Drafts | No influence | Noise only | Many exp. | Informative | Core avg. | |||
| Reference | 8 | 0.782 0.031 | 0.782 0.031 | 0.782 0.031 | 0.782 0.031 | 0.782 0.031 | 0.782 0.031 | 0.782 0.031 | 0.782 0.031 |
| , full factorial | 4 | 0.784 0.029 | 0.782 0.038 | 0.785 0.029 | 0.804 0.048 | 0.904 0.130 | 0.782 0.029 | 0.798 0.042 | 0.812 0.043 |
| , half fraction | 4 | 0.784 0.032 | 0.784 0.042 | 0.787 0.038 | 0.821 0.061 | 0.970 0.123 | 0.789 0.036 | 0.799 0.047 | 0.829 0.041 |
| , full factorial | 8 | 0.782 0.028 | 0.778 0.032 | 0.784 0.031 | 0.904 0.089 | 1.078 0.154 | 0.812 0.048 | 0.850 0.070 | 0.865 0.051 |
| , half fraction (default) | 8 | 0.783 0.028 | 0.777 0.037 | 0.794 0.045 | 0.901 0.084 | 1.085 0.114 | 0.836 0.079 | 0.862 0.072 | 0.868 0.045 |
| , full factorial | 16 | 0.780 0.030 | 0.778 0.029 | 0.792 0.049 | 0.950 0.080 | 1.143 0.134 | 0.869 0.094 | 0.903 0.085 | 0.889 0.047 |
| Trust rule | No influence | Noise only | Many exp. | Informative | |||
| MemTrial (trust test) | 2.3 6.6 | 3.0 10.1 | 13.7 20.7 | 64.5 23.3 | 84.8 8.9 | 26.1 22.6 | 44.9 22.5 |
| Per-experience learner only | 1.5 5.1 | 2.0 8.3 | 10.5 18.0 | 61.4 24.0 | 83.5 10.1 | 22.0 20.7 | 22.0 20.7 |
| In-sample F-test | 2.4 6.7 | 2.1 7.1 | 8.1 13.7 | 52.4 22.9 | 81.7 10.5 | 14.8 17.7 | 15.1 17.8 |
| Method | No influence | Noise only | Many exp. | Informative | |||
| Reference (memory-free average) | 0.788 0.032 | 0.788 0.032 | 0.788 0.032 | 0.788 0.032 | 0.788 0.032 | 0.788 0.032 | 0.788 0.032 |
| Outcome credit (MemRL†) | 0.759 0.097 | 0.658 0.089 | 0.758 0.090 | 0.755 0.098 | 0.749 0.144 | 0.760 0.085 | 0.760 0.085 |
| Counterfactual selection | 0.768 0.097 | 0.659 0.087 | 0.805 0.093 | 0.926 0.100 | 1.101 0.138 | 0.852 0.078 | 0.852 0.078 |
| MemTrial w/o trust test | 0.776 0.047 | 0.758 0.056 | 0.803 0.059 | 0.924 0.092 | 1.094 0.131 | 0.843 0.061 | 0.882 0.068 |
| MemTrial, never acting on its values | 0.788 0.028 | 0.787 0.029 | 0.788 0.028 | 0.788 0.029 | 0.788 0.029 | 0.789 0.029 | 0.789 0.029 |
| MemTrial (ours) | 0.787 0.030 | 0.782 0.035 | 0.793 0.043 | 0.912 0.088 | 1.090 0.124 | 0.821 0.050 | 0.862 0.068 |
| Method | No influence | Noise only | Many exp. | Informative | Core avg. | |||
| 1/ | 0.575 0.000 | 0.575 0.000 | 0.575 0.000 | 0.575 0.000 | 0.575 0.000 | 0.575 0.000 | 0.575 0.000 | 0.57 0.00 |
| Draft averaging | 0.782 0.032 | 0.715 0.052 | 0.777 0.031 | 0.764 0.044 | 0.719 0.080 | 0.775 0.041 | 0.775 0.041 | 0.75 0.03 |
| MemTrial, always 1/ when not trusted ( ) | 0.578 0.016 | 0.580 0.021 | 0.596 0.055 | 0.783 0.124 | 1.030 0.136 | 0.626 0.073 | 0.700 0.096 | 0.71 0.05 |
| MemTrial, never acting on its values | 0.633 0.009 | 0.608 0.017 | 0.630 0.009 | 0.624 0.014 | 0.611 0.026 | 0.626 0.013 | 0.626 0.013 | 0.62 0.01 |
| MemTrial, anchored ( ) | 0.635 0.015 | 0.612 0.024 | 0.646 0.047 | 0.805 0.114 | 1.033 0.133 | 0.669 0.063 | 0.733 0.085 | 0.75 0.05 |
| MemTrial, uniform prior ( ) | 0.735 0.022 | 0.675 0.041 | 0.736 0.039 | 0.852 0.100 | 1.043 0.130 | 0.751 0.055 | 0.797 0.073 | 0.81 0.05 |
| Method | All 71 months | 50 months before cutoff | 21 months after cutoff |
| 1/ | 0.623 0.000 | 0.554 0.000 | 0.787 0.000 |
| Minimum variance | 0.230 0.000 | 0.178 0.000 | 0.353 0.000 |
| Memory-free | 0.657 0.001 | 0.636 0.008 | 0.705 0.017 |
| Self-consistency | 0.686 0.007 | 0.672 0.010 | 0.721 0.013 |
| Similarity retr. (top-2) | 0.649 0.033 | 0.616 0.043 | 0.728 0.008 |
| Similarity retr. (top-4) | 0.643 0.039 | 0.621 0.047 | 0.696 0.027 |
| LLM, temperature | Memory-free | Best experience- learning agent | Best other LLM-based method | MemTrial (ours) | MemTrial w/o trust test | Improv. (%, ) | Trust rate (%) |
| (a) InvestorBench (bp per day) | |||||||
| gpt-4.1-mini, 0 | 7.74 0.35 | ExpeL 9.46 0.09 | ExpeL 9.46 0.09 | 9.84 0.51 | 9.69 0.22 | 6.0 (0.32) | 23.2 7.9 |
| gpt-4.1-mini, 0.3 | 7.63 0.25 | ExpeL 9.40 0.15 | ExpeL 9.40 0.15 | 10.54 0.32 | 9.33 0.28 | 4.0 (0.09) | 7.0 12.1 |
| gpt-4.1-mini, 0.7 (main) | 8.43 0.29 | ExpeL 8.93 0.32 | Hedge 8.97 0.23 | 10.40 0.04 | 9.91 0.45 | 3.2 (0.12) | 0.3 0.5 |
| gpt-4.1-mini, 1.0 | 7.83 0.47 | ExpeL 9.51 0.45 | ExpeL 9.51 0.45 | 10.36 0.01 | 10.10 0.54 | 4.8 (0.15) | 0.0 0.0 |
| gpt-4.1-nano, 0.7 | 8.45 0.74 | MemRL 7.75 0.55 | Self-consistency 9.07 0.39 | 10.30 0.06 * | 9.22 1.25 | 7.4 (0.02) | 2.1 3.6 |
| gpt-4.1-mini, temperature | ||||
| Method | 0 | 0.3 | 0.7 (main) | 1.0 |
| 1/ | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 |
| Minimum variance | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 |
| Memory-free | 7.74 0.35 | 7.63 0.25 | 8.43 0.29 | 7.83 0.47 |
| Self-consistency | 7.86 0.17 | 7.81 0.08 | 7.89 0.21 | 7.83 0.24 |
| Similarity retrieval (top-2) | 8.36 0.14 | 8.73 0.28 | 8.67 0.19 | 8.60 0.45 |
| Method | gpt-4.1-mini | gpt-4.1-nano | gpt-5-mini | Llama-3.3-70B | Qwen3-235B | DeepSeek-V3.1 | Gemini 2.5 Flash | Claude Haiku 4.5 |
| 1/ | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 | 10.45 0.00 |
| Minimum variance | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 | 5.53 0.00 |
| Memory-free | 8.43 0.29 | 8.45 0.74 | 8.36 0.54 | 8.34 0.39 | 8.31 0.21 | 8.37 0.55 | 7.66 0.23 | 8.42 0.10 |
| Self-consistency | 7.89 0.21 | 9.07 0.39 | 8.31 0.31 | 8.26 0.22 | 8.37 0.17 | 8.63 0.51 | 7.92 0.23 | 8.48 0.07 |
| Similarity retrieval (top-2) | 8.67 0.19 | 6.89 0.92 | 8.41 0.17 | 8.90 0.28 | 8.15 0.44 | 8.66 0.47 | 8.50 0.41 | 8.98 0.12 |
| Similarity retrieval (top-4) | 8.12 0.61 | 7.71 0.66 | 8.46 0.24 | 8.51 0.14 | 8.00 0.32 | 8.55 0.45 | 8.22 0.08 | 8.83 0.15 |
| Method | gpt-4.1-mini | gpt-4.1-nano | gpt-5-mini | Llama-3.3-70B | Qwen3-235B | DeepSeek-V3.1 | Gemini 2.5 Flash | Claude Haiku 4.5 |
| 1/ | 0.623 0.000 | 0.623 0.000 | 0.623 0.000 | 0.623 0.000 | 0.623 0.000 | 0.623 0.000 | 0.623 0.000 | 0.623 0.000 |
| Minimum variance | 0.230 0.000 | 0.230 0.000 | 0.230 0.000 | 0.230 0.000 | 0.230 0.000 | 0.230 0.000 | 0.230 0.000 | 0.230 0.000 |
| Memory-free | 0.657 0.001 | 0.587 0.007 | 0.713 0.013 | 0.619 0.059 | 0.688 0.022 | 0.597 0.018 | 0.616 0.012 | 0.668 0.020 |
| Self-consistency | 0.686 0.007 | 0.588 0.014 | 0.719 0.012 | 0.614 0.017 | 0.698 0.021 | 0.615 0.024 | 0.618 0.003 | 0.656 0.004 |
| Similarity retrieval (top-2) | 0.649 0.033 | 0.557 0.000 | 0.701 0.019 | 0.602 0.009 | 0.599 0.006 | 0.603 0.009 | 0.656 0.026 | 0.640 0.006 |
| Similarity retrieval (top-4) | 0.643 0.039 | 0.554 0.019 | 0.655 0.006 | 0.598 0.005 | 0.631 0.019 | 0.625 0.016 | 0.586 0.016 | 0.641 0.009 |
| Memory-free | MemTrial (ours) | |||||
| LLM | Cutoff | Months | before | after | before | after |
| gpt-4.1-mini | Jun 2024 | 50/21 | ||||
| gpt-4.1-nano | Jun 2024 | 50/21 | ||||
| gpt-5-mini | May 2024 | 49/22 | ||||
| Llama-3.3-70B | Dec 2023 | 44/27 | ||||
| Qwen3-235B | – | |||||
| Absolute recall | Relative recall | ||||
| LLM | Cutoff | before | after | before | after |
| gpt-4.1-mini | Jun 2024 | 0.02 | 0.12 | 0.01 | 0.03 |
| gpt-4.1-nano | Jun 2024 | 0.04 | 0.04 | 0.06 | 0.09 |
| gpt-5-mini | May 2024 | 0.06 | 0.09 | 0.04 | 0.13 |
| Llama-3.3-70B | Dec 2023 | 0.09 | 0.15 | 0.08 | 0.12 |
| Qwen3-235B | – | 0.03 | – | 0.05 | – |
| gpt-4.1-mini | gpt-5-mini | |||||||
| dates shown | dates hidden | dates shown | dates hidden | |||||
| Method | before | after | before | after | before | after | before | after |
| Memory-free | 0.008 | 0.017 | 0.017 | 0.027 | 0.020 | 0.025 | 0.038 | 0.006 |
| Self-consistency | 0.010 | 0.013 | 0.030 | 0.012 | 0.013 | 0.016 | 0.028 | 0.015 |
| FinMem † | 0.021 | 0.026 | 0.029 | 0.003 | 0.025 | 0.045 | 0.016 | 0.035 |
| MemRL † | 0.030 | 0.033 | 0.032 | 0.050 | 0.046 | 0.034 | 0.049 | 0.041 |
| New months (2025–2026) | All months after the cutoff | |||
| Method | Full | Raw | Full | Raw |
| Rule-based portfolios | ||||
| 1/ | 1.025 0.000 | 1.025 0.000 | 2.075 0.000 | 2.075 0.000 |
| Minimum variance | 0.029 0.000 | 0.029 0.000 | 0.078 0.000 | 0.078 0.000 |
| LLM agents without experience learning | ||||
| Memory-free | 1.825 0.228 | 1.055 0.103 | 1.290 0.277 | 1.363 0.053 |
| 1 April 2022 | 1 September 2022 | ||||||||||||
| Lessons in the draft | Eq. | Bd. | Cm. | RE | Cash | Utility | Lessons in the draft | Eq. | Bd. | Cm. | RE | Cash | Utility |
| none | 35 | 10 | 25 | 20 | 10 | 4.63 | none | 15 | 35 | 15 | 15 | 20 | 5.43 |
| L18/04, L16/08 | 35 | 10 | 20 | 25 | 10 | 5.14 | L18/04, L16/08 | 15 | 30 | 10 | 15 | 30 | 5.05 |
| L18/04, L17/05 | 35 | 10 | 15 | 30 | 10 | 5.67 | L18/04, L19/10 | 10 | 30 | 10 | 10 | 40 | 3.89 |
| L16/08, L17/05 | 35 | 15 | 10 | 30 | 10 | 6.02 | L16/08, L19/10 | 15 | 25 | 10 | 15 | 35 | 4.81 |
| L18/04, L15/12 | 35 | 15 | 15 | 25 | 10 | 5.48 | L18/04, L17/05 | 15 | 35 | 10 | 15 | 25 | 5.29 |
| Method | Memory block |
| Memory-free, Self-consistency | none |
| MemTrial, Similarity retrieval, Counterfactual selection, Uplift credit, Draft averaging | experiences block with a subset of the four retrieved lessons |
| FinMem | experiences block with four lessons ranked by similarity, recency and importance |
| MemRL | experiences block with four of the ten most similar lessons, ranked by learned value ( -greedy) |
| Reflexion | reflections block with its three most recent reflections |
| ExpeL | insights block with at most eight insights distilled from all warm-up lessons |