Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability
Figures & tables
Figure 1 : LiveMACEBench overview. Agents interact with a shared, continuously evolving live-market environment along persistent trajectories. Evaluation combines realized outcomes with mechanism-specific process diagnostics to distinguish task performance from how effectively agents use each mechanism.
Figure 2 : Capability evaluation in a shared market environment. (a) Daily cumulative-return and Tool Call Score (TCS) rankings for the same Tool-use accounts. Rank 1 is highest; annotations count pairwise reversals among comparable pairs. (b) Full-period TCS , Memory Usage Quality (MUQ) , Rule Satisfaction Score (RSS) , and Observable Collaboration Score (OCS) , all on [0,1] . Colors identify models; bold marks each dimension’s highest score, and dashes indicate unavailable scores. Rankings for the other dimensions are shown in Appendix C.7 .
Figure 3 : Tool-use capability components across four models. Left: Judge scores produced by Qwen3.5-397B-A17B over 170 traces per evaluated backbone: Tool Relevance (TR) , Invocation Efficiency (IE) , Information Coverage (IC) , and Evidence Faithfulness (EF) . Right: Routing Quality (RQ) and Valid Call Rate (VCR) . All scores are shown on a 0–1 scale, with higher values indicating better performance. Hallucination-Free Rate (HFR) is omitted because all four models score 1. Additional judge results and complete metrics are provided in Appendix C.3 .
Memory Interaction
Memory Quality
Downstream Effect
Model
#Search
#Add
#Stored
MCQ ↑
MUQ ↑
Δ MDD ↑
Δ TL ↑
Δ PnL (Full → Late)
GPT-5.4
332
1
1
.978
.762
−.10
+.13
−1.46→−1.80
DeepSeek
712
55
41
.938
.680
+2.57
+.40
+1.90→+2.14
Gemini
319
97
93
.938
.497
+.06
+.05
−1.32→+.53
Grok
291
12
12
.760
.855
+3.30
+.83
+2.91→+3.18
Qwen
415
15
11
.878
.742
+1.51
+.38
−6.50→−4.85
Table 1 : Persistent-memory lifecycle across backbones. Memory interaction statistics characterize experience accumulation; MCQ and MUQ measure construction and utilization quality; downstream metrics report percentage-point changes relative to the matched memory-free configuration. Δ PnL compares the full period with the late phase (Apr. 23 onward).
Model
Final Return (%) ↑
MDD (%) ↓
HPR ↑
RSS ↑
RAS ↑
GPT-5.4
+5.738%
5.169%
0.682
0.634
0.677
DeepSeek-V3.2
+5.765%
4.270%
0.877
0.847
0.623
Gemini-3.1-Pro
+3.503%
4.927%
0.888
0.872
0.882
Grok-4.20
+3.085%
5.165%
0.834
0.751
0.775
Qwen3-Max
+4.418%
5.674%
0.613
0.612
0.511
Table 2: Trading outcomes and rule-following metrics. HPR denotes Hard-rule Pass Rate , RSS denotes Rule Satisfaction Score , and RAS denotes Rule Auditability Score . MDD denotes maximum drawdown; arrows indicate the preferred direction.
Figure 4 : Rule execution and auditability across models. (a) RSS versus RAS , with marker color indicating final return and dashed lines denoting cross-model means. RSS denotes Rule Satisfaction Score , and RAS denotes Rule Auditability Score . (b) LLM-auditor diagnostic scores, where RC denotes Rule Coverage and CH denotes Conflict Handling .
Figure 5 : Equity trajectories under Multi-Agent Configuration and Base ReAct Configuration across four backbone models. The dotted line marks the initial capital of $10,000.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Tool
Function
get_market_snapshot
Retrieves the latest price and market status for a crypto asset or U.S. stock.
get_kline_history
Fetches candlestick history under preset short-, medium-, or long-horizon modes and saves the data to the sandbox.
get_account_state
Returns current account balances and open positions.
get_history_decisions
Retrieves recent trading decisions together with account-level profit-and-loss information.
consult_search_agent
Searches for market news, macroeconomic information, project updates, and other non-price evidence.
execute_shell_command
Executes shell commands in the sandboxed analysis environment.
Appendix
Table 5: Fixed tool set of the Base configuration.
Role
Responsibility
Available tools
Manager
Selects specialists, integrates evidence, and forms the final plan.
None
TradingAgent
Analyzes the market, technical signals, and portfolio exposure.
Retrieves recent market events and external signals.
consult_search_agent
CoderAgent
Performs targeted quantitative verification.
run_python_script file tools shell tools
AnalystAgent
Reconciles evidence and makes conflicts explicit.
None
CriticAgent
Challenges the thesis and identifies downside or veto conditions.
None
Appendix
Table 6: Roles and tool permissions in the Multi-Agent configuration.
Figure 7 : Weekly market returns during the evaluation period for the cryptocurrency basket, U.S. equity basket, and buy-and-hold reference.
Model
Judge mean
TR
IE
IC
EF
GPT-5.4
0.8371
0.8988
0.6388
0.9112
0.8994
Qwen3-Max
0.8547
0.8959
0.7312
0.9012
0.8906
DeepSeek-V3.2
0.7429
0.8118
0.5812
0.8341
0.7447
Grok-4.20
0.7254
0.7665
0.5871
0.8085
0.7397
Appendix
Table 8: Normalized LLM-judge scores (0–1) for tool use: Tool Relevance (TR) , Invocation Efficiency (IE) , Information Coverage (IC) , and Evidence Faithfulness (EF) . Judge mean is the average of these four dimensions.
Model
Calls/step
VCR
RQ
Avg. tokens
TCS
HFR
GPT-5.4
4.212
0.9983
0.9200
30.2k
0.8941
1.000
Qwen3-Max
0.947
0.9780
0.7364
18.6k
0.8641
1.000
DeepSeek-V3.2
0.967
0.8254
0.7846
50.9k
0.7894
1.000
Grok-4.20
2.197
0.6659
0.8052
21.9k
0.7599
1.000
Appendix
Table 9: Objective tool-use metrics and aggregate scores: Valid Call Rate (VCR) , Routing Quality (RQ) , Tool Call Score (TCS) , and Hallucination-Free Rate (HFR) . All four scores are reported on a 0–1 scale, with higher values indicating better performance. Calls/step and average tokens describe execution cost.
Figure 8 : Judge-specific scores for the four tool-use dimensions across Tool-Augmented ReAct agents: Tool Relevance (TR) , Invocation Efficiency (IE) , Information Coverage (IC) , and Evidence Faithfulness (EF) . Scores are shown on a 0–1 scale.
Figure 9 : Cross-judge robustness of tool-use evaluation on the 0–1 scale. (a) Average judge scores by model. (b) Agreement on Tool Relevance (TR) , Invocation Efficiency (IE) , Information Coverage (IC) , and Evidence Faithfulness (EF) , measured by mean absolute error (MAE) and the fraction of paired scores within 0.1 points.
Model
Memory
Return
TL
Max Loss Streak
MDD
GPT-5.4
w/ Mem
+ 0.56%
− 0.31%
3
1.51%
w/o Mem
+ 2.36%
− 0.43%
5
1.41%
DeepSeek-V3.2
w/ Mem
+ 0.29%
− 0.04%
2
0.10%
w/o Mem
− 1.85%
− 0.48%
6
2.23%
Gemini-3.1-Pro
w/ Mem
+ 2.89%
− 0.51%
5
1.56%
w/o Mem
+ 2.36%
− 0.66%
5
2.42%
Appendix
Table 11 : Late-phase performance from April 23 onward, after approximately ten days of memory accumulation. TL denotes Tail Loss , and MDD denotes Maximum Drawdown .
Model
Judge
MCQ
MUQ
MS
GPT-5.4
GPT
1.000
0.625
0.430
DeepSeek
1.000
0.755
0.463
Qwen
0.933
0.905
0.484
Avg
0.978
0.762
0.459
DeepSeek-V3.2
GPT
0.927
0.630
0.700
DeepSeek
0.947
0.655
0.711
Appendix
Table 12 : Per-judge persistent-memory scores. MCQ denotes Memory Content Quality , MUQ denotes Memory Usage Quality , and MS denotes Memory Score . For each model and judge, MS combines that judge’s MCQ and MUQ with the model’s fixed programmatic ΔMDDnorm and ΔTLnorm components defined in Appendix B.4.2 . Avg denotes the arithmetic mean across the three judges.
Figure 10 : Account-equity trajectories for all ten matched accounts over the full evaluation period. Solid lines denote w/ Mem and dashed lines denote w/o Mem . The largest matched downside differences are visible for Grok-4.20 and DeepSeek-V3.2;
Figure 11 : Overall trading and rule-following results across models. The left panel reports final return and maximum drawdown (MDD); the right panel reports Hard-rule Pass Rate (HPR) , Rule Satisfaction Score (RSS) , and Rule Auditability Score (RAS) .
Figure 12 : Additional diagnostics of rule-following behavior. (a) Rule-level event rates for hard-rule violations and soft-rule deviations, highlighting which R0–R2 rules account for the main execution bottlenecks. (b) Final return versus hard-rule failure rate, defined as 1−HPR , where HPR denotes Hard-rule Pass Rate . Marker color represents Rule Auditability Score (RAS) , and marker size represents maximum drawdown (MDD).
Audited agent
Qwen3.5-397B-A17B
DeepSeek-V4-Pro
GPT-5.5
Mean
GPT-5.4
0.677
0.744
0.586
0.669
DeepSeek-v3.2
0.623
0.680
0.516
0.606
Gemini-3.1-pro
0.882
0.782
0.645
0.770
Grok-4.20
0.775
0.699
0.652
0.709
Qwen3-Max
0.511
0.545
0.428
0.495
Appendix
Table 13: Cross-auditor comparison of Rule Auditability Score (RAS) . Qwen3.5-397B-A17B and DeepSeek-V4-Pro are evaluated without thinking, while GPT-5.5 uses reasoning-effort=none . The final column reports the mean RAS across the three auditors.
Model
Traces
Reported OCS
Alternative OCS
95% CI
DeepSeek-V3.2
161
0.806
0.726
[0.783, 0.830]
Gemini-3.1-Pro
168
0.738
0.695
[0.721, 0.755]
Qwen3-Max
170
0.516
0.485
[0.475, 0.559]
GPT-5.4
170
0.265
0.265
[0.255, 0.277]
Appendix
Table 14: Observable Collaboration Score (OCS) robustness. Reported OCS uses CriticAgent and excludes CoderAgent from role breadth; the alternative reverses that choice. The 95% confidence intervals (CI) apply to the reported OCS .
Figure 13 : Daily cumulative-return and capability rankings for tool use, memory, rule following, and collaboration. Capability scores are computed from all available evaluation records from the experiment start through each observation time: Tool Call Score (TCS) , Memory Usage Quality (MUQ) , Rule Satisfaction Score (RSS) , and Observable Collaboration Score (OCS) . Each row compares matched accounts and observation times; rank 1 is highest. Annotations count pairwise ranking reversals among comparable pairs.
Figure 14 : Positive tool-use case study. The upper panel shows the GPT-5.4 Tool-Augmented and Base ReAct account-equity trajectories. The lower panel indexes NVDA to the Tool Agent’s April 13 execution price; vertical lines mark the Tool and Base entries. Bars report, for each UTC day, the root mean square of observed hourly returns across the benchmark asset universe.
Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to successfully complete a task and the cost of doing so. We propose MarketBench, a benchmark for assessing whether AI agents have these capabilities. We use a 93-task subset of SWE-bench Lite, a software engineering benchmark, with six recently released LLMs as a demonstration. These LLMs are miscalibrated on both success probability and token usage, and auctions built from these self-reports diverge from a full-information allocation. A follow-up intervention where we add information about capabilities from prior experiments to the context improves calibration, but only modestly narrows the gap to a full-information benchmark. We also document the performance of a market-based scaffolding with these LLMs. Our results point to self-assessment as a key bottleneck for market-style coordination of AI agents.
Andrey Fradkin, Rohit Krishnan
Boston University and the MIT Initiative on the Digital Economy. · Independent Researcher.
Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as market context accumulates over long horizons, these mandates gradually lose their behavioral influence, a phenomenon we formalize as Mandate Salience Decay (MSD). To measure MSD objectively, we introduce FinPersona-Bench, a simulation benchmark in which a synthetic market decouples observable price from hidden fundamental value, enabling falsifiable evaluation across three failure modes: trading without signal in calm markets, panic-selling during crashes, and ignoring fundamental value during speculative bubbles. Evaluating 18 leading frontier and open-source LLMs, each assigned one of three behavioral profiles ranging from strict capital preservation to aggressive growth, shows that MSD compounds over time and is model-dependent. In crash scenarios, the behavioral gap between static agents and those receiving periodic mandate re-grounding grows 4.4x from the first to the final quarter of the simulation. The effects of mandate re-grounding are not uniformly positive: it consistently helps conservative agents in low-signal markets but actively worsens behavior for aggressive agents in the same setting. These findings suggest that reliable long-horizon deployment requires selective, mandate-aware re-grounding based on agent profile and market regime.
Muhammad Usman Safder, Ayesha Gull, Rania Elbadry +7
Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment reasoning. Second, raw returns are a noisy proxy for stock-selection ability, since positive performance may come from market beta, style exposure, or favorable regimes rather than genuine alpha. We introduce KTD-Fin (Knowing-To-Doing Financial Benchmark), an end-to-end stock-market trading benchmark that addresses both issues. KTD-Fin uses a data-side masking protocol to anonymize key identifiers and calendar information consistently across prompts and tools, separating historical market memory from investment decision-making. It also incorporates a Barra-style performance attribution framework that decomposes portfolio returns into market, style, and stock-selection alpha components. Across ten frontier LLM agents evaluated on the Chinese CSI300 over a 2024--2026 window, masking substantially changes agent rationales, pushing them towards anonymized factor-based reasoning. Attribution analysis further shows that LLM agents' cumulative returns under leakage-controlled evaluation are largely explained by passive market and style exposure, with limited evidence of persistent stock-selection alpha. These findings suggest that financial LLM benchmarks should evaluate not only whether an agent makes money, but also whether the source of returns reflects transferable investment skill. We release KTD-Fin as a reproducible template for leakage-controlled and attribution-aware evaluation of LLM trading agents.