Modern language-model agents are built around the agent loop: the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain capabilities such as long-term memory and self-improvement currently require specialized systems beyond the agent loop itself. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints and perform monitoring. Generalizing existing code-mode agent loops, invoke is the simplest loop that satisfies two defining properties: (1) the LLM can write arbitrary executable code that can include recursive invoke; (2) everything visible to the LLM - all inputs to invoke as well as its interaction history with the code environment - are variables in the code environment. We motivate our design from first principles, viewing invoke as a language primitive representing a function whose implementation is provided at runtime by an LLM every time it is called. To validate the design of our core invoke primitive, we evaluate invoke - with only prompting, no manually designed tools, harness, or external systems (e.g., memory or the file system) - on workflows traditionally implemented through specialized harnesses. On long-horizon workflows requiring recall far beyond the context window, JAZ invoke outperforms Letta (MemGPT) by 8% at half its cost on the recall-heavy portion of StuLife. On continual self-improvement, JAZ invoke outperforms ACE by 4% at a lower cost on AppWorld.
Figures & tables
Figure 1 : In JAZ, invoke acts as a function whose implementation is provided each time it’s called — by the LLM in a Python REPL. Everything — all named inputs to invoke and the REPL history itself — are variables available in the REPL.
Figure 2 : In the functional version of invoke , the LLM writes a complete implementation for every invocation. A closed agent loop arises from tail-recursive invoke .
StuLife (all)
StuLife (far recall)
prompt- only? *
Pass (%)
Score (%)
Pass (%)
Score (%)
Cost ($)
CodeAct JAZ ( Wang et al., 2024b ) (per task)
✓
52.5 ± 0.3
59.2 ± 0.1
24.8 ± 0.6
25.8 ± 0.5
4.4 ± 0.1
CodeAct+subagents smol ( Roucher et al., 2025 )
✓
30.2 ± 4.8
33.6 ± 5.2
20.6 ± 0.6
22.0 ± 0.9
9.4 ± 1.4
CodeAct+subagents JAZ ( Zhang et al., 2025a )
✓
60.0 ± 1.1
68.1 ± 1.0
32.0 ± 2.3
33.9 ± 2.0
13.1 ± 0.9
Letta (MemGPT) ( Packer et al., 2023 )
✗
70.9 ± 0.5
81.0 ± 0.5
61.8 ± 2.3
67.0 ± 2.4
42.1 ± 1.6
JAZ invoke (ours)
✓
72.6 ± 0.1
81.6 ± 0.1
69.9 ± 1.8
73.6 ± 1.5
18.3 ± 0.3
Table 1: Results on the long-horizon environment StuLife ( Cai et al., 2026 ) . Methods marked “per-task” solve each StuLife task individually with no state or memory persistence across tasks. We report both pass rate (fraction of scored tasks with a perfect score) and average score, for both the full scored set of 939 tasks (“all”) and the subset of 207 tasks that require recall of information delivered over 50 tasks ago (“far recall”). We use GPT-5.4 nano (high) across all methods. We report the mean and its standard error over 3 independent runs. See Table 3 for all individual data points.
AppWorld ( test-challenge )
prompt- only? *
TGC (%)
SGC (%)
Cost ($)
Meta $
Solver $
CodeAct AppWorld ( Trivedi et al., 2024 ) (per-task)
✗ †
48.2 ± 1.6
20.4 ± 3.2
16.6 ± 0.2
—
16.6 ± 0.2
CodeAct JAZ ( Wang et al., 2024b ) (per-task)
✓
67.5 ± 1.2
43.4 ± 2.5
10.2 ± 0.2
—
10.2 ± 0.2
CodeAct+subagents JAZ ( Zhang et al., 2025a )
✓
71.1 ± 1.3
47.6 ± 2.1
21.6 ± 1.8
7.3 ± 1.7
14.3 ± 0.9
ACE ( Zhang et al., 2025b ) on CodeAct JAZ
✗
69.9 ± 1.4
47.1 ± 1.3
30.6 ± 0.5
17.1 ± 0.3
13.5 ± 0.3
JAZ invoke (ours)
✓
74.2 ± 2.1
51.1 ± 3.7
20.9 ± 3.7
9.9 ± 3.8
10.9 ± 0.6
Table 2: Results on the full test-challenge split of AppWorld ( Trivedi et al., 2024 ) . Methods marked “per-task” solve each AppWorld task individually with no continual learning across tasks. The solver agent in all methods uses GPT-5.4 nano (high) and the meta-agent in CSI methods uses GPT-5.4 (high). We report the mean and standard error over n independent runs, where n=3 for non-self-improving methods and n=6 for self-improving methods to account for higher variance. See Table 4 for all individual data points.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
StuLife [ 5 ] (all)
StuLife (far recall)
run #
Pass (%)
Score (%)
Pass (%)
Score (%)
Cost ($)
CodeAct JAZ [ 24 ] (per task)
1
52.3
59.2
25.1
26.1
4.6
2
52.2
58.9
25.6
26.5
4.2
3
53.0
59.4
23.7
24.7
4.3
mean
52.5 ±0.3
59.2 ±0.1
24.8 ±0.6
25.8 ±0.5
4.4 ±0.1
CodeAct+subagents [ 19 ]
1
29.5
31.8
20.3
22.3
11.0
Appendix
Table 3: Expanded version of Table 1 : full results on the full sequence of tasks of StuLife [ 5 ] . Here, in addition to the mean, we also report every individual data point.
AppWorld ( test-challenge )
run #
TGC (%)
SGC (%)
Cost ($)
Meta $
Solver $
CodeAct AppWorld [ 21 ] (per-task)
1
46.8
15.8
16.6
—
16.6
2
46.5
18.7
16.6
—
16.6
3
51.3
26.6
16.6
—
16.6
mean
48.2 ±1.6
20.4 ±3.2
16.6 ±0.0
—
16.6 ±0.0
CodeAct JAZ [ 24 ] (per-task)
1
69.1
48.2
10.4
—
10.4
Appendix
Table 4: Expanded version of Table 2 : full results on the full test-challenge split of AppWorld [ 21 ] . Here, in addition to the mean, we also report every individual data point. For self-improving methods, we also report the median and its 78% confidence interval formed by the 2nd and 5th order statistics. Notes: 1) The 0.0 uncertainty in the official baseline cost is by coincidence under only 3 runs — in the main table we report 0.2 , the quadrature over tasks of the SEM cost of each task over runs. 2) The substantially higher value of the median than the mean in JAZ invoke ’s TGC and SGC is due to the outlier run 5, which scored below the CodeAct baseline.