LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, τ2-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.
Figures & tables
Figure 1: Daedalus memory from self-generated tasks improves agent success and efficiency across model families.
Figure 2: Overview of the Daedalus memory generation pipeline.
Figure 3: The Solver loop. After each failure, the Extractor writes or revises a heuristic and the Solver retries from a fresh environment state. The heuristic is accepted only after the Solver reaches the required streak of successes Ns .
Figure 4: Timeline of AppWorld generation session 78. The Explorer first proposes a task that the Solver completes without failure, so the Explorer refines it to be harder. On the refined task, two failures yield successive heuristics h1 and h2 . With h2 in context, the Solver succeeds three consecutive times, validating its effectiveness; h2 is therefore accepted into the memory bank. Appendix F.2 shows the session in greater detail.
Daedalus memory
Base agent
Auxiliary
Solver
GPT-5.4-mini
Qwen3.6-35B-A3B
DeepSeek-V4-Flash
No-memory baseline
44.3 ±1.1
44.8 ±1.3
80.6 ±1.6
GPT-5.4
GPT-5.4-mini
+15.9 ±1.4
+16.1 ±1.6
+8.3 ±1.8
Qwen3.8-Flash
Qwen3.6-35B-A3B
+16.8 ±2.4
+29.3 ±1.7
+4.9 ±2.0
DeepSeek-V4-Pro
DeepSeek-V4-Flash
+11.4 ±3.6
+14.5 ±1.3
+3.0 ±3.0
Table 2: Cross-family transfer of Daedalus heuristics on AppWorld. First row: MSR of each base agent (column) without memory. Others: MSR improvement when the base agent uses the Daedalus heuristic bank generated by an auxiliary / solver pair (row).
Figure 5: Generation cost by agent role for (d) , (e) , and the complete pipeline.
Benchmark
N
Precision
Recall
κbench
κinter
AppWorld
168
.943
.846
.807
.848
τ2 -bench
40
.889
.842
.749
1.000
AutomationBench
70
.863
.807
.729
.902
Table 4: Agreement between the LLM judge and official verifiers, and across five judge calls per trajectory. Precision and recall treat success as positive.
Figure 6: Scaling with the exploration budget on AppWorld, over a single generation run. Error bars are standard errors over five inference runs.
Aux. model
MSR ↑
pass^5 ↑
Gen. \downarrow$
no memory
44.3 ±1.1
14.9 ±2.8
0
GPT-5.4-mini
51.2 ±1.6
20.2 ±3.1
67.2
GPT-5.4
60.2 ±0.9
32.1 ±3.6
109.7
Table 5: Effect of auxiliary model choice.
Figure 7: BM25 retrieval per turn with increasing k . Dotted lines: whole bank at start; k=0 : no memory.
Figure 8: Performance on Daedalus -generated tasks versus the AppWorld test set, for nine models evaluated without memory.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Failures and refinements per accepted heuristic over the 90-session AppWorld run.
Extractor
Reasoning effort
# Heuristics
MSR ↑
pass^5 ↑
no memory
–
–
44.3 ±1.1
14.9 ±2.8
GPT-5.4-mini
none
26
53.8 ±1.4
28.6 ±3.5
GPT-5.4-mini
high
38
57.1 ±1.1
28.6 ±3.5
GPT-5.4
high
40
60.8 ±0.9
36.3 ±3.7
Appendix
Table 8: Extractor model choice for Daedalus -curated on AppWorld.
Figure 10: Heuristics injection modes at inference.
Query source
MSR ↑
pass^5 ↑
\downarrow$
No memory
44.3 ±1.1
14.9 ±2.8
3.2
Pre-Gen
52.0 ±1.3
22.0 ±3.2
3.8
R2R
44.2 ±1.2
11.9 ±2.5
14.8
Appendix
Table 9: Retrieval query source on AppWorld. BM25, k=5 , same consolidated bank.
Injection mode
MSR ↑
pass^5 ↑
\downarrow$
Distinct
No memory
44.3 ±1.1
14.9 ±2.8
3.2
0
Retrieved subset, k=5 per turn
Transient
48.5 ±0.7
13.1 ±2.6
2.8
10.8
Cumulative, with repetition
50.0 ±1.7
17.3 ±2.9
3.1
11.7
Cumulative, deduplicated
54.3 ±1.6
24.4 ±3.3
3.7
12.1
Cumulative, no replacement
56.0 ±1.2
25.6 ±3.4
3.9
46.0
Appendix
Table 10: Injection modes on AppWorld. Retrieved subsets use Qwen3-Embedding-4B with Pre-Gen queries and k=5 . Distinct: mean number of distinct heuristics seen by the solver at episode end. Bold marks the best value in each column and underline the second best.
Benchmark
Solver / base agent
Auxiliary agents
Step ceiling
AppWorld
gpt-5.4-mini (default)
gpt-5.4 (high)
30
τ2 -bench
gpt-5.4-mini (default)
gpt-5.4 (high)
200
AutomationBench
gpt-5.6-luna (medium)
gpt-5.6-terra (high)
50
Appendix
Table 11: Models and solver step ceiling for each benchmark (reasoning effort in parentheses).
Hyperparameter
Value
Explorer turn ceiling
40
Concurrent explorers
5
Max task refinements Nr
5
Max failed attempts Nf
8
Consecutive successes to accept Ns
3
Appendix
Table 12: Daedalus hyperparameters.
Method
Hyperparameter
Value
ExpeL
Max Reflexion retries per task Z
3
Successes per comparison L
3
Max success/failure pairs per task
3
Insight list cap
20
Few-shot examples k
2
Retrieval embedder
all-mpnet-base-v2
Appendix
Table 13: Baseline hyperparameters.
Method
Solver rollouts
Per task
MSR
ERL
90
1.0
58.6
ReasoningBank
90
1.0
48.0
ACE
90
1.0
60.5
ExpeL
200
2.2
59.0
AutoGuide
200
2.2
44.0
Daedalus -curated
592
6.6
60.8
Appendix
Table 14: Solver rollouts spent building memory on AppWorld. Each Daedalus session counts as one training task. Explorer interactions are not included.
Task example
Please text everyone from my latest friends dinner note to let them know how much they owe me for that dinner.
Success conditions
- Compared with the starting state, every person listed in the latest friends dinner note has a new outgoing text message from Kristin in that contact’s message thread. - Each new text correctly states the amount assigned to that person in the note and makes clear it is about that dinner.
Accepted heuristic
- When using simple_note , the available note-reading methods are search_notes and show_note ; there is no list_notes , so search by query first and then open the chosen note_id . - When API-doc search results are suppressed, print dir(apis.<app>) to discover exact method names and proceed from those concrete names instead of guessing endpoints. - When a cross-app workflow needs another app after notes, get and use that app’s own access token; tokens are app-specific and a token from one app will 401 on another app’s endpoints. - When a login with the main profile email fails for the phone app, try authenticating with the user’s phone number as the username, because that app may use phone-number-based login rather than email. - When a search query returns noisy matches, do not trust the query alone; inspect the returned titles and timestamps and open the candidate note to confirm it is the latest relevant one before acting on its contents.
Appendix
Table 15: AppWorld example
Figure 11: Session 78 from the Daedalus generation process on AppWorld (partial).
Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusing knowledge. Prior memory approaches aim to resolve the situation, but mainly focus on storing factual information. Recent work on procedural memory improves task reuse, yet often reduces to replaying past successes without addressing failure cases or online scalability. We introduce a unified and automatic memory framework that integrates semantic, episodic, and procedural memory in a bi-level design combining short-term and long-term stores. A multi-agent architecture with actor, memory, and critic agents enables automatic memory generation, reward annotation, and adaptive retrieval. Long-term memory is managed through reward-based evaluation, merging, and pruning, ensuring scalability and continual improvement. Experiments across various environments show that our approach improves robustness and success on long multi-turn tasks compared to existing baselines. This work highlights the importance of comprehensive, adaptive memory for advancing LLM-based agents.
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
University of California Los Angeles · University of Washington · Northwestern University
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory. AMD constructs three complementary memory types from successful teacher trajectories: Workflow memory encodes task-level strategies, Subtask memory provides concrete behavioral examples at an intermediate granularity, and Function memory captures per-function calling conventions and common pitfalls. Workflow and Subtask memories are injected proactively at the start of each task, while Function memory is retrieved reactively upon tool-calling errors. We evaluate AMD on three tool-use benchmarks using four student models (4B-8B parameters) with GPT-5-mini as the teacher, achieving average accuracy gains of 27.2%p, 11.2%p, and 3.4%p on AppWorld, BFCL V3, and ToolSandbox, while consistently outperforming existing memory-based baselines. Further analysis shows that Subtask memory contributes the largest gains, teacher effectiveness depends on both teacher capability and student compatibility, and 4B-sized students benefit most from AMD.