LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, τ2-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.
Figures & tables
Figure 1: Daedalus memory from self-generated tasks improves agent success and efficiency across model families.
Figure 2: Overview of the Daedalus memory generation pipeline.
Figure 3: The Solver loop. After each failure, the Extractor writes or revises a heuristic and the Solver retries from a fresh environment state. The heuristic is accepted only after the Solver reaches the required streak of successes Ns .
Figure 4: Timeline of AppWorld generation session 78. The Explorer first proposes a task that the Solver completes without failure, so the Explorer refines it to be harder. On the refined task, two failures yield successive heuristics h1 and h2 . With h2 in context, the Solver succeeds three consecutive times, validating its effectiveness; h2 is therefore accepted into the memory bank. Appendix F.2 shows the session in greater detail.
Daedalus memory
Base agent
Auxiliary
Solver
GPT-5.4-mini
Qwen3.6-35B-A3B
DeepSeek-V4-Flash
No-memory baseline
44.3 ±1.1
44.8 ±1.3
80.6 ±1.6
GPT-5.4
GPT-5.4-mini
+15.9 ±1.4
+16.1 ±1.6
+8.3 ±1.8
Qwen3.8-Flash
Qwen3.6-35B-A3B
+16.8 ±2.4
+29.3 ±1.7
+4.9 ±2.0
DeepSeek-V4-Pro
DeepSeek-V4-Flash
+11.4 ±3.6
+14.5 ±1.3
+3.0 ±3.0
Table 2: Cross-family transfer of Daedalus heuristics on AppWorld. First row: MSR of each base agent (column) without memory. Others: MSR improvement when the base agent uses the Daedalus heuristic bank generated by an auxiliary / solver pair (row).
Figure 5: Generation cost by agent role for (d) , (e) , and the complete pipeline.
Benchmark
N
Precision
Recall
κbench
κinter
AppWorld
168
.943
.846
.807
.848
τ2 -bench
40
.889
.842
.749
1.000
AutomationBench
70
.863
.807
.729
.902
Table 4: Agreement between the LLM judge and official verifiers, and across five judge calls per trajectory. Precision and recall treat success as positive.
Figure 6: Scaling with the exploration budget on AppWorld, over a single generation run. Error bars are standard errors over five inference runs.
Aux. model
MSR ↑
pass^5 ↑
Gen. \downarrow$
no memory
44.3 ±1.1
14.9 ±2.8
0
GPT-5.4-mini
51.2 ±1.6
20.2 ±3.1
67.2
GPT-5.4
60.2 ±0.9
32.1 ±3.6
109.7
Table 5: Effect of auxiliary model choice.
Figure 7: BM25 retrieval per turn with increasing k . Dotted lines: whole bank at start; k=0 : no memory.
Figure 8: Performance on Daedalus -generated tasks versus the AppWorld test set, for nine models evaluated without memory.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Failures and refinements per accepted heuristic over the 90-session AppWorld run.
Extractor
Reasoning effort
# Heuristics
MSR ↑
pass^5 ↑
no memory
–
–
44.3 ±1.1
14.9 ±2.8
GPT-5.4-mini
none
26
53.8 ±1.4
28.6 ±3.5
GPT-5.4-mini
high
38
57.1 ±1.1
28.6 ±3.5
GPT-5.4
high
40
60.8 ±0.9
36.3 ±3.7
Appendix
Table 8: Extractor model choice for Daedalus -curated on AppWorld.
Figure 10: Heuristics injection modes at inference.
Query source
MSR ↑
pass^5 ↑
\downarrow$
No memory
44.3 ±1.1
14.9 ±2.8
3.2
Pre-Gen
52.0 ±1.3
22.0 ±3.2
3.8
R2R
44.2 ±1.2
11.9 ±2.5
14.8
Appendix
Table 9: Retrieval query source on AppWorld. BM25, k=5 , same consolidated bank.
Injection mode
MSR ↑
pass^5 ↑
\downarrow$
Distinct
No memory
44.3 ±1.1
14.9 ±2.8
3.2
0
Retrieved subset, k=5 per turn
Transient
48.5 ±0.7
13.1 ±2.6
2.8
10.8
Cumulative, with repetition
50.0 ±1.7
17.3 ±2.9
3.1
11.7
Cumulative, deduplicated
54.3 ±1.6
24.4 ±3.3
3.7
12.1
Cumulative, no replacement
56.0 ±1.2
25.6 ±3.4
3.9
46.0
Appendix
Table 10: Injection modes on AppWorld. Retrieved subsets use Qwen3-Embedding-4B with Pre-Gen queries and k=5 . Distinct: mean number of distinct heuristics seen by the solver at episode end. Bold marks the best value in each column and underline the second best.
Benchmark
Solver / base agent
Auxiliary agents
Step ceiling
AppWorld
gpt-5.4-mini (default)
gpt-5.4 (high)
30
τ2 -bench
gpt-5.4-mini (default)
gpt-5.4 (high)
200
AutomationBench
gpt-5.6-luna (medium)
gpt-5.6-terra (high)
50
Appendix
Table 11: Models and solver step ceiling for each benchmark (reasoning effort in parentheses).
Hyperparameter
Value
Explorer turn ceiling
40
Concurrent explorers
5
Max task refinements Nr
5
Max failed attempts Nf
8
Consecutive successes to accept Ns
3
Appendix
Table 12: Daedalus hyperparameters.
Method
Hyperparameter
Value
ExpeL
Max Reflexion retries per task Z
3
Successes per comparison L
3
Max success/failure pairs per task
3
Insight list cap
20
Few-shot examples k
2
Retrieval embedder
all-mpnet-base-v2
Appendix
Table 13: Baseline hyperparameters.
Method
Solver rollouts
Per task
MSR
ERL
90
1.0
58.6
ReasoningBank
90
1.0
48.0
ACE
90
1.0
60.5
ExpeL
200
2.2
59.0
AutoGuide
200
2.2
44.0
Daedalus -curated
592
6.6
60.8
Appendix
Table 14: Solver rollouts spent building memory on AppWorld. Each Daedalus session counts as one training task. Explorer interactions are not included.
Task example
Please text everyone from my latest friends dinner note to let them know how much they owe me for that dinner.
Success conditions
- Compared with the starting state, every person listed in the latest friends dinner note has a new outgoing text message from Kristin in that contact’s message thread. - Each new text correctly states the amount assigned to that person in the note and makes clear it is about that dinner.
Accepted heuristic
- When using simple_note , the available note-reading methods are search_notes and show_note ; there is no list_notes , so search by query first and then open the chosen note_id . - When API-doc search results are suppressed, print dir(apis.<app>) to discover exact method names and proceed from those concrete names instead of guessing endpoints. - When a cross-app workflow needs another app after notes, get and use that app’s own access token; tokens are app-specific and a token from one app will 401 on another app’s endpoints. - When a login with the main profile email fails for the phone app, try authenticating with the user’s phone number as the username, because that app may use phone-number-based login rather than email. - When a search query returns noisy matches, do not trust the query alone; inspect the returned titles and timestamps and open the candidate note to confirm it is the latest relevant one before acting on its contents.
Appendix
Table 15: AppWorld example
Figure 11: Session 78 from the Daedalus generation process on AppWorld (partial).