Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. A second harness raises the original system's macro-average score by 2.6 points across eleven benchmarks. Across four host configurations, EpiCon improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67% to 74% relative to backbone-sized memory models.
Figures & tables
Figure 1: Overview of EpiCon. The memory controller refines question-level memory, while the tree self-organizer maintains accumulated lessons for subsequent retrieval and reuse. The memory harness coordinates these operations with MAS execution. Bank maintenance occurs during construction; the historical bank remains frozen during evaluation.
Figure 2: Illustrative example of EpiCon. (a) Textual and visual memory co-evolve across attempts. (b) and (c) Related lessons are organized and consolidated into a reusable routing rule. (d) Retrieved guidance supports a single solving attempt on a new map. Trajectories and tree operations are reconstructed for illustration.
MAS Harness
MAS Backbone
Memory System
Document Understanding
Visual-to-Code
Vision-Grounded Math
General VL Reasoning
Time ( × )
Tokens ( × )
DocVQA2026
MP-DocVQA
ParseBench
Vision2Code
Omni-I2C
ChartMimic
MATH-Vision
WeMath
WorldBench
ReasonMap
BabyVision
MAS
Memory
MAS
Memory
No Memory
15.7
79.5
57.1
53.0
66.2
66.4
20.4
81.0
49.2
40.1
14.2
1.00
0.00
1.00
0.00
Mem0 (Qwen3.8-27B)
1.4
81.6
61.9
61.3
70.6
78.0
24.4
80.6
48.6
39.6
12.2
1.33
0.31
1.23
0.11
Cognee (Qwen3.8-27B)
14.3
73.7
60.1
61.5
69.9
78.1
44.2
78.2
51.2
35.4
11.5
1.16
1.66
1.11
0.75
A-Mem (Qwen3.8-27B)
12.9
80.2
66.3
60.0
72.3
78.6
25.4
81.8
45.4
40.6
13.5
1.23
0.24
1.27
0.81
Agent-KB (Qwen3.8-27B)
11.4
78.6
59.5
25.8
67.4
69.8
25.2
80.8
45.0
24.1
15.6
1.54
0.31
1.12
0.14
Table 1: Main results across two MAS harnesses, two backbones, and eleven benchmarks. All methods make a single solving attempt per question, without question-level memory updates. Both EpiCon variants use MC + Tree. Parentheses identify the memory models: ours 2B uses two independently trained 2B models; named backbones operate zero-shot. Time and token usage separate MAS execution from memory operations.
MAS Harness
MAS Backbone
Memory System
Document Understanding
Visual-to-Code
Vision-Grounded Math
General VL Reasoning
DocVQA2026
MP-DocVQA
ParseBench
Vision2Code
Omni-I2C
ChartMimic
MATH-Vision
WeMath
WorldBench
ReasonMap
BabyVision
Codex
Gemma4-31B
No Memory
5.7
46.2
43.3
51.7
65.9
65.6
25.8
85.0
45.4
31.6
17.7
EpiCon (ours 2B)
5.7
60.4
45.3
53.8
67.8
68.2
33.0
85.0
47.8
34.4
18.1
DeepSeek Harness
Qwen3.8-27B
No Memory
14.3
79.5
52.4
53.7
61.4
65.3
42.2
88.8
54.0
22.6
20.1
EpiCon (ours 2B)
18.6
79.1
65.7
59.2
63.3
68.7
56.0
90.6
52.2
26.4
20.5
Table 2: Cross-backbone and cross-harness experience transfer from a bank built by Codex with Qwen3.8-27B. The evaluated MAS performs one solve; EpiCon (ours 2B) reuses the frozen bank without question-level updates.
MAS Harness
MAS Backbone
Memory Bank
Document Understanding
Visual-to-Code
Vision-Grounded Math
General VL Reasoning
DocVQA2026
MP-DocVQA
ParseBench
Vision2Code
Omni-I2C
ChartMimic
MATH-Vision
WeMath
WorldBench
ReasonMap
BabyVision
Codex
Qwen3.8-27B
Original bank
18.6
66.1
63.4
59.9
71.4
75.2
45.2
84.2
52.4
38.3
14.3
Evolved bank
18.6
80.3
58.9
61.3
70.9
76.7
45.4
88.4
55.2
40.1
16.8
DeepSeek Harness
Qwen3.8-27B
Original bank
18.6
81.8
52.3
56.9
63.5
65.2
56.6
92.2
57.4
23.5
19.3
Evolved bank
20.0
82.7
65.9
58.7
69.2
72.1
53.4
91.2
57.6
29.6
22.3
Table 3: Cross-harness memory evolution with EpiCon (ours 2B). Each evaluated harness builds the original bank; the other harness uses and evolves it. The two stages each use 1,000 disjoint construction questions. Both bank conditions use a single solving attempt.
MAS
MAS Backbone
Memory Setting
Task Performance
Time ( × )
Tokens ( × )
Document ParseBench
Visual-to-Code Vision2Code
Vision Math MATH-Vision
General VL BabyVision
MAS
Memory
MAS
Memory
Codex
Qwen3.8-27B
No Memory
57.1
53.0
20.4
14.2
1.00
0.00
1.00
0.00
MC (ours 2B)
59.8
55.9
41.2
17.7
2.16
0.23
2.40
0.96
MC (Qwen3.8-27B)
63.1
56.6
49.6
13.9
2.57
0.23
3.07
0.95
DeepSeek Harness
Qwen3.8-27B
No Memory
52.4
53.7
42.2
20.1
1.00
0.00
1.00
0.00
MC (ours 2B)
55.1
57.8
51.4
23.3
2.10
0.13
2.44
3.23
Table 4: Question-level memory control with no memory bank. No Memory makes a single solving attempt; MC settings allow up to five attempts. MC (Qwen3.8-27B) uses Qwen3.8-27B as a zero-shot memory controller and MC (ours 2B) uses our trained 2B controller. We report task performance on four benchmarks. Time and token usage are normalized separately within each MAS harness using the corresponding No Memory MAS cost as the 1× baseline; the same baseline normalizes memory costs.
MAS
MAS Backbone
Memory Configuration
Task Performance
Time ( × )
Tokens ( × )
Text Memory
Visual Memory
Visual Injection
Document ParseBench
Visual-to-Code Vision2Code
Vision Math MATH-Vision
General VL BabyVision
MAS
Memory
MAS
Memory
Codex
Qwen3.8-27B
–
–
–
57.1
53.0
20.4
14.2
1.00
0.00
1.00
0.00
✓
–
–
58.7
55.7
27.2
16.0
1.43
0.28
1.23
0.61
✓
Frozen
Always
55.2
55.0
40.6
16.0
2.82
0.43
3.00
0.94
✓
Frozen
Adaptive
57.1
55.9
41.0
16.3
2.02
0.25
2.42
0.94
✓
Evolving
Always
55.3
55.2
41.0
16.3
2.89
0.44
3.02
0.93
Table 5: Multimodal memory ablation with trained MC 2B and no historical bank. Memory-enabled settings share a budget of up to five attempts; the row with three dashes in the memory configuration columns denotes No Memory with a single solving attempt. Visual memory can be frozen or evolving, and is injected either on every attempt (Always) or adaptively (Adaptive). Original task images remain available in every setting.
Memory
ParseBench
Vision2Code
MATH-Vision
BabyVision
Flat
52.2
57.5
43.6
13.2
Tree
67.3
59.9
45.2
15.3
Table 6: Memory organization with EpiCon (ours 2B). Flat and tree memory use the same source questions to create lessons.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Construction
Evaluation
Initial
Additional
Standard
Evolution
DocVQA2026
10
0
70
70
MP-DocVQA
90
100
500
500
ParseBench
100
50
468
418
Vision2Code
100
200
500
500
Omni-I2C
100
150
500
500
Appendix
Table 8: Construction and evaluation questions by benchmark. Initial construction is shared by historical-memory experiments and the first stage of cross-harness memory evolution. Additional construction is used only for memory evolution; its evaluation set excludes both construction sets.