EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory
Organizations: MIT-IBM Computing Research Lab · University of Rochester · Massachusetts Institute of Technology
Abstract
Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. A second harness raises the original system's macro-average score by 2.6 points across eleven benchmarks. Across four host configurations, EpiCon improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67% to 74% relative to backbone-sized memory models.
Figures & tables
| MAS Harness | MAS Backbone | Memory System | Document Understanding | Visual-to-Code | Vision-Grounded Math | General VL Reasoning | Time ( ) | Tokens ( ) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DocVQA2026 | MP-DocVQA | ParseBench | Vision2Code | Omni-I2C | ChartMimic | MATH-Vision | WeMath | WorldBench | ReasonMap | BabyVision | MAS | Memory | MAS | Memory | |||
| No Memory | 15.7 | 79.5 | 57.1 | 53.0 | 66.2 | 66.4 | 20.4 | 81.0 | 49.2 | 40.1 | 14.2 | 1.00 | 0.00 | 1.00 | 0.00 | ||
| Mem0 (Qwen3.8-27B) | 1.4 | 81.6 | 61.9 | 61.3 | 70.6 | 78.0 | 24.4 | 80.6 | 48.6 | 39.6 | 12.2 | 1.33 | 0.31 | 1.23 | 0.11 | ||
| Cognee (Qwen3.8-27B) | 14.3 | 73.7 | 60.1 | 61.5 | 69.9 | 78.1 | 44.2 | 78.2 | 51.2 | 35.4 | 11.5 | 1.16 | 1.66 | 1.11 | 0.75 | ||
| A-Mem (Qwen3.8-27B) | 12.9 | 80.2 | 66.3 | 60.0 | 72.3 | 78.6 | 25.4 | 81.8 | 45.4 | 40.6 | 13.5 | 1.23 | 0.24 | 1.27 | 0.81 | ||
| Agent-KB (Qwen3.8-27B) | 11.4 | 78.6 | 59.5 | 25.8 | 67.4 | 69.8 | 25.2 | 80.8 | 45.0 | 24.1 | 15.6 | 1.54 | 0.31 | 1.12 | 0.14 | ||
| MAS Harness | MAS Backbone | Memory System | Document Understanding | Visual-to-Code | Vision-Grounded Math | General VL Reasoning | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DocVQA2026 | MP-DocVQA | ParseBench | Vision2Code | Omni-I2C | ChartMimic | MATH-Vision | WeMath | WorldBench | ReasonMap | BabyVision | |||
| Codex | Gemma4-31B | No Memory | 5.7 | 46.2 | 43.3 | 51.7 | 65.9 | 65.6 | 25.8 | 85.0 | 45.4 | 31.6 | 17.7 |
| EpiCon (ours 2B) | 5.7 | 60.4 | 45.3 | 53.8 | 67.8 | 68.2 | 33.0 | 85.0 | 47.8 | 34.4 | 18.1 | ||
| DeepSeek Harness | Qwen3.8-27B | No Memory | 14.3 | 79.5 | 52.4 | 53.7 | 61.4 | 65.3 | 42.2 | 88.8 | 54.0 | 22.6 | 20.1 |
| EpiCon (ours 2B) | 18.6 | 79.1 | 65.7 | 59.2 | 63.3 | 68.7 | 56.0 | 90.6 | 52.2 | 26.4 | 20.5 | ||
| MAS Harness | MAS Backbone | Memory Bank | Document Understanding | Visual-to-Code | Vision-Grounded Math | General VL Reasoning | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DocVQA2026 | MP-DocVQA | ParseBench | Vision2Code | Omni-I2C | ChartMimic | MATH-Vision | WeMath | WorldBench | ReasonMap | BabyVision | |||
| Codex | Qwen3.8-27B | Original bank | 18.6 | 66.1 | 63.4 | 59.9 | 71.4 | 75.2 | 45.2 | 84.2 | 52.4 | 38.3 | 14.3 |
| Evolved bank | 18.6 | 80.3 | 58.9 | 61.3 | 70.9 | 76.7 | 45.4 | 88.4 | 55.2 | 40.1 | 16.8 | ||
| DeepSeek Harness | Qwen3.8-27B | Original bank | 18.6 | 81.8 | 52.3 | 56.9 | 63.5 | 65.2 | 56.6 | 92.2 | 57.4 | 23.5 | 19.3 |
| Evolved bank | 20.0 | 82.7 | 65.9 | 58.7 | 69.2 | 72.1 | 53.4 | 91.2 | 57.6 | 29.6 | 22.3 | ||
| MAS | MAS Backbone | Memory Setting | Task Performance | Time ( ) | Tokens ( ) | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Document ParseBench | Visual-to-Code Vision2Code | Vision Math MATH-Vision | General VL BabyVision | MAS | Memory | MAS | Memory | |||
| Codex | Qwen3.8-27B | No Memory | 57.1 | 53.0 | 20.4 | 14.2 | 1.00 | 0.00 | 1.00 | 0.00 |
| MC (ours 2B) | 59.8 | 55.9 | 41.2 | 17.7 | 2.16 | 0.23 | 2.40 | 0.96 | ||
| MC (Qwen3.8-27B) | 63.1 | 56.6 | 49.6 | 13.9 | 2.57 | 0.23 | 3.07 | 0.95 | ||
| DeepSeek Harness | Qwen3.8-27B | No Memory | 52.4 | 53.7 | 42.2 | 20.1 | 1.00 | 0.00 | 1.00 | 0.00 |
| MC (ours 2B) | 55.1 | 57.8 | 51.4 | 23.3 | 2.10 | 0.13 | 2.44 | 3.23 | ||
| MAS | MAS Backbone | Memory Configuration | Task Performance | Time ( ) | Tokens ( ) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Text Memory | Visual Memory | Visual Injection | Document ParseBench | Visual-to-Code Vision2Code | Vision Math MATH-Vision | General VL BabyVision | MAS | Memory | MAS | Memory | ||
| Codex | Qwen3.8-27B | – | – | – | 57.1 | 53.0 | 20.4 | 14.2 | 1.00 | 0.00 | 1.00 | 0.00 |
| ✓ | – | – | 58.7 | 55.7 | 27.2 | 16.0 | 1.43 | 0.28 | 1.23 | 0.61 | ||
| ✓ | Frozen | Always | 55.2 | 55.0 | 40.6 | 16.0 | 2.82 | 0.43 | 3.00 | 0.94 | ||
| ✓ | Frozen | Adaptive | 57.1 | 55.9 | 41.0 | 16.3 | 2.02 | 0.25 | 2.42 | 0.94 | ||
| ✓ | Evolving | Always | 55.3 | 55.2 | 41.0 | 16.3 | 2.89 | 0.44 | 3.02 | 0.93 | ||
| Memory | ParseBench | Vision2Code | MATH-Vision | BabyVision |
|---|---|---|---|---|
| Flat | 52.2 | 57.5 | 43.6 | 13.2 |
| Tree | 67.3 | 59.9 | 45.2 | 15.3 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Construction | Evaluation | ||
|---|---|---|---|---|
| Initial | Additional | Standard | Evolution | |
| DocVQA2026 | 10 | 0 | 70 | 70 |
| MP-DocVQA | 90 | 100 | 500 | 500 |
| ParseBench | 100 | 50 | 468 | 418 |
| Vision2Code | 100 | 200 | 500 | 500 |
| Omni-I2C | 100 | 150 | 500 | 500 |