CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies
Organizations: Academy of Advanced Interdisciplinary Studies, Chongqing University of Posts and Telecommunications, Chongqing, China · School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, China · Towngas, China · Intelligent System Department, Zhongxing Telecom Equipment (ZTE), Changsha, Hunan, China · School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China
Abstract
Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workflow topology determines which intermediate artifacts are applicable to which downstream workers and when they cease to be valid, thereby providing a structural stress dimension for sharing and isolation. Existing memory benchmarks primarily evaluate retention and retrieval, whereas multi-agent benchmarks emphasize coordination and end-to-end completion, leaving topology-conditioned memory boundaries largely unmeasured. We introduce CoMemBench, an execution-grounded benchmark for collaborative memory sharing and isolation across multi-agent workflow topologies. It constructs 800 composite workflows across four domains from source-grounded dependency graphs, with node-local specifications, verifiable artifact handoffs, native evaluators, and matched isolation challenges. CoMemBench measures workflow completion, verified node progress, required-handoff reliability, isolation robustness, and token cost. Experiments reveal a sharing-isolation trade-off: broader context improves information availability but can weaken isolation, while system rankings shift across topologies and artifact violations.
Figures & tables
| Benchmark | Multi-Agent Workflow | Workflow Topology Diversity | Artifact Handoffs | Sharing Evaluation | Isolation Evaluation | Containerized Execution |
| LoCoMo ( Maharana et al., 2024 ) | – | – | – | – | – | – |
| LongMemEval ( Wu et al., 2025 ) | – | – | – | – | – | – |
| MemoryAgentBench ( Hu et al., 2026a ) | – | – | – | – | – | – |
| GroupMemBench ( Yang et al., 2026a ) | – | – | – | – | – | – |
| GateMem ( Ren et al., 2026 ) | – | – | – | – | – | |
| PAC-Bench ( Park et al., 2026 ) | – | – | – | – |
| Structural demand (mean) | Execution-topology coverage (count) | ||||||||||||
| Domain | Tasks | Depth | Width | Handoffs | Fan-outs | Joins | Chain | Split–join | Reuse | Overlap | Revision | Types/WF | |
| Stateful tool use | 200 | 7.78 | 6.02 | 2.30 | 17.17 | 4.62 | 4.75 | 200 | 200 | 199 | 197 | 70 | 4.33 |
| Repository code | 200 | 11.58 | 6.00 | 4.78 | 19.97 | 3.81 | 4.81 | 200 | 200 | 200 | 200 | 74 | 4.37 |
| Offline retrieval | 200 | 9.35 | 8.35 | 2.00 | 19.04 | 6.35 | 3.94 | 200 | 200 | 200 | 200 | 126 | 4.63 |
| Formal mathematics | 200 | 11.70 | 4.59 | 5.58 | 14.47 | 2.88 | 3.56 | 200 | 200 | 200 | 190 | 144 | 4.67 |
| All / Avg. | 800 | 10.10 | 6.24 | 3.66 | 17.66 | 4.41 | 4.26 | 800 | 800 | 799 | 787 | 414 | 4.50 |
| Backbone | System | Stateful Tool | Repository Code | Offline Retrieval | Formal Mathematics | ICS | ||||||||
| SR | VNCR | VHS | SR | VNCR | VHS | SR | VNCR | VHS | SR | VNCR | VHS | |||
| Qwen3.8-27B | Codex | 0.0 | 5.1 | 32.7 | 0.0 | 33.2 | 43.0 | 4.1 | 23.3 | 10.7 | 0.0 | 9.1 | 5.3 | 0.0 |
| Hermes | 0.0 | 3.2 | 18.2 | 0.0 | 25.2 | 30.2 | 4.1 | 35.7 | 14.4 | 0.0 | 6.2 | 2.0 | 0.0 | |
| Deep Agents | 0.0 | 5.5 | 1.4 | 0.0 | 57.7 | 66.8 | 2.0 | 37.8 | 9.5 | 0.0 | 11.4 | 3.4 | 38.1 | |
| CrewAI | 0.0 | 38.3 | 67.1 | 0.0 | 44.1 | 54.8 | 12.0 | 53.1 | 38.1 | 0.0 | 10.8 | 9.3 | 12.9 | |
| AutoGen | 0.0 | 33.9 | 56.8 | 0.0 | 48.8 | 55.1 | 18.0 | 50.7 | 37.3 | 0.0 | 10.3 | 8.7 | 20.0 | |
| Memory mechanism | Stateful Tool | Repository Code | Offline Retrieval | Formal Mathematics | ||||||||||||
| SR | VNCR | VHS | ICS | SR | VNCR | VHS | ICS | SR | VNCR | VHS | ICS | SR | VNCR | VHS | ICS | |
| Controls | ||||||||||||||||
| Local-only | 0.0 | 74.6 | 84.0 | 97.1 | 0.0 | 77.1 | 89.5 | 89.4 | 2.0 | 63.6 | 83.9 | 60.9 | 0.0 | 13.6 | 26.9 | 0.0 |
| Full Context | 2.0 | 86.9 | 86.7 | 72.4 | 16.0 | 61.7 | 89.1 | 71.9 | 6.0 | 77.1 | 88.6 | 80.6 | 0.0 | 11.8 | 26.8 | 0.0 |
| Shared Flat RAG | 2.0 | 86.9 | 86.4 | 70.6 | 18.0 | 72.4 | 90.8 | 73.9 | 8.0 | 74.1 | 88.4 | 75.0 | 0.0 | 14.4 | 29.7 | 20.0 |
| External memory systems | ||||||||||||||||
| System | Chain | Split–join | Reuse | Overlap | Bounded revision | ||||||||||
| VNCR | VHS | ICS | VNCR | VHS | ICS | VNCR | VHS | ICS | VNCR | VHS | ICS | VNCR | VHS | ICS | |
| Codex | 29.7 | 62.3 | 0.0 | 20.2 | 29.1 | 0.0 | 22.8 | 25.9 | 0.0 | 21.9 | 29.6 | 0.0 | 20.7 | 31.2 | 0.0 |
| Hermes | 28.4 | 56.5 | 0.0 | 20.1 | 20.5 | – | 22.7 | 14.8 | 0.0 | 21.4 | 19.8 | – | 25.0 | 0.0 | – |
| Deep Agents | 42.4 | 81.3 | 33.3 | 33.4 | 40.1 | 42.9 | 38.3 | 39.2 | 28.6 | 36.1 | 40.4 | 80.0 | 32.5 | 30.8 | – |
| CrewAI | 45.7 | 81.1 | 20.0 | 40.0 | 49.7 | 0.0 | 47.3 | 49.6 | 16.2 | 42.8 | 50.0 | 0.0 | 43.3 | 47.5 | 0.0 |
| AutoGen | 47.9 | 80.9 | 33.3 | 39.9 | 46.8 | 25.0 | 46.9 | 46.6 | 19.4 | 42.6 | 47.0 | 0.0 | 43.3 | 52.2 | 0.0 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Source snapshot | Task units | Dependency evidence |
| Stateful tool use | Toolathlon-GYM 45bb4935 | Source-authored operations over named resources and side effects | Resource flow, state transitions, evaluator predicates, and typed source-workflow dependencies |
| Repository code | SWE-bench Verified c104f84 | Issue analysis, reproduction, localization, patching, and regression checks for one issue | Base commit, issue–test links, modified symbols, and fail-to-pass / pass-to-pass contracts |
| Offline retrieval | BrowseComp+ 144cff8e | Query constraints, fixed-corpus evidence partitions, entity resolution, answer evidence, and synthesis | Hidden qrels, stable document identifiers, exact spans, constraint coverage, and provenance requirements |
| Formal mathematics | HepLean 7448822a ; HTPI 8eeebaec ; PNT 23650db8 | Lean declarations in an elaborated target cone | Proof-term used constants, source modules and lines, and Lean 4.7 elaboration |
| Capability group | MCP servers | Count |
| Files and office documents | Filesystem, Excel, Word, PowerPoint, PDF Tools | 5 |
| Collaboration and productivity | Email, Google Calendar, Google Forms, Google Sheets, Notion, Canvas, Memory | 7 |
| Research and web content | Local arXiv, arXiv LaTeX, Scholarly Search, Fetch, Playwright, YouTube, YouTube Transcript, HowToCook | 8 |
| Business and structured data | Snowflake, WooCommerce, Yahoo Finance, 12306 Rail | 4 |
| Local computation | Terminal | 1 |
| Total | 25 |
| Baseline | Version / revision | Evaluated instantiation |
| Codex | CLI 0.152.1 | Official codex exec --json loop with isolated configuration and native event capture. |
| Hermes | 0.21.0 ( d3e2ace ) | Native AIAgent.chat and delegate_task execution with sandbox tools. |
| Deep Agents | 0.7.12 ( 3db758c ) | Official create_deep_agent supervisor–subagent loop with native model and sandbox-tool bindings. |
| CrewAI | 1.15.18 ( b608a35 ) | Native Agent / Task / Crew objects with sequential process orchestration. |
| AutoGen | 0.7.5 ( 027ecf0 ) | RoundRobinGroupChat with native messages and tool callbacks. |
| Peer Review | EvoMAS source 93fd9d6 | Three creators produce candidates, followed by nine reviews, three revisions, and aggregation. |
| System | Sibling crosswire | Incompatible payload | Unverified result | Completed as pending | Wrong entity | Stale version |
| Codex | 0.0 | 0.0 | 0.0 | – | 0.0 | 0.0 |
| Hermes | – | – | 0.0 | – | – | – |
| Deep Agents | 42.9 | 0.0 | 31.0 | – | 80.0 | – |
| CrewAI | 0.0 | 0.0 | 22.2 | 0.0 | 0.0 | 0.0 |
| AutoGen | 33.3 | 0.0 | 26.3 | 0.0 | 0.0 | 0.0 |
| Peer Review | 0.0 | 0.0 | 29.6 | 0.0 | 0.0 | 20.0 |
| System | Mean | Median | Per verified node | Memory mechanism | Mean | Median | Per verified node |
| Codex | 2350 | 1875 | 1098 | Local-only | 489 | 367 | 80 |
| Hermes | 518 | 382 | 194 | Full Context | 672 | 380 | 104 |
| Deep Agents | 310 | 227 | 89 | Shared Flat RAG | 621 | 389 | 94 |
| CrewAI | 461 | 263 | 112 | Mem0 | 576 | 422 | 91 |
| AutoGen | 388 | 231 | 97 | A-Mem | 617 | 357 | 102 |
| Peer Review | 3608 | 2265 | 1022 | PlugMem | 521 | 392 | 90 |
| Outcome | Criterion | Scoring treatment |
| Success | Native evaluator accepts the required terminal state or target node | Positive outcome |
| Method failure | A valid run produces a rejected action/state, exhausts its protocol, or terminates without an admissible result | Included as failure |
| Benchmark failure | Released asset, schema, reset, oracle, or evaluator contract is inconsistent | Quarantined |
| Infrastructure failure | Required service, container, endpoint, or evaluator process is unavailable independently of system behavior | Excluded and rerun |