Raven: The Harness of Harnesses for Composable Agentic Intelligence
Organizations: EverMind AI
Abstract
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
Figures & tables
| Group | Checks |
|---|---|
| Format | Arguments parse as complete JSON. Fields have the required types and values. No unknown field is present. Identifiers and instance handles match their patterns. Required node fields are nonblank. |
| Graph structure | Identifiers are unique within the graph and unused in the session. Dependencies name a node of the graph or a completed earlier node with output. Every output placeholder is backed by a dependency. Input values are well formed, and declared and referenced inputs agree. Path forms apply only to file or node inputs. File and reference paths stay within the permitted roots. The graph is acyclic. |
| Agent capability | Shared instances use stateful agents. Skill overrides appear only at the head of an instance chain. Path forms are not sent to agents without local-file access. |
| Agent status | Each named agent is registered and enabled. |
| Environment | The call does not come from inside a sub-agent run. Delegation is not paused, and the dispatch budget is available. The user approves the graph when confirmation is requested. |
| Actor | Agent Libraries | ||
|---|---|---|---|
| Host | read, write | none | read and write every mapped library |
| Mapped worker | read via prefetch | none | its own verdicts via prefetch |
| Unmapped agent | none | none | none |
| Domains per task | Subsets covered | Tasks |
|---|---|---|
| 2 | 6 / 6 | 66 |
| 3 | 4 / 4 | 47 |
| 4 | 1 / 1 | 27 |
| Total | 11 / 11 | 140 |
| System | Accuracy | Input (M) | Output (k) | Cost (USD) |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | ||||
| MiroFlow | 47.7 | 1.493 | 11.3 | – |
| DeepSeek-Harness | 49.7 | 4.587 | 16.2 | – |
| Raven-Research | 56.3 | 1.706 | 30.0 | – |
| Qwen3.5-397B-A17B | ||||
| MiroFlow | 48.0 | 1.799 | 15.3 | – |
| Case | Deliverable | Task Graph | Design Time (min:s) |
|---|---|---|---|
| (a) Song dynasty | Deck (Chinese) | 18:45 | |
| (b) Greek polychromy | Deck (English) | 55:34 | |
| (c) Pop music | Deck (English) | 47:03 | |
| (d) Abstract art | Deck (English) | 64:50 | |
| (e) Orchestration frameworks | Comparison board | 6:57 | |
| (f) Light pollution | Poster series | 2:58, 13:57 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Execution policy, underlying model, and harness. | |
| Permitted histories, actions, and probability distributions. | |
| Host controller, its nondelegating executor, and the pool including . | |
| A generic system, either the composed system or a standalone agent, and the composed system . | |
| Task, request, initial law, response kernel, and objective correctness. | |
| Reference conditions, their read-only state projection, full state, immutable-record ledger, and final record. |
| Symbol | Meaning |
|---|---|
| Submitted graph and its analytical expansion. | |
| Prompt template, dependency identifiers, and slot bindings. | |
| Invocation capability overrides and stateful-instance handle. | |
| Scheduler-event index and node status at that event. | |
| Rendered task, task plus memory paths, and final dispatched input. | |
| Advertised ancestor-memory paths and a node’s reserved memory-file location. |
| Symbol | Meaning |
|---|---|
| Candidate identity and executable components. | |
| Attempts per task, survivor cap, and round cap. | |
| Archive cell and the candidate’s assigned descriptor. | |
| Fixed evaluation protocol and paired-gain screening threshold. | |
| Editable and protected harness locations. | |
| Executable harness patch. |
| Benchmark | Initial | Evolved | Gain |
|---|---|---|---|
| Terminal-Bench-2 | 36.1 | 45.4 | +9.3 |
| LiveCodeBench | 58.1 | 71.8 | +13.7 |
| Omni-MATH | 54.3 | 66.0 | +11.7 |
| BrowseComp+ | 16.9 | 30.8 | +13.9 |
| GDPval | 43.7 | 52.9 | +9.2 |
| AppWorld | 41.3 | 56.7 | +15.4 |
| SkillsBench | GDPval | QwenClawBench | |||||||
| Cell | None | Skills | Gain | None | Skills | Gain | None | Skills | Gain |
| Raven Qwen3.5-27B | 10.0 | 16.5 | +6.5 | 82.6 | 83.8 | +1.2 | 66.9 | 70.8 | +3.9 |
| Raven Qwen3.5-397B-A17B | 9.2 | 22.6 | +13.4 | 84.0 | 85.2 | +1.2 | 68.8 | 73.2 | +4.4 |
| OpenClaw Qwen3.5-27B | 8.8 | 13.0 | +4.2 | 81.2 | 83.1 | +1.9 | 65.2 | 66.7 | +1.5 |
| OpenClaw Qwen3.5-397B-A17B | 11.1 | 16.9 | +5.8 | 82.2 | 84.0 | +1.8 | 65.7 | 67.0 | +1.3 |
| Pooled gain | |||||||||
| Catalog | Retrieval Stack | Pass@1 | Gain |
|---|---|---|---|
| No skills | 9.2 | – | |
| Curated | Fine-tuned | 22.6 | +13.4 |
| Curated | Off-the-shelf Qwen3 | 13.8 | +4.6 |
| Raw crawl | Fine-tuned | 14.9 | +5.7 |