Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving
Organizations: Harbin Institute of Technology (Shenzhen)
Abstract
LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not systematically connect these views, limiting workflow-aware execution. HEAR, a bidirectional Harness--Engine Pairing protocol for agentic LLM serving. HEAR standardizes how the harness communicates workflow intent and execution requirements and how the engine returns runtime state, capabilities, and outcomes. By separating protocol semantics from optimization policies, HEAR supports diverse coordination strategies without changing workflow or model semantics. We instantiate HEAR for online cache-aware runtime coordination and workload-aware execution-mode selection for agent roles. Across four conversational and research-agent benchmarks under memory-constrained, concurrent serving, HEAR achieves a batch speedup and reduces median time-to-first-token by on SCBench. Mooncake shows that workflow intent and live engine state provide complementary benefits across load regimes. On BrowseComp-Plus and DeepResearchBench, workload-specific configurations yield and end-to-end speedups, respectively, without observed task-quality degradation. These results establish HEAR as a reusable coordination substrate for efficient agentic LLM serving.
Figures & tables
| Category | Representative Contents |
| Harness Engine | |
| Execution Description and Intent | Agent roles and stages; workflow dependencies and readiness; request sizes and output limits; context lifecycles; known future uses and explicitly marked predictions. |
| Execution Requirements and Control | Scheduling preferences and waiting-time constraints; cache preparation, retention, and release requests; allowed or requested execution configurations and serving targets. |
| Engine Harness | |
| State and Capabilities | Reusable prefix ranges and residency tiers; allocatable capacity, queue state, and resource pressure; supported controls and configurations; calibrated or estimated execution costs. |
| Execution Outcomes | Status updates for inference and control requests (accepted, completed, rejected, unsupported, or failed); actual configuration, permitted fallback, and cache reuse; latency, resource use, retries, and errors. |
| Distinction | Protocol Interpretation |
| Intent vs. Control | Descriptions expose current or anticipated execution conditions; only explicit control requests ask the receiver to act. For example, reporting that context will be reused is distinct from requesting that its KV cache be retained or prepared. |
| Preference vs. Requirement | Preferences are best-effort, whereas requirements constrain valid execution. An unsatisfied requirement must be rejected or handled by an explicitly permitted fallback. |
| For example, a priority hint may affect request ordering, but it cannot override a mandatory protection condition. | |
| Observation vs. Guarantee | State reports describe observed conditions, and capability reports describe supported operations. Neither reserves resources nor guarantees that a future request will be admitted. |
| For example, GPU-residency feedback may guide scheduling without reserving the prefix. | |
| Acceptance vs. Completion | Acceptance means that an operation has beenadmitted for processing; completion confirms that its stated effect has been realized. Rejected, unsupported, and failed operations remain distinct outcomes. |
| Configuration | TTFT | Batch Completion | Reuse | ||
| P50 | P95 | MAX | |||
| FCFS | 63.1 1.00× | 74.2 1.00× | 79.0 1.00× | 481 1.00× | 19.2 |
| Cache-Aware | 4.6 13.6× | 75.8 0.98× | 93.9 0.84× | 224 2.15× | 87.6 |
| Cache-Aware + Guard-40 | 28.3 2.23× | 51.1 1.45× | 65.0 1.22× | 299 1.61× | 66.1 |
| Cache-Aware + Guard-60 | 4.5 13.9× | 64.0 1.16× | 80.4 0.98× | 230 2.09× | 84.7 |
| Load | Configuration | TTFT | Session | Throughput | Reuse | ||
| P50 | P95 | MAX | |||||
| 50% | FCFS | 0.7 1.00 | 2.8 1.00 | 6.2 1.00 | 13.7 1.00 | 86.1 1.00 | 19.5 |
| Cache-Aware | 0.6 1.21 | 2.4 1.17 | 5.9 | 13.3 1.02 | 86.0 1.00 | 20.2 | |
| Cache-Aware + Guard | 0.6 1.10 | 2.8 1.01 | 7.4 0.84 | 13.6 1.01 | 86.2 1.00 | 19.6 | |
| Session-Aware | 0.6 1.16 | 2.5 1.14 | 6.3 0.99 | 13.0 1.05 | 86.6 1.01 | 31.9 | |
| Session-Aware + Cache-Aware + Guard | 0.5 | 2.3 | 6.1 1.02 | 12.7 | 86.8 | 32.1 | |
| Main | Reader | Browscomp-plus | Deep-research-bench | ||||
| Wall (h) | P95 E2E (min) | Acc. (%) | Makesp (h) | P95 E2E (min) | RACE (%) | ||
| vLLM | vLLM | 5.69_{\color[rgb]{0.5,0.5,0.5}1.00\times} | 40.0_{\color[rgb]{0.5,0.5,0.5}1.00\times} | 46.15 | 5.93_{\color[rgb]{0.5,0.5,0.5}1.00\times} | 39.8_{\color[rgb]{0.5,0.5,0.5}1.00\times} | 40.6 |
| OmniKV | vLLM | 5.37_{\color[rgb]{0.5,0.5,0.5}1.06\times} | 40.3_{\color[rgb]{0.5,0.5,0.5}0.99\times} | 47.12 | 6.17_{\color[rgb]{0.5,0.5,0.5}0.96\times} | 39.9_{\color[rgb]{0.5,0.5,0.5}1.00\times} | 41.1 |
| vLLM | \mathbf{4.62}_{\color[rgb]{0.3,0.3,0.3}\mathbf{1.23\times}} | 39.6_{\color[rgb]{0.5,0.5,0.5}1.01\times} | 46.63 | 2.76_{\color[rgb]{0.5,0.5,0.5}2.15\times} | 15.6_{\color[rgb]{0.5,0.5,0.5}2.55\times} | 40.5 | |
| OmniKV | 4.98_{\color[rgb]{0.5,0.5,0.5}1.14\times} | 37.2_{\color[rgb]{0.5,0.5,0.5}1.08\times} | 43.27 | \mathbf{2.42}_{\color[rgb]{0.3,0.3,0.3}\mathbf{2.45\times}} | 15.8_{\color[rgb]{0.5,0.5,0.5}2.52\times} | 41.1 | |
| vLLM | SnapKV | 4.83_{\color[rgb]{0.5,0.5,0.5}1.18\times} | 40.4_{\color[rgb]{0.5,0.5,0.5}0.99\times} | 44.23 | 2.85_{\color[rgb]{0.5,0.5,0.5}2.08\times} | 17.7_{\color[rgb]{0.5,0.5,0.5}2.25\times} | 40.0 |
| Benchmark | Main Output Tok./Req. | Main Latency Alt. Selected | Reader/Sub Latency Dense Selected | Cache Evidence |
| BrowseComp-Plus | K | OmniKV vLLM s 1.14× | vLLM H 2 O s 5.60× | Prefill Reuse |
| DeepResearchBench | K | vLLM OmniKV s 1.08× | vLLM H 2 O s 7.22× | Chain reuse |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Frequency | Additional Online Work | Included in Formal E2E Time |
| KV-state update | Per cache event | Residency-directory metadata update | Yes |
| Scheduling decision | Per dispatch point | Prefix lookup, priority construction, and Guard check | Yes |
| Role-to-mode lookup | Per model request | Lookup in the frozen role mapping | Yes |
| Outcome reporting | Per inference/control outcome | Status, actual mode, reuse, retry, and error metadata | Yes |
| Profile calibration | Offline development stage | Candidate profiling, filtering, ranking, and verification | No |
| Mode | Cache and Token Settings | Execution Capacity | Mode-specific Parameters |
| Main Role | |||
| vLLM | Prefix caching Max. context 202,752; 32,768 batched tokens | 16 sequences | Dense execution |
| OmniKV | Radix cache 32,768-token prefill chunks | 4 batch / 4 decode / 16 GPU sequences | Full-attention layers: automatic; Retention: 8 sink / 256 recent / 4,096 decode tokens; Prefill workspace: 8 GiB |
| Reader Role | |||
| vLLM | Prefix caching Max. context 202,752; 16,384 batched tokens | 16 sequences | Dense execution |
| H 2 O | Chain cache 16,384-token prefill chunks | 12 batch / 12 decode / 36 GPU sequences | Budgets: 16,384 prefill / 8,192 decode tokens; Recent ratio: 0.2; Scoring: 128-token window with logits-based prefill scoring |
| Mode | Cache and Token Settings | Execution Capacity | Mode-specific Parameters |
| Main Role | |||
| vLLM | Prefix caching Max. context 200,000; 32,768 batched tokens | 10 sequences | Dense execution |
| OmniKV | Radix cache 32,768-token prefill chunks | 10 batch / 10 decode / 16 GPU sequences | Full-attention layers: |
| Sub-Agent Role | |||
| vLLM | Prefix caching Max. context 200,000; 32,768 batched tokens | 10 sequences | Dense execution |
| H 2 O | Chain cache 32,768-token prefill chunks | 10 batch / 10 decode / 44 GPU sequences | Budgets: 8,192 prefill / 4,096 decode tokens; Recent ratio: 0.2; Scoring: 128-token window with logits-based prefill scoring |
| Main | Reader | Efficiency and Latency | Quality and Completion Outcomes | ||||||||||
| Wall (h) | Speedup | Throughput | Median | P95 | Max | Acc. (%) | Retr. Recall (%) | C | I | U | F | ||
| vLLM | vLLM | 5.69 | 36.57 | 14.63 | 40.00 | 41.32 | 46.15 | 69.22 | 96 | 46 | 54 | 12 | |
| H 2 O | 4.62 | 45.00 | 10.98 | 39.58 | 41.81 | 46.63 | 70.17 | 97 | 59 | 45 | 7 | ||
| SnapKV | 4.83 | 43.03 | 10.75 | 40.38 | 43.26 | 44.23 | 70.56 | 92 | 63 | 40 | 13 | ||
| OmniKV | vLLM | 5.37 | 38.72 | 13.67 | 40.28 | 46.54 | 47.12 | 68.75 | 98 | 42 | 53 | 15 | |
| H 2 O | 4.98 | 41.73 | 12.64 | 37.24 | 42.37 | 43.27 | 68.33 | 90 | 61 | 49 | 8 | ||
| Main | Sub-Agent | Efficiency and Latency | Quality and Completion Outcomes | |||||||||||||
| Makespan | Speedup | Throughput | E2E Latency (min) | RACE (%) | Comp. | Insight | Instr. | Read. | Errors | Empty | ||||||
| Mean | Median | P95 | Max | |||||||||||||
| vLLM | vLLM | 5.93 | 16.7 | 29.46 | 30.64 | 39.81 | 41.18 | 97 | 40.65 | 39.93 | 38.68 | 42.98 | 42.91 | 1 | 4 | |
| 2.76 | 36.3 | 13.17 | 13.17 | 15.64 | 21.27 | 96 | 40.46 | 39.65 | 38.68 | 42.81 | 42.57 | 0 | 4 | |||
| SnapKV | 2.85 | 35.1 | 15.13 | 15.19 | 17.70 | 20.26 | 99 | 40.01 | 39.33 | 37.78 | 42.51 | 42.45 | 0 | 1 | ||
| OmniKV | vLLM | 6.17 | 16.2 | 30.41 | 30.64 | 39.93 | 40.64 | 98 | 41.14 | 40.30 | 39.26 | 43.57 | 43.20 | 0 | 2 | |
| Benchmark | Role | Req./Item | Input/Req. (K) | Output/Req. (K) | Workflow Pattern |
| BrowseComp-Plus | Main | 3.0 | 52.5 | 0.36 | Dynamic coordination and answer synthesis |
| Reader | 13.3 | 59.1 | 0.36 | Question-dependent document reads | |
| DeepResearchBench | Main | 5.4 | 32.5 | 4.42 | Round-synchronized review and synthesis |
| Sub-Agent | 12.0 | 135.3 | 0.51 | Four parallel agents over three fixed rounds |
| Benchmark | Role | Alternative Selected | Mean Latency (s) | Gain | Execution Evidence |
| BrowseComp-Plus | Main | OmniKV vLLM | Selected Main output: 0.36K tokens/request | ||
| Reader | vLLM H 2 O | Prefill ; reuse | |||
| DeepResearchBench | Main | vLLM OmniKV | Selected Main output: 4.42K tokens/request | ||
| Sub-Agent | vLLM H 2 O | Chain reuse: |
| Main | Reader | Outcome Counts | Answer Quality | ||||
| Correct | Incorrect | Unresolved | Failure | Acc. (%) | 95% CI | ||
| vLLM | vLLM | 96 | 46 | 54 | 12 | 46.15 | |
| OmniKV | vLLM | 98 | 42 | 53 | 15 | 47.12 | |
| vLLM | 97 | 59 | 45 | 7 | 46.63 | ||
| OmniKV | 90 | 61 | 49 | 8 | 43.27 | ||
| vLLM | SnapKV | 92 | 63 | 40 | 13 | 44.23 | |
| Main | Reader | Answers Produced | Evidence Quality (%) | ||
| Retrieval Recall | Citation Precision | Citation Recall | |||
| vLLM | vLLM | 142 | 69.22 | 78.01 | 30.82 |
| OmniKV | vLLM | 140 | 68.75 | 77.71 | 33.90 |
| vLLM | 156 | 70.17 | 75.44 | 28.62 | |
| OmniKV | 151 | 68.33 | 74.78 | 27.87 | |
| vLLM | SnapKV | 155 | 70.56 | 72.63 | 28.10 |
| Main | Sub-Agent | Evaluation Coverage | RACE (%) | ||||||
| Judged | Empty | Task Errors | Overall | Comp. | Insight | Instr Follow. | Read. | ||
| vLLM | vLLM | 97 | 4 | 1 | 40.65 | 39.93 | 38.68 | 42.98 | 42.91 |
| OmniKV | vLLM | 98 | 2 | 0 | 41.14 | 40.30 | 39.26 | 43.57 | 43.20 |
| vLLM | 96 | 4 | 0 | 40.46 | 39.65 | 38.68 | 42.81 | 42.57 | |
| OmniKV | 99 | 1 | 0 | 41.05 | 39.91 | 39.59 | 43.26 | 43.10 | |
| vLLM | SnapKV | 99 | 1 | 0 | 40.01 | 39.33 | 37.78 | 42.51 | 42.45 |
| Benchmark | Selected | Paired | Quality Diff. (%) | 95% Bootstrap CI (%) | Paired Test |
| BrowseComp-Plus | vLLM–H 2 O | 208 | McNemar | ||
| DeepResearchBench | OmniKV–H 2 O | 89 | Paired -test |