cs.AIOct 5, 2026

Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving

Authors: Jiaqi Zhao, Haodong Chen, Jitai Hao, Wei Zhao, Jinghao Pang, Qiang Huang, Jun Yu

Organizations: Harbin Institute of Technology (Shenzhen)

Abstract

LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not systematically connect these views, limiting workflow-aware execution. HEAR, a bidirectional Harness--Engine Pairing protocol for agentic LLM serving. HEAR standardizes how the harness communicates workflow intent and execution requirements and how the engine returns runtime state, capabilities, and outcomes. By separating protocol semantics from optimization policies, HEAR supports diverse coordination strategies without changing workflow or model semantics. We instantiate HEAR for online cache-aware runtime coordination and workload-aware execution-mode selection for agent roles. Across four conversational and research-agent benchmarks under memory-constrained, concurrent serving, HEAR achieves a 1.61×1.61\times batch speedup and reduces median time-to-first-token by 2.23×2.23\times on SCBench. Mooncake shows that workflow intent and live engine state provide complementary benefits across load regimes. On BrowseComp-Plus and DeepResearchBench, workload-specific configurations yield 1.23×1.23\times and 2.45×2.45\times end-to-end speedups, respectively, without observed task-quality degradation. These results establish HEAR as a reusable coordination substrate for efficient agentic LLM serving.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving

    Apr 28, 2026Shan Yu, Junyi Shu, Yuanjiang Ni +14WorkloadCache

  2. ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System

    Feb 14, 2026Hao Kang, Ziyang Li, Weili Xu +7Agentic InferenceLLM Inference Optimization

  3. Harnesses for Inference-Time Alignment over Execution Trajectories

    May 15, 2026Boyuan Wang, Bochao Li, Minghan Wang +2Agent HarnessAgentic Inference