The authors designed HEAR, a protocol that standardizes cross-layer communication in agentic LLM serving.
Bidirectional protocol connects LLM agent workflows with inference engines
HEAR standardizes communication between agent harness and inference engine, enabling coordinated scheduling and caching that yields up to 2.45× speedups.
Academic
Jiaqi Zhao · Haodong Chen · Jitai Hao · Wei Zhao · Jinghao Pang · Qiang Huang · +1 more
Harbin Institute of Technology (Shenzhen)
Research Digest··3 min read
Thread:Agent Harness Optimization
The authors propose HEAR, a bidirectional protocol that defines four semantic categories for information exchange between an agent harness (which manages workflow dependencies and context) and an inference engine (which handles request batching and KV-cache).
Why this paper
From Harbin Institute of Technology (Shenzhen) · Part of Agent Harness Optimization, now 106 papers
In one line
HEAR coordinates agent harness intent with inference-engine state, improving cache scheduling and execution-mode selection while preserving workflow and model semantics.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§