cpp, vLLM, and SGLang, examining whether each stack accepted the request and returned calls as text or through a native structured tool_calls channel.
Serving stacks can distort local language-model tool-use evaluations
Tests across four inference stacks show that request handling, output protocols, decoding constraints, and aggregation choices can substantially alter measured tool-call fidelity.
Industry
Lijuan Tang · Yuemeng Zheng
Northeastern University, Seattle
Research Digest··2 min read
Thread:Agent Harness Optimization
Tang and Zheng examine the protocol step in which a coding model must produce a parseable tool call before an agent harness can act.
Why this paper
From Northeastern University, Seattle · Part of Agent Harness Optimization, now 61 papers
In one line
Tool-use evaluation outcomes can depend on the serving stack, not just the model, causing misclassification of failures as model non-calls.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§