Serving stacks can distort local language-model tool-use evaluations

Tests across four inference stacks show that request handling, output protocols, decoding constraints, and aggregation choices can substantially alter measured tool-call fidelity.

Industry
Lijuan Tang · Yuemeng Zheng

Northeastern University, Seattle

Research Digest··2 min read
Tang and Zheng examine the protocol step in which a coding model must produce a parseable tool call before an agent harness can act.

cpp, vLLM, and SGLang, examining whether each stack accepted the request and returned calls as text or through a native structured tool_calls channel.

Why this paper

From Northeastern University, Seattle · Part of Agent Harness Optimization, now 61 papers

In one line

Tool-use evaluation outcomes can depend on the serving stack, not just the model, causing misclassification of failures as model non-calls.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.