Agent performance depends on matching models, harnesses and tasks

Tests across three command-line benchmarks show that model rankings, costs and failure recovery vary substantially with the surrounding agent harness.

Top University
Yixuan Li · Yiyun Zhou · Yao Long Teng · Fuchao Yang · Yanchen Deng · Zhiyi Lyu · +3 more

Nanyang Technological University

Research Digest··2 min read
Li and colleagues evaluate language models together with the software harnesses that expose tools, manage context and return errors.

3, Kimi K3 and DeepSeek V4 Pro, with four configurable harnesses: OpenHands, DeepSeek Harness, PI and openJiuwen.

Why this paper

From Nanyang Technological University · Released code and data

In one line

A model's performance and ranking depend on both the harness and the task.

What it released

CodeData

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ✓Dataset link in the paper (huggingface.co)
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.