3, Kimi K3 and DeepSeek V4 Pro, with four configurable harnesses: OpenHands, DeepSeek Harness, PI and openJiuwen.
Agent performance depends on matching models, harnesses and tasks
Tests across three command-line benchmarks show that model rankings, costs and failure recovery vary substantially with the surrounding agent harness.
Top University
Yixuan Li · Yiyun Zhou · Yao Long Teng · Fuchao Yang · Yanchen Deng · Zhiyi Lyu · +3 more
Nanyang Technological University
Research Digest··2 min read
Li and colleagues evaluate language models together with the software harnesses that expose tools, manage context and return errors.
Why this paper
From Nanyang Technological University · Released code and data
In one line
A model's performance and ranking depend on both the harness and the task.
What it released
CodeData
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ✓Dataset link in the paper (huggingface.co)
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§