Validation-gated co-evolution improves agent models and execution harnesses

Alternating reinforcement learning with validated harness revisions outperformed weight-only training and ungated updates on two agent benchmarks.

Chinese Tech
Jiexing Qi · Yu He · Jun Liu · Qichen Huang · Shaohua Hu · Zhan Dang · +6 more

ICT AI Competence Center, Huawei Technologies Co., Ltd. · Shanghai Jiao Tong University

Research Digest··2 min read
Qi and colleagues introduce VACE, a training loop that alternates updates to a language model’s weights with revisions to its harness, the instructions, skills and execution rules surrounding the model.

VACE alternates agentic reinforcement learning with trajectory-driven harness refinement.

Why this paper

From ICT AI Competence Center, Huawei Technologies Co., Ltd. and Shanghai Jiao Tong University

In one line

Alternating model training with checkpoint-specific, validation-gated harness updates outperforms weight-only training and ungated alternation on two agent benchmarks.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.