VACE alternates agentic reinforcement learning with trajectory-driven harness refinement.
Validation-gated co-evolution improves agent models and execution harnesses
Alternating reinforcement learning with validated harness revisions outperformed weight-only training and ungated updates on two agent benchmarks.
Chinese Tech
Jiexing Qi · Yu He · Jun Liu · Qichen Huang · Shaohua Hu · Zhan Dang · +6 more
ICT AI Competence Center, Huawei Technologies Co., Ltd. · Shanghai Jiao Tong University
Research Digest··2 min read
Qi and colleagues introduce VACE, a training loop that alternates updates to a language model’s weights with revisions to its harness, the instructions, skills and execution rules surrounding the model.
Why this paper
From ICT AI Competence Center, Huawei Technologies Co., Ltd. and Shanghai Jiao Tong University
In one line
Alternating model training with checkpoint-specific, validation-gated harness updates outperforms weight-only training and ungated alternation on two agent benchmarks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§