The researchers replayed successful coding-agent trajectories, running the benchmark tests after every code-changing action.
Successful Agent Traces Predict Which Base Models Merit Post-Training
Three verifier-based probes ranked ten base checkpoints similarly to their post-trained descendants on SWE-bench Verified.
Big Tech
Tan Yu · Alexander Bukharin · Khushi Bhardwaj · Jennifer Williams · Zirui Liu · Jonathan Lingjie Li · +16 more
NVIDIA · University of Minnesota – Twin Cities · University of California, Berkeley
Research Digest··3 min read
The authors propose evaluating base models at the decisive point in successful coding-agent trajectories, avoiding the tool-use failures that make end-to-end agent benchmarks nearly unusable before post-training.
Why this paper
From NVIDIA and 2 others
In one line
Base model potential for agentic coding can be predicted by probing at the decisive step in successful post-trained trajectories.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§