7 Code, which activates 32 billion parameters per inference.
Reinforcement learning helps coding agents finish tasks across benchmarks
Training on 1,700 expert-built tasks improved one open-weight model across six external coding benchmarks and three agent harnesses.
Industry
Sushant Mehta · Logan Ritchie · Edwin Chen
Surge AI
Research Digest··2 min read
7 Code using reinforcement learning with verifiable rewards, without a supervised warm-up.
Why this paper
From Surge AI
In one line
RL on 1,700 expert coding tasks improves a 1T-parameter model on six benchmarks and shortens agent steps.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (6 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§