The authors constructed difficult software tasks from major GitHub release pull requests containing more than 1,000 lines of new code.
Post-training helps an agent sustain complex work for over 10 hours
Marathoner combines synthesized software tasks, filtered teacher trajectories and sandboxed reinforcement learning to train agents for unusually long executions.
Chinese Tech
Zhang Ruiyang · Ou Jinpeng · Xie Yifan · Zhou Jingang · Pan Lirui · Guo Qingpei · +1 more
FIC, University of Macau · Ant Group · Peking University
Research Digest··2 min read
Zhang et al.
Why this paper
From Ant Group and 2 others
In one line
Marathoner achieves ultra-long-horizon autonomous execution with 1000+ tool calls over 10+ hours, beating proprietary models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§