Post-training helps an agent sustain complex work for over 10 hours

Marathoner combines synthesized software tasks, filtered teacher trajectories and sandboxed reinforcement learning to train agents for unusually long executions.

Chinese Tech
Zhang Ruiyang · Ou Jinpeng · Xie Yifan · Zhou Jingang · Pan Lirui · Guo Qingpei · +1 more

FIC, University of Macau · Ant Group · Peking University

Research Digest··2 min read
Zhang et al.

The authors constructed difficult software tasks from major GitHub release pull requests containing more than 1,000 lines of new code.

Why this paper

From Ant Group and 2 others

In one line

Marathoner achieves ultra-long-horizon autonomous execution with 1000+ tool calls over 10+ hours, beating proprietary models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.