Search agents improve by deciding when to compress and reconsider

Traverse combines answer criteria, independent verification and agent-controlled memory with a reinforcement-learning method designed to stabilize context management.

Chinese Tech
Jingyuan Ma · Lynx Aster · He Zhang · Siyao Song · Weijie Yuan · Zhe Zhang · +2 more

State Key Laboratory of Multimedia Information Processing · Peking University · ByteDance

Research Digest··3 min read
Ma et al.

Traverse organizes long web searches into three states: Rubric, Answer and Verify.

Why this paper

From ByteDance and 2 others

In one line

A three-state search agent with self-managed memory and final-segment-only reinforcement learning improves long-horizon web search and reaches 72.83 on BrowseComp.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 35B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.