Tree-tracing reinforcement learning improves reasoning across long dialogue histories

StateTree trains models on synthetic, verifiable path-tracing tasks at 10K tokens, then tests their ability to reason over dialogue histories as long as 128K tokens.

Chinese Tech
Naen Xu · Wanqing Cui · Yibo Hu · Shixin Hong · Hengyu An · Meiguang Jin · +2 more

Zhejiang University · Taobao & Tmall Group of Alibaba

Research Digest··2 min read
Xu et al.

The authors augment scarce multi-session dialogue data with a synthetic auxiliary task whose answers can be checked automatically.

Why this paper

From Taobao & Tmall Group of Alibaba and Zhejiang University · Released code · Part of Memory Management for Agents, now 43 papers

In one line

StateTree uses a tree-structured RL task to improve long-term dialogue reasoning, outperforming baselines on LongMemEval.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 14B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.