The authors augment scarce multi-session dialogue data with a synthetic auxiliary task whose answers can be checked automatically.
Tree-tracing reinforcement learning improves reasoning across long dialogue histories
StateTree trains models on synthetic, verifiable path-tracing tasks at 10K tokens, then tests their ability to reason over dialogue histories as long as 128K tokens.
Chinese Tech
Naen Xu · Wanqing Cui · Yibo Hu · Shixin Hong · Hengyu An · Meiguang Jin · +2 more
Zhejiang University · Taobao & Tmall Group of Alibaba
Research Digest··2 min read
Thread:Memory Management for Agents
Xu et al.
Why this paper
From Taobao & Tmall Group of Alibaba and Zhejiang University · Released code · Part of Memory Management for Agents, now 43 papers
In one line
StateTree uses a tree-structured RL task to improve long-term dialogue reasoning, outperforming baselines on LongMemEval.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 14B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§