The authors represent each sampled agent trajectory using task-specific descriptors chosen by the practitioner.
Explicit diversity training gives LLM agents distinct successful strategies
TJPO rewards task-specific variation across groups of trajectories while preserving competitive success rates.
Big Tech
Huaiyu Fu · Heng Cao · Hao Wang · Jian Ya · Tao Chen
Microsoft · tuyoogame
Research Digest··2 min read
Thread:RL for Tool Agents
Fu et al.
Why this paper
From Microsoft and tuyoogame · Part of RL for Tool Agents, now 13 papers
In one line
Trajectory-guided Joint Policy Optimization (TJPO) improves task-relevant behavioral diversity in LLM agents.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§