CIPO consists of three stages: budget-constrained mining of recurring tool subsequences from successful trajectories using a BPE-style algorithm, instantiation of these subsequences as callable skills with parameter binding and execution tracing, and policy training via counterfactual imagination.
LLM agents learn when to use atomic tools versus composite skills via counterfactual outcome comparisons
CIPO mines executable skills from successful trajectories and trains granularity decisions by comparing paired outcomes of atomic and skill actions.
Chinese Tech
Yu Li · Yunlu Wan · Zijian Zhu · Han Luo · Chao Ren · Long-Fei Li · +1 more
Southeast University · KTH Royal Institute of Technology · Huawei Noah's Ark Lab
Research Digest··2 min read
The authors propose CIPO, a framework that enables LLM agents to adaptively select between atomic tools and composite skills.
Why this paper
From Huawei Noah's Ark Lab and 2 others
In one line
CIPO improves LLM agents' tool use by training them to choose between atomic tools and composite skills based on state.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§