LLM agents learn when to use atomic tools versus composite skills via counterfactual outcome comparisons

CIPO mines executable skills from successful trajectories and trains granularity decisions by comparing paired outcomes of atomic and skill actions.

Chinese Tech
Yu Li · Yunlu Wan · Zijian Zhu · Han Luo · Chao Ren · Long-Fei Li · +1 more

Southeast University · KTH Royal Institute of Technology · Huawei Noah's Ark Lab

Research Digest··2 min read
The authors propose CIPO, a framework that enables LLM agents to adaptively select between atomic tools and composite skills.

CIPO consists of three stages: budget-constrained mining of recurring tool subsequences from successful trajectories using a BPE-style algorithm, instantiation of these subsequences as callable skills with parameter binding and execution tracing, and policy training via counterfactual imagination.

Why this paper

From Huawei Noah's Ark Lab and 2 others

In one line

CIPO improves LLM agents' tool use by training them to choose between atomic tools and composite skills based on state.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.