Mined agent skills need domain-specific evidence and format choices

A controlled study of skill mining from production traces shows that workflow and ontology representations, and the availability of outcome labels, have varying effects on downstream task success across enterprise benchmarks.

Big Tech
Yue Ran Kang · Colton Mikolajczyk · Chhaya Methani · Hazel Mak · Sahil Bhatnagar · Susheel Suresh · +1 more

Massachusetts Institute of Technology · Microsoft Corporation

Research Digest··2 min read
The authors compare six combinations of mining evidence (successful traces only, successes and failures with labels, or the same mix without labels) and skill form (ordered workflow plans or declarative ontologies) while holding the mining pipeline fixed.

The authors study how skill-mining evidence and representation affect downstream agent performance in enterprise settings.

Why this paper

From Microsoft Corporation and Massachusetts Institute of Technology

In one line

Mined skill performance depends on task domain; workflows beat ontologies on one benchmark, outcome labels help on another, but no configuration wins everywhere.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.