Workspace-grounded synthesis scales verifiable training for long-horizon work agents

WorkForge generated 16,700 realistic workspaces whose tasks and evaluators were derived from facts found directly in their files.

Chinese Tech
Jiazheng Zhang · Long Ma · Yunxian Yang · Zhiheng Xi · Zhikai Lei · Yajie Yang · +20 more

Tencent · Fudan University

Research Digest··2 min read
Zhang and colleagues introduce WorkForge, a pipeline for creating training environments in which agents must navigate multiple files and produce professional deliverables over extended interactions.

WorkForge begins with expert-defined workflow blueprints describing the resources, decisions and outputs involved in professional tasks.

Why this paper

From Tencent and Fudan University

In one line

WorkForge builds verifiable work-agent environments from real-world files by anchoring tasks and verifiers in observable workspace facts, improving agent performance across benchmarks.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (3 noted)
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.