According to the case study, two mechanisms carried the work. The first was a quality signal: the CodeHealth MCP Server gave agents a deterministic score to optimize and to judge whether a transformation had helped. The second was correctness: a replay-trace harness compared the rollback state hash frame by frame, so behavior could be checked after every change.
Rather than applying a fixed catalogue, the agents accumulated a refactoring playbook, ending with 22 recipes and 82 supporting notes. Familiar transformations appear, including Extract Function and Guard Clauses, but so do recipes specific to this codebase. Shared Index Range captures repeated loops differing only in start and end ranges. Action Parameter handles duplicated control structures differing mainly in which function they invoke. Uniform Step Table converts heterogeneous calls into table-driven dispatch. Failed attempts were recorded too, including transformations that made Code Health worse.
Model choice mattered. The team settled on Claude Opus for the bulk of the work, reporting that Claude Code with Opus was significantly better than Codex with Sol at capturing and documenting the emerging patterns. Files often plateaued when smaller models ran the task, appearing to reach a local optimum they could not move past.
The InfoQ article reporting the case study is titled "Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves", a headline that captures some of the open questions the result has generated.