AI Agents Refactor 300,000-Line Codebase in Three Weeks, Case Study Shows

CodeScene reports 2,903 commits and a Code Health jump from 5.6 to 10.0 at a token cost of roughly $4,000

By LineZotpaper
Published
Read Time2 min
CodeScene has published a case study in which coding agents refactored a 300,000-line C codebase in three weeks at a token cost of roughly $4,000. The work produced 2,903 commits across 726 files, modified 252,055 lines and improved the codebase's Code Health score from 5.6 to 10.0. The codebase is an open-source decompilation of Street Fighter III: 3rd Strike, and Adam Tornhill, CodeScene's founder, said it was the first time he had seen what he called "superhuman AI performance at scale".

According to the case study, two mechanisms carried the work. The first was a quality signal: the CodeHealth MCP Server gave agents a deterministic score to optimize and to judge whether a transformation had helped. The second was correctness: a replay-trace harness compared the rollback state hash frame by frame, so behavior could be checked after every change.

Rather than applying a fixed catalogue, the agents accumulated a refactoring playbook, ending with 22 recipes and 82 supporting notes. Familiar transformations appear, including Extract Function and Guard Clauses, but so do recipes specific to this codebase. Shared Index Range captures repeated loops differing only in start and end ranges. Action Parameter handles duplicated control structures differing mainly in which function they invoke. Uniform Step Table converts heterogeneous calls into table-driven dispatch. Failed attempts were recorded too, including transformations that made Code Health worse.

Model choice mattered. The team settled on Claude Opus for the bulk of the work, reporting that Claude Code with Opus was significantly better than Codex with Sol at capturing and documenting the emerging patterns. Files often plateaued when smaller models ran the task, appearing to reach a local optimum they could not move past.

The InfoQ article reporting the case study is titled "Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves", a headline that captures some of the open questions the result has generated.

§

Analysis

Why This Matters

  • The case study provides a concrete data point on the speed and token cost of AI agents for large-scale refactoring, a task traditionally done by teams of engineers over months.
  • The combination of a deterministic quality signal and a verification harness offers a possible template for delegating risky code changes to agents.
  • The results also highlight the importance of model choice and raise questions about how these methods transfer to codebases without such thorough test coverage.

Background

Refactoring legacy code is one of the harder tasks in software maintenance because changes must preserve behavior while improving structure. AI coding assistants have evolved from suggesting snippets to agentic systems that can plan and execute multi-step changes. CodeScene, which maintains the Code Health metric, applied agentic refactoring to an open-source decompilation of a classic arcade game, which includes a replay harness that can verify behavior frame by frame. That harness is a key advantage for safe automation and is not typical of most production code.

Key Perspectives

CodeScene founder Adam Tornhill: Describes the result as "superhuman AI performance at scale" and highlights the agents' ability to discover and document codebase-specific refactoring recipes. Practitioners: The InfoQ article's headline indicates some practitioners are asking what the case study proves, a sign that the result, while striking, may not be taken as a general template without further evidence. Critics: May note that the token cost excludes the development of the harness and MCP server, and that the reliance on a frontier model like Claude Opus may not be available to all teams.

What to Watch

  • Whether CodeScene or others replicate the approach on production codebases with existing test suites.
  • Additional comparisons between model families for agentic refactoring, especially where smaller models plateau.
  • Adoption of MCP-based quality signals and replay harnesses as standard components of AI coding workflows.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.