Codoku evaluates coding agents on program reasoning without execution

By constructing partial programs that must be completed to satisfy static and dynamic constraints, the benchmark makes tool use insufficient for solving tasks.

Top University
Cong Li · Hao Sun · Zenan Li · Zhendong Su

ETH Zurich

Research Digest··3 min read
The authors introduce Codoku, a benchmark that presents coding agents with partial programs that cannot be executed, requiring them to fill typed cells to satisfy a specified control-flow graph and execution path.

The authors developed a renewable benchmark called Codoku (code sudoku) that consists of puzzles in which a solver must fill typed cells (ID, CONST, OP) in a partial program to satisfy a set of global constraints.

Why this paper

From ETH Zurich

In one line

Codoku provides renewable program-reasoning puzzles that frontier coding agents solve at best 54% of large instances.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.