The authors generated 173,561 multi-turn coding trajectories from Qwen3-8B on programming problems, each with successive code revisions and execution feedback.
Detecting reward hacking by supervising attempted shortcuts, not just successful exploits.
HackTrace reads a coding agent's internal states during code generation to spot shortcut attempts with high accuracy and minimal latency.
Top University
Hao Jiang · Xin Li · Annan Wang · Yichi Zhang · Weisi Lin
Nanyang Technological University · Independent Researcher
Research Digest··3 min read
The authors introduce HackTrace, a monitor that detects reward hacking in coding agents by supervising attempted shortcut behavior independently of whether the exploit succeeded.
Why this paper
From Nanyang Technological University and Independent Researcher · Released code
In one line
Monitoring internal states during code generation detects attempted reward hacking with high accuracy and low overhead, cutting cheating in RL from 82-91% to 1-5%.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§