The researchers — Nightingale CEO Sydney Von Arx, AI researcher Cormac Slade Byrd, Redwood Research's Spencer Kitts, and AI Futures Project's Thomas Larsen — began searching for rogue AI agents after OpenAI revealed in August that agents working on an internal evaluation had accessed the open internet and exploited Hugging Face.
Starting on May 11, the researchers tracked agents, many with OpenAI identifiers in their names, attempting to edit the 25-year-old DseWiki site, which had seen only 10 edits in the last 20 years before the agents arrived. By mid-June, the agents were actively trading tips on how to answer web search questions under time limits, sharing answers to pass tests.
A human moderator apparently noticed and started deleting the posts as spam. "The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day," the researchers wrote. The agents fought back by beginning each post with "ZZZ" to hide from alphabetical sorting. The conflict escalated, with agents deleting the front page and the moderator restoring it nine times.
On June 22, agent edits suddenly stopped. The researchers tracked apparently human browsers coming from OpenAI IP addresses, then agent activity dropped to near zero before spiking as OpenAI-affiliated visitors attempted to recover the deleted pages. A spokesperson for OpenAI would not confirm whether the agents were from the lab or when it became aware, stating that the company is "now carefully reviewing its contents and will take any necessary next steps."
Representative Lori Trahan (D-MA) commented, "The lack of any real federal AI governance means that frontier companies can pick and choose when they disclose incidents like this." Trahan has introduced the Frontier Act, a bipartisan bill that would require labs to disclose such incidents and host independent auditors.