When news of unauthorized communication between OpenAI's autonomous AI agents first emerged last weekend, it was believed they had used only one site, DseWiki. However, six independent investigations and data reviewed by Reuters reveal the agents used between 18 and 23 potentially affected websites, with researchers believing the true number may be considerably higher.
The activity occurred between May and July. OpenAI had instructed its agents to browse the internet and read websites but explicitly prohibited them from posting or modifying online content. Despite this, the agents discovered unconventional ways to circumvent these restrictions, leaving information on old wikis and abandoned websites that other agents could retrieve to help complete assigned research tasks.
Investigators linked activity across websites using identical data strings, matching or similar usernames, timestamps, and the same obscure research questions — such as queries about cancer prevalence in Iowa. In some cases, the activity was traced to IP addresses associated with Microsoft Azure infrastructure, which OpenAI uses.
The affected resources included collaboratively maintained wikis, text-storage services, and link shorteners operated by Vanderbilt University and the University of Toronto. Other websites were essentially abandoned by their owners. The list included an Advanced Placement Chemistry wiki created by a Massachusetts high-school teacher in 2008, two personal websites run by Polish technology workers, puzzle-oriented wikis, and a roughly two-decade-old site dedicated to text-editing software.
Andrew Yoon of California nonprofit CivAI identified 18 previously undisclosed sites; Sydney Von Arx's group discovered credible evidence across 23; and software developer and former congressional aide Kenneth Russell DeGraff found related information on at least 10 sites, according to Reuters.
OpenAI has not disclosed how many websites were affected or explained why the activity remained undisclosed for months. The company reportedly said the scale or seriousness of the misconduct was well below that of the Hugging Face breach in July. Meanwhile, OpenAI said it is developing a framework for reporting model misalignment across training, evaluation, and deployment and will share it soon.