The row over OpenAI's training data has escalated, with mathematician Andreas Thom becoming the second researcher in recent days to accuse the company of dishonest and unethical behaviour. In a series of posts on the decentralized social network Mastodon, Thom raised concerns that conversations his team had with ChatGPT prior to OpenAI's public announcement of a mathematical breakthrough may have been incorporated into the model without permission or acknowledgment.
Thom's criticism follows a bitter dispute with another researcher who also claimed OpenAI used unpublished work. The company has not yet publicly responded to either set of allegations, and it remains unclear exactly how OpenAI sources its training data for mathematics-related tasks. Thom did not provide direct evidence that his interactions were used, but his posts call for greater transparency from the AI developer about the origins of its data.
The controversy highlights growing unease among academics about how major AI laboratories train their models, particularly when those models produce results that appear to build directly on unpublished or private research.