A Second Mathematician Accuses OpenAI of 'Dishonesty' Over Its AI Math Breakthroughs

Andreas Thom says OpenAI never explained whether his ChatGPT conversations shaped the very result that built on his own research, and its answers were 'materially misleading.'

AI2Day Newsdesk4 min read
A futuristic AI model represented visually as interconnected nodes and pathways, symbolizing complex data processing
Share

Key points

  • Mathematician Andreas Thom publicly accused OpenAI of dishonesty in May 2025 over its use of researcher data to train its AI models.
  • OpenAI acknowledged its non-sofic groups result, one of ten recent math breakthroughs, built heavily on prior work by Thom and fellow mathematician Gábor Kun.
  • OpenAI quietly updated its writeup after criticism that it failed to credit Thom and Kun's recent contributions.
  • OpenAI admits it "cannot rule out" that anonymised user data indirectly improved its models, but denies accessing specific conversations to solve the problems.
  • Researchers warn the episode could push mathematics into greater secrecy if even rumours of progress trigger a race with well-funded tech companies.

A second mathematician has publicly accused OpenAI of being dishonest about where its AI systems get their ideas, deepening a controversy over the company's recent string of claimed mathematical discoveries.

Andreas Thom, a mathematician at TU Dresden, posted a series of messages on the social network Mastodon saying that conversations he and colleagues had with ChatGPT, OpenAI's widely used AI chatbot, may have contributed to a breakthrough the company announced last month. The result in question concerns what are called non-sofic groups: infinite mathematical structures that behave in ways too complex to be captured by any finite approximation. Thom is one of the world's leading experts on precisely this topic.

What did OpenAI actually say?

OpenAI acknowledged the result built heavily on earlier work by Thom and by mathematician Gábor Kun, but only after first publishing a writeup that omitted that credit. The company quietly amended it following criticism from the mathematical community.

Thom emailed OpenAI researchers Sébastien Bubeck and Mark Sellke, also a statistician at Harvard University, asking whether his ChatGPT conversations had entered OpenAI's training data, the vast collection of text and information used to teach its models. The reply addressed only whether his specific conversations could be accessed directly. It did not say whether the content of those conversations had been absorbed into training data in an anonymised form.

"No such qualification, explanation, or evidence was given," Thom wrote. "I take this as dishonesty to say the least."

His concern mirrors one raised days earlier by Tristan Buckmaster, a mathematics professor at New York University, who questioned whether OpenAI's models had benefited from his own use of Codex, an OpenAI coding tool, while working on a separate prize-winning problem. First reported by The Verge AI, that dispute centres on OpenAI's claimed solution to a Navier-Stokes problem, a famous unsolved question about how fluids move.

OpenAI's public statement on that case drew a careful line: it denied accessing any specific user data to solve the problem, but admitted it "cannot rule out that de-identified data," meaning data stripped of names and other identifying details, may have indirectly improved its models.

Thom finds that distinction unconvincing. "De-identification may remove a name; it does not remove the intellectual content of a mathematical idea," he wrote.

What does this mean for researchers and ordinary users?

For working scientists, the implications are direct. If sharing ideas with an AI tool can later help that same tool race you to a discovery, without your knowledge or consent, the rational response is to stop sharing. Several researchers told The Verge AI they worry the field will grow more secretive as a result.

For everyone else, the case is a useful reminder of something easy to overlook: conversations with AI chatbots can, in principle, feed back into the training of future AI systems. Most services offer settings to opt out of this, and it is worth checking those settings if you work with sensitive or original ideas.

"Only OpenAI has the relevant data," Thom wrote. The burden of proof, he argues, belongs with them.

OpenAI had not responded to requests for comment at the time of publication.

Common questions

Can my conversations with ChatGPT be used to train future AI models?

OpenAI's default settings have historically allowed user conversations to be used for training, though names and other personal details may be stripped out first. You can opt out through ChatGPT's data controls in the settings menu.

Is OpenAI's non-sofic groups result actually verified?

OpenAI announced the result last month, but independent mathematical verification of results this complex takes time. As of publication, no full peer-reviewed confirmation has been published.

Why does credit matter in mathematics?

Priority and credit in research determine careers, funding and professional reputation. A company announcing a result without acknowledging the researchers whose work made it possible is, in academic terms, a serious breach of norms.

© 2026 AI2Day