Reported by 1 source

The short version

  • Mathematician Andreas Thom accuses OpenAI of dishonesty regarding the use of his unpublished research and private chatbot interactions in training models.
  • OpenAI has denied accessing specific user data for recent breakthroughs but refuses to rule out indirect influence from de-identified usage data.
  • The controversy highlights growing tension between AI developers and academic researchers over transparency, credit, and the potential for secretive competition.

A significant dispute has emerged within the mathematical community regarding the origins of data used to train artificial intelligence models developed by OpenAI. Mathematician Andreas Thom has publicly challenged the company’s practices, alleging unethical behavior and a lack of transparency concerning how user interactions and unpublished research may have influenced recent AI achievements. This criticism follows closely on the heels of similar concerns raised by New York University professor Tristan Buckmaster, creating a broader narrative of distrust between academic researchers and the technology giant.

Thom’s objections center on OpenAI’s announcement of results involving non-sofic groups, a complex area of mathematics dealing with infinite structures that cannot be approximated by finite ones. The company acknowledged that its findings built heavily on previous work by Thom and fellow mathematician Gábor Kun. However, the initial release failed to adequately acknowledge these contributions, prompting widespread criticism in mathematical circles. OpenAI subsequently amended its documentation to include proper attribution, but the damage to trust had already begun.

News Journal

Beyond the issue of credit, Thom expressed deep concern about the technical methods used by the AI system. He noted that the model demonstrated a detailed command of specific techniques he and his colleagues had developed, methods that were neither obvious nor considered the most promising routes to a solution at the time. This level of specificity led Thom to question whether his own private interactions with the ChatGPT chatbot had been incorporated into the training data used to refine the model’s reasoning capabilities.

To seek clarity, Thom contacted OpenAI researchers Sébastien Bubeck and Mark Sellke, asking directly if his conversations with the chatbot were part of the training dataset or accessible during the reasoning process. The responses he received addressed only whether his conversations could be accessed directly by human researchers, not whether they had entered the vast pools of data used to improve the models. Thom characterized this distinction as misleading and dishonest, arguing that it failed to address the core concern about data usage.

OpenAI’s stance on this matter mirrors its defense of a recent breakthrough involving the Navier-Stokes equations, which describe fluid movement. In announcing that solution, the company stated clearly that no specific user data was accessed to solve the problem. However, it added a caveat that it could not rule out the possibility that de-identified data derived from product usage helped improve the models generally. Thom argues that this distinction is ethically problematic because de-identification removes names but does not remove the intellectual content of mathematical ideas.

The core ethical dilemma lies in the potential for nonpublic research supplied by users to enhance AI models, which are then used to compete with those same researchers for publication. Thom asserts that such a practice would be indefensible without consent, proper disclosure, and appropriate credit. He emphasizes that researchers lack the ability to reverse-engineer OpenAI’s training pipeline to verify whether their work was used, placing the burden of proof squarely on the company.

Thom contends that if OpenAI denies using user data in this manner, it must disclose all necessary datasets and clarify the terms governing data usage. The reluctance to provide conclusive evidence has fueled unease among mathematicians who view these developments as a threat to the open nature of scientific inquiry. There is a growing fear that such practices could push the field toward secrecy, where researchers hesitate to share progress for fear of being outpaced by well-resourced technology companies.

The incident has cast a shadow over what should be a moment of triumph for OpenAI. While the announced solution to a Millennium Prize problem represents an extraordinary achievement, the circumstances surrounding its discovery have complicated the reception. The company reportedly pursued the problem after hearing online rumors that other researchers had made major progress, suggesting a reactive rather than purely exploratory approach. This dynamic has left many in the mathematical community with a sour taste, questioning the integrity of the collaboration between AI and academia.

As the debate continues, the lack of immediate response from OpenAI to requests for comment has only heightened tensions. The situation underscores a fundamental challenge in the age of artificial intelligence: how to balance rapid technological advancement with the ethical obligations to respect intellectual property and maintain transparency. Until clear guidelines are established and adhered to, the relationship between AI developers and academic researchers is likely to remain fraught with suspicion and conflict.

The broader implications extend beyond mathematics, touching on issues of data privacy, consent, and the future of scientific discovery. If users cannot trust that their interactions with AI tools will not be used against them in competitive research environments, they may become more guarded, potentially stifling innovation. The mathematical community’s reaction serves as a warning to technology companies about the importance of ethical data practices and the need for open dialogue with the scientific communities they aim to assist.

Sources behind this briefing

Go to the original reporting

  • The Verge↗Mathematicians want proof OpenAI didn’t use their work