OpenAI Accused of Stealing Mathematicians' Research: A Deep Dive

Two mathematicians allege OpenAI used their unpublished research drafts to train new models, sparking debate over AI data ownership.
Two mathematicians published a joint statement claiming that after feeding a year's worth of research drafts into OpenAI Codex, OpenAI released new models containing nearly identical solutions — and refused to confirm whether private user data had been used for training. The incident has sparked broad discussion about AI training data ownership, platform transparency, and user intellectual property rights. The article examines the legal gray areas around AI data use, advocates for local deployment of open-source models like Llama and Mistral to mitigate data exposure risks, and calls on the AI industry to adopt clearer data disclosures, optional privacy modes, and independent auditing mechanisms.
Two mathematicians have accused OpenAI of using their private conversation data — without authorization — to train new models. The allegation has reignited intense debate around AI training data sources and privacy protection. This incident touches not only on academic ethics, but on a far more fundamental question: who owns the data users generate when working with AI tools?

What Happened: A Year of Work Allegedly Hijacked
According to a joint statement published by the two mathematicians, they spent an entire year solving a challenging mathematical problem, feeding each draft iteration into OpenAI's Codex model for assistance along the way. Then, just days before they planned to publish their findings, OpenAI released new models called Sol and Astra — containing solutions nearly identical to their own.
What made the situation even more troubling was OpenAI's response when the mathematicians asked whether their private conversation data had been used to train these new models: the company dodged the question entirely, refusing to give a direct answer. That evasiveness only deepened public skepticism about how transparently AI companies handle user data.
The full statement has been published on the NYU Courant Institute of Mathematical Sciences website (https://cims.nyu.edu/~tristanb/statement.pdf), detailing the full timeline of events and supporting evidence.
OpenAI Codex is a code generation model fine-tuned on GPT-3's architecture, released in 2021. It can understand natural language and translate it into code, and also supports mathematical reasoning tasks. When users interact with online AI tools like Codex, their prompts and conversation history are typically transmitted via HTTPS to cloud servers for processing, and may be retained under the platform's terms of service. OpenAI's enterprise API terms and its consumer-facing ChatGPT terms differ in how they handle data retention and training rights — but neither explicitly prohibits using user interaction data for model improvement. Sol and Astra, mentioned in the statement as subsequently released math-enhanced models from OpenAI, have not been fully confirmed by independent sources, and the complete technical details of this incident remain to be verified.
The Core Dispute: Who Owns What Users Put Into AI?
At the heart of this controversy is a fundamental question: when users process their own creative work or research using AI tools, who actually owns that data?
From a legal and ethical standpoint, original content submitted to a model by a researcher should be protected by intellectual property law. Yet most AI service agreements include data usage clauses that permit the platform to use user interaction data for model improvement. These clauses are often buried in lengthy legal documents that the average user is unlikely to read or fully understand.
This incident seems to confirm a troubling trend: major AI companies may operate under the implicit assumption that "everything you accomplish using our models fundamentally belongs to us." If that logic holds, the consequences for academic research, commercial innovation, and individual creativity could be severe.
From an intellectual property law perspective, original text or research drafts submitted by users to an AI system remain the copyright of their creator. However, whether using such data to train a model constitutes infringement is an area where no clear legal precedent yet exists in most jurisdictions. In the United States, the Digital Millennium Copyright Act (DMCA) primarily addresses the copying and distribution of copyrighted content — the legality of using copyrighted material for machine learning training remains actively debated among legal scholars and in the courts. The EU's General Data Protection Regulation (GDPR) offers stronger protection from a personal data processing standpoint, requiring companies to obtain explicit consent before using personal data for new purposes. This lag between law and technology means users who attempt to assert their rights often face a dual obstacle: difficulty proving their case and ambiguity over which jurisdiction applies.
Local AI Deployment: An Effective Solution for Data Privacy
This controversy once again highlights the importance of running large language models (LLMs) on local hardware. For users handling sensitive information, proprietary research, or trade secrets, uploading data to cloud-based services always carries a risk of exposure.
Core Advantages of Local Deployment
Complete data control: All interactions stay on the local device, fundamentally eliminating any possibility of data leakage.
Absolute privacy protection: No risk of your data being used for model training or any other purpose — your research remains entirely your own.
Regulatory compliance: For industries subject to sector-specific regulation (such as healthcare or finance), local deployment makes it significantly easier to meet compliance requirements.
With the rapid advancement of open-source models, locally deployable high-performance options like Llama and Mistral are becoming the go-to choice for privacy-conscious users. While these models may trail closed-source solutions like GPT-4 in certain capabilities, their advantages in data security are unmatched.
Local LLM deployment is typically achieved using inference frameworks such as Ollama, LM Studio, or llama.cpp. Users can run quantized models ranging from 7B to 70B parameters on consumer-grade GPUs with sufficient VRAM (such as the NVIDIA RTX 4090) or on high-memory CPU setups. Quantization compresses model weights from 16-bit floating point to 4-bit or 8-bit integers, dramatically reducing VRAM requirements while preserving most inference capability — making local deployment increasingly practical. Open-source models including Llama 3 (Meta), Mistral, and Qwen 2.5 have achieved performance on coding and mathematical reasoning benchmarks that approaches early GPT-4 versions, meeting the demands of most professional research use cases. The core security advantage of local deployment is that inference runs entirely within the device's memory — no data is transmitted over the network, physically eliminating the risk of interception or platform misuse.
How the AI Industry Can Rebuild User Trust
This incident should prompt the entire AI industry to reassess its data usage policies. Building genuine user trust requires:
Clear data usage disclosures: Explain in plain language how user data will be used — not buried in complex legal boilerplate.
Optional privacy modes: Offer a "strict privacy" setting with a genuine commitment not to use user data for training.
Independent audit mechanisms: Allow third-party organizations to verify AI companies' actual data usage practices.
Stronger legal protections: Legislative frameworks that clearly define users' rights over their AI interaction data.
How Researchers and Creators Can Protect Themselves
For researchers and creators using commercial AI services, a higher level of vigilance is now essential:
- Prioritize local models when working with unpublished original research
- Carefully read the data usage clauses in any service agreement before using the platform
- For critical creative work, avoid inputting complete materials into AI systems before formal publication
AI Data Privacy Protection Can't Wait
Regardless of how this specific case ultimately resolves, OpenAI's alleged theft of research findings has sounded a clear warning. In an era of rapidly accelerating AI development, data privacy and intellectual property rights must not be treated as afterthoughts. Users have a right to know where their data goes, and the creative labor of researchers deserves to be respected.
This dispute may be just the beginning. As AI penetrates deeper into specialized professional fields, similar conflicts will almost certainly keep emerging. Only by establishing transparent, fair, and trustworthy rules for data use can AI technology truly become a force that advances human progress — rather than one that threatens it.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.