Sam Altman has practically admitted that OpenAI has run models trained entirely on synthetic data. He also said that when it comes to math, human data may no longer be necessary.
In a podcast with Atlantic CEO Nicholas Thompson, Altman discussed the role of AI-generated data in training future models. Thompson noted that the internet is now filled with AI-generated content, and even human writing styles are starting to mimic AI. "GPT-4 is the last model that didn't use much AI data," Thompson said, and Altman nodded in agreement.
Thompson then pressed directly: has OpenAI ever run a model trained purely on synthetic data — using AI outputs to train the next generation? Altman paused and said, "I'm not sure I should say." That was as close to a confirmation as you'll get.
He went on to argue that the core of these models is reasoning, and pure synthetic data can handle that just fine. He used math as an example: could a model that has never seen human data outperform a human at math? "I think yes." But a model that has never encountered human culture understanding human values? "That probably won't work."
The conversation echoes a long-standing concern in AI: the "mad cow disease" problem — if AI keeps feeding on its own outputs, does the information degrade over generations? Altman's answer suggests a split: math can be taught without us, but culture and values still need a human touch.
Altman's comments come as OpenAI continues to push boundaries, including merging its Codex coding model into GPT-5.4. He has also recently argued that token pricing is a dying model, suggesting customers care only about results, not how tokens are consumed.