The AI Data Wall Is Real. Synthetic Data Won't Save Us.
We're running out of human-generated text to train on. The industry's solution — training on AI output — creates a quality death spiral.
The AI industry has a quiet problem that nobody in a keynote wants to acknowledge. We’re running out of training data.
The data cliff
Estimates from multiple research groups converge on the same finding: the total volume of high-quality human-generated text available online is roughly 10-15 trillion tokens. That includes books, articles, code, academic papers, and the cleaned web corpus.
Current frontier models train on 5-10 trillion tokens. The next generation aims for 15-20 trillion. We’re approaching the ceiling.
The web is already mined. Common Crawl, the backbone of most training datasets, has been scraped exhaustively. Duplicate removal and quality filtering reduce the useful corpus further.
Books are gated. Copyright enforcement has restricted access to the full text of published books. The Books3 dataset — used by Meta and others — was taken down after legal action.
Fresh content is increasingly AI-generated. By some estimates, 40% of new web content in 2026 is AI-generated. Training on this creates a well-documented problem.
The synthetic data solution — and its flaw
The industry’s answer is synthetic data: using AI to generate training data for the next generation of AI. It sounds elegant. It’s actually a trap.
Model collapse. When models train on their own output, quality degrades with each generation. This isn’t theoretical — it’s been demonstrated in multiple papers across different modalities. Text, images, code — all degrade when trained on synthetic data from the same model family.
Diversity loss. Human text is messy, inconsistent, and surprising. AI text is optimized, bland, and predictable. Training on AI output amplifies the mode-seeking behavior that makes models less creative over time.
Distribution narrowing. The edges of human knowledge — niche expertise, unusual perspectives, dialectal variation — are exactly what AI models underrepresent. Training on AI output removes these edges entirely.
What might actually work
Multimodal data. Video, audio, and sensor data represent vastly more information than text. The shift toward multimodal training isn’t just about capability — it’s about accessing new data sources.
Proprietary data. Companies with unique datasets (Reddit’s conversation logs, GitHub’s code history, internal enterprise data) have a significant advantage. This is why data licensing deals have accelerated.
Better data curation. Instead of training on more tokens, train on better tokens. Microsoft’s Phi models demonstrated that small, high-quality datasets can compete with massive ones.
Active learning. Selectively sampling the most informative training examples rather than dumping the entire internet into the model. More efficient, but harder to implement at scale.
The uncomfortable truth
The data wall means the era of “just scale up training data” is ending. The next breakthroughs won’t come from more data — they’ll come from better architectures, more efficient training, and new data sources.
That’s a harder problem. And nobody has solved it yet.

