CourionAI
EN
Newsletter
← All news
research 2 min read

Reinforcement Learning Pioneer Rich Sutton Calls Synthetic Data a Big Mistake

Sutton argues that a world of endless complexity cannot be learned from data a model generates for itself. It is a direct challenge to one of the industry's favourite shortcuts.

A tangled maze of lines beside a small closed loop feeding into itself

Rich Sutton, one of the founding figures of reinforcement learning and co-author of the field’s standard textbook, has called synthetic data a “big mistake” given how complex the real world is. The remark, reported on 20 August, lands squarely on a practice that has quietly become standard across the industry.

Synthetic data is exactly what it sounds like: training material a model generates itself, or that another model generates for it, instead of text and images collected from the world. Labs lean on it heavily because real data is finite, expensive, legally fraught and increasingly picked clean. Sutton’s objection connects to an idea he has argued for years, the big world hypothesis: any learning system is always far smaller than the environment it operates in, so it can never hold a complete model of that world. If that is true, then a system generating its own training material is drawing from the part of the world it has already captured, which is by definition the small part. You get more of what the model already knows, not more of what it does not.

This is not a fringe position, but it is a contested one, and it fits into a broader argument that has picked up in recent weeks. Mathematicians Timothy Gowers and Peter Sarnak have separately said that language models combine known methods well but struggle to invent genuinely new foundations. A DeepMind paper argued the same bottleneck from a different direction. The counter-case is straightforward and has evidence behind it: synthetic data demonstrably works for teaching specific skills, particularly reasoning steps and code, where a correct answer can be checked automatically. What Sutton is questioning is not whether it helps on benchmarks but whether it can ever produce something genuinely new, and those are different claims that often get argued as if they were one.

What this means for you: nothing practical, and everything about how you read the next round of announcements. When a lab reports a jump on a benchmark, it is worth knowing whether the gain came from new information about the world or from a model practising against itself. Both are real improvements, but only one of them tells you the system now knows something it did not before. If you use these tools daily, Sutton’s argument also explains a familiar frustration: models are excellent at recombining what exists and noticeably weaker the moment you need something with no precedent.

Sources

Source: https://the-decoder.com/ki-pioneer-sutton-calls-synthetic-data-a-big-mistake-in-the-face-of-an-infinitely-complex-world/

Next story

AI Could Write Like a Human. Safety Training Is What Gives It Away

Pangram's CTO argues that language models are detectable not because they lack ability but because post-training narrows how they express themselves. The effect has a name: mode collapse.

Rows of identical typewriters funneling into one narrow bottleneck shape