Field notes
October 10, 2026
Why AI Models Get Worse When Trained on Their Own Outputs, and What Real Human Data Is For
Training on model outputs isn't automatically bad. Training on model outputs instead of real data is.
If you've heard that AI models get worse when they're trained on text written by other AI models, that's broadly true, with an important catch.
Models degrade when generated text replaces the real data they learned from. When generated text is added while the real data stays in the mix, the damage can be contained.
That real data is what people mean by a human floor. This post is about why the floor matters and what it's actually doing.
Why anyone trains on model outputs at all
Models learn from huge amounts of text. A lot of it comes from people: books, articles, forums, code.
There's a long-running argument about whether good human text is running out. One analysis people often cite, Villalobos and colleagues' "Will we run out of data?", projects that fresh human text may not keep up with how fast training sets grow. It's a projection and a live debate, not a settled deadline.
Either way, the pressure is real. And there's an obvious-looking fix. Models can produce text in nearly unlimited amounts. So why not train the next model on what the last one wrote?
Sometimes that's a reasonable thing to do. Generated data can fill gaps, create practice examples, or stand in for private data that can't be used directly. The trouble starts when it stops being a supplement and becomes a substitute.
There's also a version of this that nobody chooses. As more of the public web is written by models, anyone who collects text from the web ends up training on model outputs whether they meant to or not. The scraper can't reliably tell who wrote what.
What goes wrong
The best-known paper on this is Shumailov and colleagues in Nature in 2024. They called the effect model collapse.
Their finding, in plain terms: when models are trained over and over on data generated by earlier models, they develop defects that don't go away. The rare parts of the original data disappear first. The range of outputs narrows. Later generations end up with a distorted picture of what the real data looked like.
Here's why that happens, without any math.
A model is an estimate. It learns how often things appear in its training data, and it gets the common things roughly right. Rare things are harder. Something that showed up only a handful of times might get a tiny probability, or none.
Now generate a big pile of text from that model. Common things show up plenty. Rare things show up less than they did in the real data, or not at all, because the model already underrated them.
Train the next model on that pile. It sees even fewer of the rare things, so it underrates them more. Its outputs drop them further.
Each round, the edges get trimmed a little. Nothing dramatic happens in a single step. But the errors stack, and they stack in one direction: toward the middle, toward the common, toward a narrower picture of the world.
The rare cases are usually the ones you care about. Unusual dialects. Uncommon tools. Edge cases in a user's request. The odd question that comes up once in ten thousand. Those are exactly what recursive training wears away first.
Replace versus accumulate
The most useful distinction I've found in this research is between two ways of handling generated data.
Replace. Each new model trains mostly or only on what the previous model produced. The real data shrinks or disappears from the mix.
Accumulate. Each new model keeps the original real data and adds generated data on top. The real set doesn't get thrown out.
Gerstgrasser and colleagues compared these two directly in a 2024 paper, across language models and other kinds of models. Replacing tended toward collapse. Accumulating avoided it. In a simplified mathematical setting, they showed that with replacement the error keeps growing with each round, while with accumulation it stays bounded no matter how many rounds you run.
A 2025 paper by Kazdan, Schaeffer, and colleagues, "Collapse or Thrive," pushed the same comparison further. They confirmed that replacing collapses. They also found that accumulating stays stable even as real data becomes a smaller and smaller share of the growing pile, as long as the real data is still there.
That second point is easy to misread, so it's worth slowing down on.
"Real data is a shrinking percentage of everything we have" is not the same as "we deleted the real data." The first can be fine. The second is the dangerous one.
Sampling still matters
There's a wrinkle in the Kazdan paper that matters for anyone running training in practice.
Keeping the real data in storage isn't the whole story. What matters is whether it actually gets used in each training run.
They looked at a setup where all the data is kept, but each training run only draws a fixed-size sample from the full pile. As generated data grows, the real data gets drawn less often. In that case they saw slow degradation. Not a sudden crash, but a gradual slide.
So the question isn't just "do we still have the human data somewhere?" It's "how much of it is in this run?"
What the human floor is for
Put those findings together and the job of real human data gets clearer. It isn't seasoning. It does three specific things.
It anchors the model to the real distribution. Generated data is a copy of a model's estimate. Real data is the thing being estimated. Without it, each round can only learn from the last round's mistakes.
It keeps the rare cases alive. Real data still contains the unusual examples that recursive training trims away. As long as it's in the mix, those cases get seen.
It keeps the grading honest. If you test a model only on text written by models, you're checking it against its own habits. A held-out set of real data, one the model never trains on, is how you find out whether it still handles the world and not just its own echo.
How much is enough?
This is the question everyone wants a number for. I don't have one, and as far as I can tell the research doesn't offer a universal one either.
The safe amount depends on the kind of model, the data, how the training mix is sampled, and what you're trying to preserve. Theory papers are still arguing about how to even measure it in high-dimensional settings. Kazdan's group also found that the value of generated data depends on context. When real data is scarce, a modest amount of generated data can help. When real data is plentiful, piling on more generated data often hurts.
So if someone gives you a fixed percentage of human data that's always enough, be skeptical. That's ahead of the evidence.
What that means in practice
None of this argues against using generated data. It argues against using it carelessly. A few habits follow from the research:
- Don't replace. Accumulate. Generated data gets added. The real set stays, and you can point to it.
- Track the real-to-generated ratio in every training run. Treat it like any other setting you log, not as a footnote.
- Check what each run actually samples, not just what sits in storage.
- Test on held-out real data after every major run.
- Protect the rare cases on purpose. If something matters and is uncommon, make sure it's represented.
- Assume your outputs will come back. If you publish model-written text at scale, expect it to end up in someone's training data, possibly your own.
The short version
Training on model outputs isn't automatically bad. Training on model outputs instead of real data is.
Each generation of copying loses a little of the edges, and the edges are where the hard, useful cases live. Real human data is what keeps those edges in view. It's also the only fair test of whether a model still understands anything beyond itself.
I write fiction about what's lost when a record of the past gets erased, beginning with Shadows of the Forgotten.