CR § 08 · THE CURRENT RECORD
The Model Eats Itself
Systems trained on human writing are now trained partly on their own output, and the failure mode has a name and a demonstrated shape.
The archive has argued that systems trained on the full corpus of human textual output function as mirrors of what has been written rather than as sources of independent judgement. That argument has a consequence that is now measurable.
As generated text accumulates on the open web, it enters the training data for subsequent systems. Models are increasingly trained on material produced by models. The published work on this describes a degradation the authors have called model collapse: with each generation trained substantially on synthetic output, the tails of the distribution thin, variance narrows, and rare or unusual material disappears first.
The mechanism is not mysterious. A model trained on a distribution produces samples that under-represent the extremes, because extremes are rare and generation regresses toward the typical. Train the next model on those samples and the extremes thin further. Iterate and the distribution converges on its own centre.
What is lost first is precisely what this archive cares about. Minority positions. Idiosyncratic phrasing. Unusual arguments held by few people. Regional and non-dominant-language material. The rare thing is rare in the training data, so it is under-generated, so it is rarer in the next round.
This is a suppression mechanism with no author at all, and it is the furthest along that spectrum the archive has travelled. The Index had compilers. Ranking has an objective function somebody chose. This requires only that the process run.
The comparison to centralisation is exact and it is the argument of THE RECORD arriving in the present. The Qin gathered every copy into one authorised library and created a single point of failure. A monoculture of training data is the same structure: apparent abundance, actual convergence, and no redundancy in the tails.
There are mitigations under discussion and some in use: provenance tracking, watermarking generated text, preferentially weighting verified human corpora, and preserving snapshots of the pre-2022 web as a reference distribution. That last one is essentially an archival response, and it is the one this archive would recognise.
The obvious caution applies. This is a young literature, the demonstrations have mostly been in controlled settings with high proportions of synthetic data, and real training pipelines filter aggressively. The effect at realistic mixtures is not established. It should not be stated as a settled outcome.
But the direction is not in dispute, and the incentive structure is unhelpful. Generated text is cheap and abundant. Verified human text is expensive and finite. Every economic pressure points toward the first.
The practical implication for anyone writing is small and concrete. Material that exists only inside a system that trains on itself is at risk of being averaged away. Material that exists as a document, on a domain, in a feed, in print, is not.