Loading slide
Loading contents...
"Most of the readable internet" is the phrase everyone uses, and it slides past a decision that matters.
The internet is not a training set. Poured in raw it is mostly unusable: navigation menus, cookie banners, adverts, the same article copied across four hundred sites, automatically generated filler, spam, and enormous quantities of text that is simply gibberish.
So before any training happens, the pile is worked over. Roughly, and the details differ between organisations:
Removed. Boilerplate, markup, pages that are mostly links, text that fails basic quality checks.
Deduplicated. The same passage appearing many times gets collapsed, because otherwise the model sees a popular paragraph ten thousand times and treats it as far more important than it is.
Sorted by quality. Some sources get weighted more heavily than others. A curated encyclopedia might be counted several times over; a random forum thread might be counted once or dropped.
Filtered for content. Categories of material are excluded by policy.
Every one of those steps is a judgment call made by people. Somebody decided what counts as good text. Somebody decided which sources deserved extra weight, which languages were worth including at what proportion, and where the quality threshold sat.
Those decisions are upstream of everything the model ends up being. A model trained mostly on English text from a handful of countries will be better at that English and will carry the assumptions common in it. Not because anyone intended that, but because the pile was assembled that way, by people making reasonable-seeming choices at a scale where nobody can read what they are choosing.
It is where the answer usually begins when someone asks why a model carries a particular bias, long before training started.
Back to the curve, then, and to what was actually so surprising about its shape.