Loading slide
The honest answer: more than you can picture. So instead of a number, let's feel the growth.
One quick note on words first. When people count the numbers in a model, they usually call them parameters rather than weights. For now, treat the two as the same thing. More on that word soon.
In the late 1980s and 1990s, when neural networks first started working, a large model might have had a few thousand parameters. Enough to fill a few pages of text. Researchers were proud of that.
By the early 2000s, the number crept into the millions. Still small enough that you could imagine printing them out, given enough paper.
Then something changed. Computing got faster. Data got bigger. And the numbers started climbing in a way that felt almost violent. Within a few years, models crossed a billion parameters. Then ten billion. Then a hundred billion. Then more.
Large language models can contain billions of parameters, or sometimes far more, though the exact numbers are not always made public.
To put a hundred billion in perspective: if you read one number per second, without stopping, it would take you over three thousand years to read them all.
When a company says they have released a new model, this is what they mean: they have produced a new file, with a new configuration of numbers that behaves differently from the ones before it. Different models have different sizes. A smaller model might run on a laptop. A larger one usually requires specialised hardware and a lot of electricity to run quickly. That is part of why you connect to their servers rather than running it yourself.
The file holding those numbers is not a document or a spreadsheet. It is closer in size to a film archive. And yet it fits on a hard drive, travels across the internet, and produces answers in milliseconds.
Which raises an obvious question. Why would anyone want that many numbers in the first place? What do you even get for them?