All the way up

Four adjustable numbers is a very small model. Press scale up and keep pressing.

More going in. Grey units join your four, which is the year-and-maker problem made visible: a model that accounts for more has to be able to sense more.

Rows in the middle. These are new, and they are the subject of the next chapter. They let a model build patterns that no straight sum of inputs can express, which is a real change in what it is able to learn. They also break the tidy rule you have been using: error x input x learning rate cannot tell a weight in an early row how much it contributed to the final error. Both of those are chapter two's business.

Keep pressing until it stops fitting, and then know that it is still nothing. A real language model has hundreds of billions of adjustable numbers. To draw one at that last density you would need a page millions of times taller than the one you are looking at. Your four sit somewhere in the corner of it, too small to find.

One core idea survives every press: examples produce errors, and those errors guide changes to adjustable weights. Guess, check the error, nudge, repeat. You watched every step of it, one bottle at a time.

So a large language model is not merely this wine model with more dials. It shares the central idea of knowledge shaped into weights through error, while adding many hidden layers, different training goals, vast datasets, and far more elaborate machinery.

Which is exactly why four was worth your afternoon. Nobody can hold the large ones in their head. You held this one completely, watched its numbers start meaningless, and watched them end up holding a rule nobody had written down.