Loading slide

Loading contents...

[███░░░░░░░░░░░░░░░][█████░░░░░░░░░░░░░░░░░░░░░░░]6 / 37
<back>next

The improvement had a shape

You have met the scaling law already. Measured in 2020, it is the finding that a model's prediction error falls along a smooth curve as parameters, training data, and computing power grow. Now that you know what a parameter is and what training costs, it is worth being precise about what that curve does and does not promise.

What it measures is prediction error, the model's average surprise at the next word. Nothing else. It is a curve about one number.

Think of measuring fuel use before a long drive. A reliable curve lets you estimate what a longer trip will cost. It says nothing about whether the destination is worth visiting.

Two limits follow. Model size alone was never the recipe. A later study found the large models of the day were badly undertrained, and that a smaller model given more text beat a larger one built on the same budget. Size and data have to grow together.

And a falling loss is not the same as a better assistant. The curve promises that the model's next-word guesses improve on average. It does not promise that any particular skill you care about arrives at any particular size.

Which leaves a plainer question underneath all of it. Why would piling on more numbers make any difference at all?

# citations(2)↓
  1. [1]arxiv.org
  2. [2]arxiv.org