Take a breath
Two ideas so far, and the rest of the module leans on both. Pin them down.
The first is the smooth curve. With model size, data, and compute balanced, prediction loss often changes in a regular way. That makes training runs easier to forecast.
The second is the apparent jump. A pass-or-fail task score can stay flat, then rise suddenly. Sometimes that marks an important threshold. Sometimes the scoring rule hides gradual progress underneath.
Broad prediction improvement can be smooth while a particular benchmark looks sudden. We should inspect the measurement before claiming a new power appeared from nowhere.
Now for what happened when one lab pushed the size argument as far as it would go, three times in a row.
# citations