Loading slide
Loading contents...
A parameter is one adjustable number, and on its own it does almost nothing. What a pile of them buys is room. The more a model has, the more shape it can bend itself into, and so the more of the pattern in its examples it is able to follow.
Watch it happen beside these examples. The dots are the examples, and they never move. Only the machine trying to fit them changes.
With very few parameters the model can manage a straight line. It catches the general drift and misses everything else, because a line is all the shape it has. Give it more and it can follow the actual curve the examples describe.
Then keep going. With far more parameters than the pattern needs, the model starts bending through every single dot, chasing the scatter rather than the shape underneath it. It now fits these examples beautifully and would be worse at anything new, because it has learned the accidents of this particular handful of points. That failure has a name,
So parameters raise the ceiling on how much detail a model can hold. They do not decide what it puts up there. The training data, the design, and the training process still settle that.
That explains why a bigger model can be better at what it already does. It does not explain the other thing scale was reported to do.