Loading slide
In 2020, researchers at OpenAI published something that sounds dry but changed how the whole field placed its bets. They measured how a language model improves as you scale it, more parameters, more text, more computing power, and found the improvement was not jumpy or random. It followed a smooth, predictable curve.
Think of what that means. If you knew how a model performed at one size, you could draw the curve forward and estimate, before spending a penny, how much better it would get at ten times the size. The link between scale and skill had become legible, like a recipe where doubling the ingredients reliably doubles the loaf.
That predictability was worth a fortune. It turned a gamble into a plan. A company could look at the curve and decide that yes, a hundred million dollars of computing power would buy a model this much better, and be roughly right. The race to scale was on, because for the first time the payoff could be forecast.
But the smooth curve was only half the story, and the tamer half. Scaling did make models better at things they already did, exactly as the curve promised. What no curve predicted was that scaling would also make them do things they had never done at all.