Loading slide

Scale and Emergence

  1. 01Scale and Emergence
  2. 02Memory card
  3. 03The loop that was being scaled
  4. 04What happens when you just make it bigger
  5. 05The improvement had a shape
  6. 06Hold this contrast
  7. 07Abilities that appear from nowhere
  8. 08Take a breath
  9. 09The GPT-3 moment
  10. 10A raw model is not yet helpful
  11. 11Teaching it to answer
  12. 12Letting humans steer it
  13. 13The side effect: confident and wrong
  14. 14Reinforce your understanding
  15. 15Question: What emergence means
  16. 16Question: Making the model helpful
  17. 17Question: Confident and wrong
  18. 18Quiz: answer
  19. 19Bigger changed what was possible
  20. 20Want to go deeper?
5 / 19
BackNext

The improvement had a shape

In 2020, researchers at OpenAI published something that sounds dry but changed how the whole field placed its bets. They measured how a language model improves as you scale it, more parameters, more text, more computing power, and found the improvement was not jumpy or random. It followed a smooth, predictable curve.

Think of what that means. If you knew how a model performed at one size, you could draw the curve forward and estimate, before spending a penny, how much better it would get at ten times the size. The link between scale and skill had become legible, like a recipe where doubling the ingredients reliably doubles the loaf.

That predictability was worth a fortune. It turned a gamble into a plan. A company could look at the curve and decide that yes, a hundred million dollars of computing power would buy a model this much better, and be roughly right. The race to scale was on, because for the first time the payoff could be forecast.

But the smooth curve was only half the story, and the tamer half. Scaling did make models better at things they already did, exactly as the curve promised. What no curve predicted was that scaling would also make them do things they had never done at all.

Citations(1)↓
  1. 1. arxiv.org
Citations(1)↓
  1. 1. arxiv.org