Loading slide

Scale and Emergence

  1. 01Scale and Emergence
  2. 02Memory card
  3. 03The loop that was being scaled
  4. 04What happens when you just make it bigger
  5. 05The improvement had a shape
  6. 06Hold this contrast
  7. 07Abilities that appear from nowhere
  8. 08Take a breath
  9. 09The GPT-3 moment
  10. 10A raw model is not yet helpful
  11. 11Teaching it to answer
  12. 12Letting humans steer it
  13. 13The side effect: confident and wrong
  14. 14Reinforce your understanding
  15. 15Question: What emergence means
  16. 16Question: Making the model helpful
  17. 17Question: Confident and wrong
  18. 18Quiz: answer
  19. 19Bigger changed what was possible
  20. 20Want to go deeper?
13 / 19
BackNext

The side effect: confident and wrong

The human example answers were written with confidence. Clear, authoritative, well-structured. That is what "good" looked like to the people writing them, and what raters kept preferring. So that is what the model learned to sound like: confident, fluent, certain.

Here is the trap. The model learned to sound confident, full stop. Not confident when it has a basis and hesitant when it does not. Confident always. It picked up the style of sure-footed human writing, and sure-footed human writing almost never says "I am not certain." So the model rarely says it either. It produces a smooth, authoritative answer whether or not anything real sits underneath.

This is what people mean when they say a model hallucinates. It is not lying, because lying needs a sense of the truth to push against, and the model has none. It is doing precisely what it was trained to do: produce the most plausible-sounding continuation. Sometimes the most plausible-sounding continuation happens to be false. There is no inner gauge tracking what it really knows versus what it is inventing, so a fabricated citation rolls out with the same easy confidence as a true one.

This matters enough that we will return to it in this module's final chapter, where we look at what the weights actually contain. For now, hold the shape of it: the very training that made the model helpful and fluent is the same training that made it sound just as sure when it is wrong as when it is right. The polish is not evidence. It was the goal.

Citations
Citations