Loading slide

A Network That Learns

  1. 01A Network That Learns
  2. 02Stacking neurons
  3. 03Quick recall
  4. 04What stacking does
  5. 05Hidden layers
  6. 06The problem of credit
  7. 07Backpropagation
  8. 08Sharing the blame
  9. 09The search problem
  10. 10Fog on a hill
  11. 11Gradient descent
  12. 12Just to be clear
  13. 13What the hidden layer learns
  14. 14What the hidden layer discovers
  15. 15Did it actually learn?
  16. 16Generalization
  17. 17Limits in the 1990s
  18. 18Reinforce your understanding
  19. 19Question: Backpropagation
  20. 20Question: What the hidden layer learns
  21. 21Question: Memorizing vs. generalizing
  22. 22Quiz: answer
  23. 23The method existed. What was missing was scale.
  24. 24Want to go deeper?
BackNext chapter

Here's what we just covered.

A network starts out knowing nothing. Every connection between neurons has a weight, and at the beginning those weights are random.

The network makes a prediction, compares it to the right answer, measures how far off it was, and uses that gap to figure out which weights to nudge, and by how much.

Do that millions of times, across thousands of examples, and the weights slowly settle into a configuration that produces good answers.

Nobody programs what the hidden layers look for. Nobody writes down the rules. The network discovers them on its own, because certain internal patterns turn out to be useful for getting the right answer.

The measure of success is generalization: not getting the training examples right, but getting examples it has never seen right.

That's the whole mechanism. It sounds almost too simple for something that can recognize a face, translate a sentence, or write a paragraph. But this is genuinely how it works.

What comes next is a stranger question. Once training is done, what remains? The network has "learned" something. But where is that knowledge? What form does it take? And in what sense does a collection of numbers actually know anything at all?

Citations