Loading slide

A Network That Learns

  1. 01A Network That Learns
  2. 02Stacking neurons
  3. 03Quick recall
  4. 04What stacking does
  5. 05Hidden layers
  6. 06The problem of credit
  7. 07Backpropagation
  8. 08Sharing the blame
  9. 09The search problem
  10. 10Fog on a hill
  11. 11Gradient descent
  12. 12Just to be clear
  13. 13What the hidden layer learns
  14. 14What the hidden layer discovers
  15. 15Did it actually learn?
  16. 16Generalization
  17. 17Limits in the 1990s
  18. 18Reinforce your understanding
  19. 19Question: Backpropagation
  20. 20Question: What the hidden layer learns
  21. 21Question: Memorizing vs. generalizing
  22. 22Quiz: answer
  23. 23The method existed. What was missing was scale.
  24. 24Want to go deeper?
10 / 23
BackNext

Fog on a hill

To understand why this is hard, think about the numbers.

A small network might have a few thousand weights. A larger one has millions. Each weight is a number that can be set to anything: 0.3, 1.7, -0.4.

When you have a million weights and each one can take any value, the number of possible combinations is not just large. It's so large that if you tried one combination every second, you would not finish before the universe ended.

Trying them all is not an option.

What you need is a smarter approach. And here's the key insight: you don't need to find the perfect combination from scratch. You just need to keep moving in the right direction.

Imagine you're standing on a hillside in thick fog. You can't see the valley. You can't see the whole landscape. But you can feel which way the ground slopes under your feet.

So you take a step downhill. Then check again. Then another step.

You don't need to see the whole mountain. You just need to know which way is down from where you're standing.

Backpropagation tells you which direction is downhill. What's left is knowing how to walk.

Citations