Loading slide

A Network That Learns

  1. 01A Network That Learns
  2. 02Stacking neurons
  3. 03Quick recall
  4. 04What stacking does
  5. 05Hidden layers
  6. 06The problem of credit
  7. 07Backpropagation
  8. 08Sharing the blame
  9. 09The search problem
  10. 10Fog on a hill
  11. 11Gradient descent
  12. 12Just to be clear
  13. 13What the hidden layer learns
  14. 14What the hidden layer discovers
  15. 15Did it actually learn?
  16. 16Generalization
  17. 17Limits in the 1990s
  18. 18Reinforce your understanding
  19. 19Question: Backpropagation
  20. 20Question: What the hidden layer learns
  21. 21Question: Memorizing vs. generalizing
  22. 22Quiz: answer
  23. 23The method existed. What was missing was scale.
  24. 24Want to go deeper?
17 / 23
BackNext

Limits in the 1990s

Three things were missing.

The first was depth. Networks with many layers were much more powerful in theory, but training them didn't work. The correction signal that traveled backwards through the network to update the weights would fade as it passed through each layer. By the time it reached the early layers, it had almost disappeared. Those weights barely changed. The network couldn't really learn.

The second was computing power. Even training a small network on a modest dataset took a very long time on the hardware available in the 1990s. Researchers often had to wait days for results. Bigger, more capable networks were simply out of reach.

The third was data. For a network to learn to recognize a cat, it needs to see thousands of examples of cats, each one labeled. Someone has to go through every single image and write down the answer. In the 1990s, that kind of labeled data barely existed. Collecting it was slow, expensive, and largely done by hand.

The idea was right. The math worked. But without the computing power to train large networks, and without the data to feed them, neural networks stayed a promising theory rather than a practical tool.

Citations(2)↓
  1. 1. deeplearningbook.org
  2. 2. developers.google.com