Loading slide

Why Shallow Networks Failed

  1. 01Why Shallow Networks Failed
  2. 02Memory card
  3. 03What "deep" actually means
  4. 04The vanishing gradient
  5. 05A small step inside each neuron
  6. 06Why bother bending the number?
  7. 07The bend they chose
  8. 08How the flat ends starve the signal
  9. 09The absurdly simple fix
  10. 10Take a breath
  11. 11Other fixes
  12. 12Depth vs. width
  13. 13Memory card
  14. 14The research gap
  15. 15Reinforce your understanding
  16. 16Question: The vanishing gradient
  17. 17Question: Why does a simple fix work?
  18. 18Question: Why does depth matter?
  19. 19Quiz: answer
  20. 20The one idea to keep
  21. 21The gradient problem was solved. What was still missing was scale.
  22. 22Want to go deeper?
BackNext chapter

By the mid-2000s, the tools existed. ReLU. Better initialization. Batch normalization. Skip connections. The vanishing gradient was a solvable problem, and researchers had solved it, piece by piece.

But solving it in theory wasn't the same as making it work in practice. Training a deep network on anything interesting still took too long. The hardware was too slow. The datasets were too small to let deep networks prove their advantage over the shallower alternatives.

The technique was ready. What was missing was scale.

The next chapter is about where that scale came from, and why it had been sitting, unused, inside gaming computers the whole time.

Citations