Loading slide

Why Shallow Networks Failed

  1. 01Why Shallow Networks Failed
  2. 02Memory card
  3. 03What "deep" actually means
  4. 04The vanishing gradient
  5. 05A small step inside each neuron
  6. 06Why bother bending the number?
  7. 07The bend they chose
  8. 08How the flat ends starve the signal
  9. 09The absurdly simple fix
  10. 10Take a breath
  11. 11Other fixes
  12. 12Depth vs. width
  13. 13Memory card
  14. 14The research gap
  15. 15Reinforce your understanding
  16. 16Question: The vanishing gradient
  17. 17Question: Why does a simple fix work?
  18. 18Question: Why does depth matter?
  19. 19Quiz: answer
  20. 20The one idea to keep
  21. 21The gradient problem was solved. What was still missing was scale.
  22. 22Want to go deeper?
20 / 21
BackNext

That was a dense chapter, so let us throw almost all of it away and keep only what you actually need going forward. Three plain sentences.

Deep just means many layers. A network with lots of layers stacked up, each one finding a slightly richer pattern than the one below. Nothing more exotic than that.

For years, going deep broke training. The signal that tells the front of the network how to fix itself faded a little at every layer it travelled back through, like a whisper passed down a long line. By the time it reached the front, there was nothing left to hear, so the front never learned.

A few simple fixes let the signal survive the trip. The main one, ReLU, was almost embarrassingly simple: stop squashing the signal at each step, and it stops fading. Once it could reach the front intact, deep networks finally trained.

If you remember only this, you are ready. Deep was always more powerful in theory. These fixes are what made it work in practice. Everything else in this chapter was just the story of how the signal faded and how it was rescued.

Citations