Loading slide

Why Shallow Networks Failed

  1. 01Why Shallow Networks Failed
  2. 02Memory card
  3. 03What "deep" actually means
  4. 04The vanishing gradient
  5. 05A small step inside each neuron
  6. 06Why bother bending the number?
  7. 07The bend they chose
  8. 08How the flat ends starve the signal
  9. 09The absurdly simple fix
  10. 10Take a breath
  11. 11Other fixes
  12. 12Depth vs. width
  13. 13Memory card
  14. 14The research gap
  15. 15Reinforce your understanding
  16. 16Question: The vanishing gradient
  17. 17Question: Why does a simple fix work?
  18. 18Question: Why does depth matter?
  19. 19Quiz: answer
  20. 20The one idea to keep
  21. 21The gradient problem was solved. What was still missing was scale.
  22. 22Want to go deeper?
1 / 21
BackNext

Module V

Going Deeper

Chapter I

Why Shallow Networks Failed

In this chapter

  • The vanishing gradientwhy adding more layers made training worse
  • A small fixthe change to a single neuron that made depth trainable
  • Depth over widthwhy stacking layers was the right shape for the problem
  • The waiting gamewho kept going, and what they were waiting for

Some ideas are right before the world is ready for them.

By the end of the last module, we had a picture of what a trained network is: layers of simple neurons, tuned by training until the whole thing becomes useful. The idea was sound. The mathematics worked.

And yet, for decades, these networks kept falling short in practice. Stack on more layers, hoping for more power, and training got worse, not better. The signal that was supposed to guide the learning seemed to dissolve somewhere on its way through the network.

This was not a flaw in the idea. It was a single technical obstacle, the kind that looks fundamental right up until someone finds the way around it. This chapter is the story of that obstacle, and the small fix that finally let networks go deep.

Citations