Loading slide

Why Shallow Networks Failed

  1. 01Why Shallow Networks Failed
  2. 02Memory card
  3. 03What "deep" actually means
  4. 04The vanishing gradient
  5. 05A small step inside each neuron
  6. 06Why bother bending the number?
  7. 07The bend they chose
  8. 08How the flat ends starve the signal
  9. 09The absurdly simple fix
  10. 10Take a breath
  11. 11Other fixes
  12. 12Depth vs. width
  13. 13Memory card
  14. 14The research gap
  15. 15Reinforce your understanding
  16. 16Question: The vanishing gradient
  17. 17Question: Why does a simple fix work?
  18. 18Question: Why does depth matter?
  19. 19Quiz: answer
  20. 20The one idea to keep
  21. 21The gradient problem was solved. What was still missing was scale.
  22. 22Want to go deeper?
3 / 21
BackNext

What "deep" actually means

One word is about to do a lot of work in this module, so let's pin it down now.

A network's layers sit between the input and the answer. In the last module, we stacked a few of them and watched each one find a slightly more specific pattern than the one before. A shallow network has just one or two of those layers. A deep network has many: ten, fifty, sometimes hundreds, stacked one after another.

That is the entire meaning of the term. Deep learning is just this idea: training a network with many layers, so each layer can build on the patterns the layer before it found.

The word makes it sound exotic. It isn't. "Deep" is a comment about the shape of the network, nothing more. Tall, if you prefer. A lot of layers, one on top of the next.

So why didn't people just stack the layers and walk away? Because for years, the moment you went deep, training broke. The next slide is why.

Citations(2)↓
  1. 1. deeplearningbook.org
  2. 2. ibm.com