Loading slide
That was a dense chapter, so let us throw almost all of it away and keep only what you actually need going forward. Three plain sentences.
Deep just means many layers. A network with lots of layers stacked up, each one finding a slightly richer pattern than the one below. Nothing more exotic than that.
For years, going deep broke training. The signal that tells the front of the network how to fix itself faded a little at every layer it travelled back through, like a whisper passed down a long line. By the time it reached the front, there was nothing left to hear, so the front never learned.
A few simple fixes let the signal survive the trip. The main one, ReLU, was almost embarrassingly simple: stop squashing the signal at each step, and it stops fading. Once it could reach the front intact, deep networks finally trained.
If you remember only this, you are ready. Deep was always more powerful in theory. These fixes are what made it work in practice. Everything else in this chapter was just the story of how the signal faded and how it was rescued.