Module V
Going Deeper
Chapter I
Why Shallow Networks Failed
Some ideas are right before the world is ready for them.
By the end of the last module, we had a picture of what a trained network is: layers of simple neurons, tuned by training until the whole thing becomes useful. The idea was sound. The mathematics worked.
And yet, for decades, these networks kept falling short in practice. Stack on more layers, hoping for more power, and training got worse, not better. The signal that was supposed to guide the learning seemed to dissolve somewhere on its way through the network.
This was not a flaw in the idea. It was a single technical obstacle, the kind that looks fundamental right up until someone finds the way around it. This chapter is the story of that obstacle, and the small fix that finally let networks go deep.