Loading slide

A Network That Learns

  1. 01A Network That Learns
  2. 02Stacking neurons
  3. 03Quick recall
  4. 04What stacking does
  5. 05Hidden layers
  6. 06The problem of credit
  7. 07Backpropagation
  8. 08Sharing the blame
  9. 09The search problem
  10. 10Fog on a hill
  11. 11Gradient descent
  12. 12Just to be clear
  13. 13What the hidden layer learns
  14. 14What the hidden layer discovers
  15. 15Did it actually learn?
  16. 16Generalization
  17. 17Limits in the 1990s
  18. 18Reinforce your understanding
  19. 19Question: Backpropagation
  20. 20Question: What the hidden layer learns
  21. 21Question: Memorizing vs. generalizing
  22. 22Quiz: answer
  23. 23The method existed. What was missing was scale.
  24. 24Want to go deeper?
7 / 23
BackNext

Backpropagation

In 1986, , , and published a solution. The method was called backpropagation.

Think of the whisper game. A message passes down a line of people, each one distorting it a little. By the end, the message is wrong. Now play it in reverse: the last person whispers the correction back up the line. Each person hears what they got wrong and by how much.

A network does the same thing. It makes a prediction, gets it wrong, then traces the error backwards through every layer. The mistake shows up at the very end, in the final answer. But the weights that caused it are scattered all through the network, including ones buried many layers back. Backpropagation is the procedure that carries the error back from the output, layer by layer, until it reaches every last one of them.

That backward journey matters. The error does not jump straight to each weight. It retraces, in reverse, the exact path the signal took on the way forward, passing back through the same connections that produced the answer. That is the part of the story we want to slow down and look at properly.

Did you know?

Backpropagation required no new mathematics. The key tool was the chain rule, a technique from calculus that had existed since the 1700s. The chain rule describes how a change in one part of a sequence ripples through to the end. If A affects B, and B affects C, it tells you how much A affects C. That is exactly what backpropagation needs: a weight buried several layers back is several steps away from the output, and the chain rule is what lets you trace its influence all the way through.
Citations(2)↓
  1. 1. nature.com
  2. 2. developers.google.com
Citations(2)↓
  1. 1. nature.com
  2. 2. developers.google.com