Loading slide

A Network That Learns

  1. 01A Network That Learns
  2. 02Stacking neurons
  3. 03Quick recall
  4. 04What stacking does
  5. 05Hidden layers
  6. 06The problem of credit
  7. 07Backpropagation
  8. 08Sharing the blame
  9. 09The search problem
  10. 10Fog on a hill
  11. 11Gradient descent
  12. 12Just to be clear
  13. 13What the hidden layer learns
  14. 14What the hidden layer discovers
  15. 15Did it actually learn?
  16. 16Generalization
  17. 17Limits in the 1990s
  18. 18Reinforce your understanding
  19. 19Question: Backpropagation
  20. 20Question: What the hidden layer learns
  21. 21Question: Memorizing vs. generalizing
  22. 22Quiz: answer
  23. 23The method existed. What was missing was scale.
  24. 24Want to go deeper?
6 / 23
BackNext

The problem of credit

In the last chapter, we saw what training a single neuron looks like: get something wrong, nudge the weights, repeat. We also saw why a whole network makes that harder: the mistake shows up at the output, but the weights that caused it are buried deep inside.

Now let's name that problem properly.

The network looks at an image and says "dog", the answer was "cat." There are dozens of neurons across many layers, each with their own weights, each contributing a little to the final answer. Which weights caused the mistake? The ones near the output? The ones buried three layers back? The wrong answer is visible. The cause is hidden.

This became known as the credit assignment problem: when the network gets it wrong, which weights were responsible? Which weights needed adjusting?

Researchers knew the multi-layer architecture was more powerful. They just couldn't figure out how to train it. To train a network, you need to know which weights to adjust. But the error only shows up at the output. There was no way to trace it back through the layers and know what caused it. The problem sat unsolved for two decades.

Citations(2)↓
  1. 1. nature.com
  2. 2. ojs.aaai.org
Citations(2)↓
  1. 1. nature.com
  2. 2. ojs.aaai.org