Loading slide

A Network That Learns

  1. 01A Network That Learns
  2. 02Stacking neurons
  3. 03Quick recall
  4. 04What stacking does
  5. 05Hidden layers
  6. 06The problem of credit
  7. 07Backpropagation
  8. 08Sharing the blame
  9. 09The search problem
  10. 10Fog on a hill
  11. 11Gradient descent
  12. 12Just to be clear
  13. 13What the hidden layer learns
  14. 14What the hidden layer discovers
  15. 15Did it actually learn?
  16. 16Generalization
  17. 17Limits in the 1990s
  18. 18Reinforce your understanding
  19. 19Question: Backpropagation
  20. 20Question: What the hidden layer learns
  21. 21Question: Memorizing vs. generalizing
  22. 22Quiz: answer
  23. 23The method existed. What was missing was scale.
  24. 24Want to go deeper?
11 / 23
BackNext

Gradient descent

That's gradient descent. The name is just plain description: a gradient is the slope of the ground at any point, meaning how steep it is and which way it tilts. Descent means going down. So gradient descent means: find the slope, step downward, repeat.

At each step, check which direction the ground slopes downward and take a small step that way. Then check again. Repeat, thousands of times, across millions of examples. The weights gradually settle toward a valley, a combination that performs well.

The size of each step matters. Too large and you overshoot the valley entirely. Too small and training takes forever. The step size is called the learning rate, and finding the right one is part of the craft of building networks.

Backpropagation tells you which direction is downhill. Gradient descent tells you how to walk.

Citations(2)↓
  1. 1. developers.google.com
  2. 2. developers.google.com
Citations(2)↓
  1. 1. developers.google.com
  2. 2. developers.google.com