Loading slide

The GPU Moment

  1. 01The GPU Moment
  2. 02One thing first
  3. 03Why training is slow
  4. 04What a GPU actually does
  5. 05The realization
  6. 06The scale this enabled
  7. 07Something unexpected
  8. 08Take a breath
  9. 09Data and compute
  10. 10Memory card
  11. 11Reinforce your understanding
  12. 12Question: Why GPUs and not faster CPUs?
  13. 13Question: Three ingredients
  14. 14Question: Why the field sped up
  15. 15Quiz: answer
  16. 16The hardware was ready. The data existed.
  17. 17Want to go deeper?
2 / 16
BackNext

This chapter is about a machine that made training fast. Before we meet it, lock in the one thing it speeds up, because the whole chapter rests on it.

Recall what training a network actually is, underneath. The network makes a guess, sees how wrong it was, and nudges its dials a little. Then it does it again. And again. Each nudge is a pile of small, simple sums, multiply these numbers, add them up, adjust. Not hard sums. Just an enormous number of them, repeated billions upon billions of times until the dials settle.

So training is not one clever, complicated calculation. It is the same tiny arithmetic, done over and over, more times than you can picture.

Hold exactly that, "training is mountains of small sums, repeated", and the rest of this chapter falls into place. Because the machine we are about to meet is not smarter than the ones before it. It is just built to do mountains of small sums all at once.

Citations