Loading slide
This chapter is about a machine that made training fast. Before we meet it, lock in the one thing it speeds up, because the whole chapter rests on it.
Recall what training a network actually is, underneath. The network makes a guess, sees how wrong it was, and nudges its dials a little. Then it does it again. And again. Each nudge is a pile of small, simple sums, multiply these numbers, add them up, adjust. Not hard sums. Just an enormous number of them, repeated billions upon billions of times until the dials settle.
So training is not one clever, complicated calculation. It is the same tiny arithmetic, done over and over, more times than you can picture.
Hold exactly that, "training is mountains of small sums, repeated", and the rest of this chapter falls into place. Because the machine we are about to meet is not smarter than the ones before it. It is just built to do mountains of small sums all at once.