Loading slide
We spent the whole last module on what training does. Here is what it costs: relentless arithmetic.
Every training example gets passed through every layer. Every weight gets multiplied by something. Every error gets traced back through the whole network, and every weight gets nudged. Then the next example. Then the next. Across millions of examples, across millions of weights, repeated thousands of times over. The guess-measure-nudge loop from the last module, run an almost unimaginable number of times.
The chip at the center of an ordinary computer, the
Training doesn't need that. Most of its calculations don't depend on each other at all. In principle, you could run all of them
A CPU does one thing at a time, very fast. Training needed something else entirely.