Loading slide
The speedup didn't just make the old work faster. It opened up work that had never been possible at all.
Before GPU training, you could train a network on tens of thousands of examples. Now you could train on millions. You could build networks dozens of layers deep instead of two or three. You could run in an afternoon an experiment that would once have taken months, which meant you could try ten ideas in the time it used to take to try one.
That last part matters more than it sounds. Research moves at the speed of how fast you can test an idea and see if it worked. Collapse the training time from months to hours, and the whole field speeds up with it. More people, trying more things, learning faster.
So the obstacles from the last chapter fell away. The deep networks that had been too slow to train were suddenly practical. And the moment researchers could actually build them at size, they noticed something nobody had predicted.