Several weeks became about one day
That 2004 network was small, and getting it onto the chip took a disguise. With CUDA in hand, a different team could send ordinary number work straight to the GPU, and go much bigger.
They did, in 2009, with the 100 million parameter network from earlier in this chapter. Researchers estimated that a CPU run would take several weeks. Their GPU system completed the work in about one day.
Across the tested network setups, the GPU ran about 12 to 72 times faster than the CPU code they compared it with. The exact speedup depended on the job and hardware.
The GPU did not make the network smaller or skip the learning loop. It divided suitable calculations among many workers. Training on a very large scale became something researchers could actually complete, inspect, and repeat.