Loading slide
Through the 1990s and into the 2000s, neural networks existed, but they stayed shallow. One or two hidden layers at most. Anything deeper was unreliable.
Other approaches dominated. Statistical methods and classical machine learning algorithms consistently beat neural networks on the benchmarks that mattered. Neural networks weren't the clear answer. They were a minority bet.
A small group of researchers kept working on them anyway. Geoffrey Hinton at Toronto, who you met back when backpropagation was first invented.
The problem wasn't that the ideas were wrong. The gradient fixes, ReLU, better initialization, normalization, were all within reach. The deeper issue was that none of it mattered much at the scale available then. The networks were too small. The datasets were too small. The computers were too slow.
The ideas needed the world to catch up. It was starting to.
Did you know?