ImageNet showed what could happen when trainable depth, parallel hardware, a large human-labelled dataset, and a shared test came together.

Look closely at what made that possible, though. Every training picture arrived with a human-written answer attached. school bus. coffee mug. The challenge kept test answers hidden, but training still depended on millions of supplied labels.

That arrangement has a limit, and it is not the cost of the labelling.

Some problems do not offer one supplied target for every decision. A Go position may have several strong moves, and a driving situation may allow several safe steering choices. Human examples can still help, but they do not provide a unique correct action for every possible moment.

What can be done, for problems like those, is to try something and see how it turns out.

That is a different way to learn, and the next chapter builds it.

The whole chapter, simply

Researchers could not fairly compare picture-reading machines when everyone trained and tested them differently.

A huge shared picture collection and one shared contest showed which ideas worked best.

In 2012, one system won by a huge margin and showed that several older ideas finally worked well together.