Loading slide

Loading contents...

[████████░░░░░░░░░░][████████████░░░░░░░░░░░░░░░░]10 / 24
<back>next

The generality was the point

The central result was reuse, not any one score. Programs had beaten humans at individual games for decades.

DeepMind used the same learning algorithm, network architecture, and training settings across all 49 games.

Before training, each game began with the same kind of system. Training produced different learned weights for Breakout, Space Invaders, and Enduro.

Compare that with the systems from the long winter. Deep Blue used hand-built chess knowledge and search for one game. Expert systems contained rules written for one field. The Atari work reused one learning setup across dozens of games, although it still trained a separate set of weights for each one.

Here was one training setup applied to dozens of games using pixels, available actions, and changes in score.

That does not make it broadly intelligent. It played Atari games, and it was poor at the ones needing a plan that pays off much later. The result showed a limited kind of reuse: one learning method could acquire several narrow skills without game-specific rules being written for each one.

It also began a remarkable run for DeepMind, founded in 2010 by Demis Hassabis, Shane Legg, and Mustafa Suleyman. Atari was followed by AlphaGo and later AlphaFold. For much of the 2010s, those systems made DeepMind one of the laboratories setting the public pace of AI research.

The laboratory's long bet was larger than winning games. Games supplied clear rules and feedback. The aim was to build agents that could learn about a world, act in it, and eventually help with scientific problems.

# citations(2)↓
  1. [1]nature.com
  2. [2]deepmind.google