Loading slide

Loading contents...

[█████████░░░░░░░░░][██████████████░░░░░░░░░░░░░░]12 / 24
<back>next

Four ideas at once

In March 2016, DeepMind's system, AlphaGo, beat Lee Sedol, one of the strongest Go players alive, four games to one.

AlphaGo combined several established ideas in a new system. Four of them matter for the picture here.

Neural networks looked at a board and produced two judgements: which moves are worth considering, and who is probably winning. That is the unwritable knowledge from the last slide, learned from examples rather than stated as rules.

Search still looked ahead through possible continuations, exactly as chess programs had for decades.

Reinforcement learning improved those judgements through games played and won or lost, not through positions labelled by people.

Self-play supplied the opponent, which is the next slide.

Imagine AlphaGo has about 250 possible moves. Its networks make two quick guesses: which few moves look promising, and who seems ahead. Search then looks several moves into the future for those promising moves. It does not have to follow every possible path.

The two parts cover each other's weakness. The networks narrow the choices, but their guesses can be wrong. Search tests those guesses by checking what could happen next. Without the networks, search has too many paths to follow. Without search, AlphaGo has to trust its first guess.

# citations(1)↓
  1. [1]nature.com