Loading slide
Loading contents...
In 2013, DeepMind published a paper about teaching one system to play old Atari games.
Look at the game. You see a ball, a paddle, and blocks. The machine received none of those names. It received the pixels on the screen, the controller buttons it could press, and the score.
It pressed a button. The screen changed. Sometimes the score changed too. If the score went up, the machine treated that choice as a good result.
After millions of attempts, the network inside the loop learned which buttons were more likely to lead to future points. This was deep reinforcement learning in action: pixels in, action out, score back.
The first paper covered seven games. A larger 2015 study tried the system on 49.
# did you know?
While learning Breakout, the system discovered a useful trick. It knocked a tunnel through one side of the brick wall, then let the ball bounce behind the wall and clear bricks from above.
Nobody wrote that strategy into the game instructions. It appeared because the move led to more points.