Loading slide

Loading contents...

[░░░░░░░░░░░░░░░░░░][░░░░░░░░░░░░░░░░░░░░░░░░░░░░]0 / 24
<back>next

Module 5 Chapter 4

Learning from Consequences

Nobody teaches bicycle balance by providing the correct handlebar angle for every possible moment. You act, the world responds, and the consequence gives evidence about what to try next.

Everything so far has needed an answer key. Here is the input, here is what you should have said, adjust. That works whenever somebody already knows the answers, and it stops working the moment nobody does.

Learning from consequences asks for much less. No correct answer, just a signal afterwards saying that went well or that went badly. It is a far weaker form of instruction, and it is also the form most of life actually offers.

The weakness is the hard part. A reward tells you the outcome, not which of the two hundred things you did was responsible for it, and the move that lost you the game may have been played twenty minutes before you lost. Something has to work backwards from a single verdict to a long history of choices, most of which were fine.

Machines that learn this way have produced some of the strangest results in the field: play that experts called creative, moves nobody had considered in centuries of study, arrived at by something with no concept of the game beyond a number going up. What that says about the thing doing it is less obvious than either the excitement or the dismissal suggests.

In this chapter

  • Two ways to show a mistakehow answer keys and rewards provide different feedback
  • Learning by actinghow observation, action, consequence, and update form a loop
  • Which action caused the resultwhy delayed rewards are hard to trace
  • Ideas meeting across decadeshow Sutton, Tesauro, and DeepMind joined feedback to networks
  • Surprising moves from both sideswhat Move 37 and Move 78 looked like on the board
  • Beyond games and rewardswhy AlphaFold belongs to deep learning but not reinforcement learning
# citations