Act, consequence, adjust
Strip the bicycle down to its shape and a loop appears.
The learner is in some situation. It takes an action. The world responds, and somewhere in that response is a signal about whether things got better or worse. The learner adjusts what it will do next time it meets a situation like that one. Then it is in a new situation, and the loop runs again.
Three things make this loop harder than it looks.
The first is that the signal can arrive late. Lean wrong and you fall two seconds later. Which of the hundred small adjustments in those two seconds was the culprit? The consequence names the outcome, not the cause.
The second is that the learner has to try things to find out anything. An answer key can be studied. Consequences have to be provoked. A learner that only ever repeats its current best guess never discovers the better move sitting one step away, so some willingness to try something worse is part of the arrangement, not a flaw in it.
Modules ago, a machine searching a maze hit the same wall. It always stepped towards the goal, so it never took the route that began by heading away, the way the road to the next bridge does when the near one is closed. That machine could at least see the map. This one cannot. It has no way of knowing what an untried action is worth until it spends a real attempt finding out.
The third is that there is no ceiling supplied. With a labelled example, matching the label is the goal and the goal is known. Here the learner only knows the scores it has seen so far. Something better may always be out there.