Loading slide
Loading contents...
That exact difficulty had been solved once already, on a problem with nothing to do with language.
In 2017, researchers at OpenAI and DeepMind, working together, took on a problem in reinforcement learning. That is the kind of training where a machine is given no examples of correct behaviour at all. It tries something, receives a number saying how well that went, and adjusts. Then again, thousands of times.
Their trouble was writing the number. Some goals score themselves, like points in a game. Many do not, and a machine given a badly written score will pursue the score rather than the thing you wanted.
So they stopped writing scores. They showed a person two short recordings of the machine attempting the task and asked one question. Which of these two was better?
Nothing else. No explanation of why, no marks out of ten. Enough of those comparisons and the machine worked out what was being asked of it. It learned to control simulated robots this way, and to play Atari games, from roughly an hour of a person's time, with humans looking at under one percent of what it did.
The machine could not be told what good looked like. It could be shown, one pair at a time, which of two attempts came closer.
Language is full of tasks with exactly that shape.