Loading slide
Loading contents...
A model trained to give the answers people prefer raises an obvious question. Which people?
The InstructGPT paper answers it directly. The team hired 40 contractors, selected by a screening test. They were mostly English-speaking, living in the United States or Southeast Asia, hired through Upwork and Scale AI. They worked to written guidelines describing what a helpful, honest and harmless answer looks like.
Two things follow, and the researchers say both themselves.
The raters did not agree with each other. They matched about 73% of the time, meaning that on roughly one comparison in four, two people looking at the same pair of answers wanted different things.
And the guidelines they worked to were written by the researchers. The paper puts it plainly: they were aligning the model to their own preferences, as the people running the project.
None of that is a scandal. Somebody has to decide, the decisions were documented, and 40 people is not a small effort. But it is worth knowing what a model's manners actually are. Not a discovered fact about good conversation. The written preferences of a specific group of people, at a specific company, in a specific year, generalised by a machine into a disposition that millions now talk to.
When a model declines something, or hedges, or has a house style you can recognise, this is where that came from.