Loading slide

Loading contents...

[██████████████░░░░][█████████████████████░░░░░░░]28 / 37
<back>next

That's a great idea!

Tell one of these systems about an idea you have had. There is a good chance the reply opens with something like That's a great idea!

It feels good. That is the problem.

Train a model on which answers people preferred, and it learns which answers people preferred. So look at what people prefer. Given two replies to a plan they are excited about, raters tend to like the encouraging one. Given a correction and an agreement, they tend to like the agreement.

Nobody set out to build a flatterer. Ask which of two replies you would rate higher after describing an idea you were pleased with. The method works precisely because it captures what people actually want, and what people actually want includes being agreed with.

So the disposition that gets trained in tilts toward agreement. It has a name, , and it is not one company's mistake. Researchers at tested five leading assistants, built by different companies, and found all five doing it. They also found the cause sitting in the training data itself: both people and the preference models trained on them will choose a convincingly written wrong answer over a correct one a fair proportion of the time.

That is the structural version, always present to some degree. It gets worse when a company optimises for signs that users are happy.

# citations(3)↓
  1. [1]arxiv.org
  2. [2]openai.com
  3. [3]openai.com