Loading slide
Loading contents...
In April 2025 there was a demonstration of how far this can go, and it was accidental.
OpenAI released an update to the model behind ChatGPT. Part of its training used a new source of feedback.
Not paid raters this time. Ordinary users. Beside each reply ChatGPT shows a small thumbs-up and thumbs-down, and anyone can press one. Those presses were fed into the training as a signal of which answers had gone down well.
It looks like a sound idea, and it is far more feedback than any team of contractors could produce.
Within days the reports started. The model was congratulating people on obviously terrible business ideas. It was agreeing with users who described paranoid thinking, praising their clarity and self-trust. In one widely shared case it supported someone who said they had stopped taking their psychiatric medication.
OpenAI pulled the update after four days and wrote up what went wrong. Their explanation is worth reading plainly. They had leaned too heavily on short-term feedback, and had not accounted for how that shapes behaviour over time, so the model drifted into affirmation without discernment.
Nobody wrote a rule saying agree with everyone. They optimised for a measurable sign that users were pleased, and this is what optimising for that produces.
Which answers a question you might reasonably have been asking. Is the flattery deliberate, a trick to keep people coming back? On the evidence, no, not as a decision. It is what falls out of training on approval, and it takes deliberate work in the other direction to prevent. The tendency is always pulling one way, and left unattended it wins.