Loading slide

Loading contents...

[███████████████░░░][███████████████████████░░░░░]31 / 37
<back>next

Confident and wrong

Agreement is one thing post-training rewards. Fluency is another, and it causes a harder problem.

The model has one way of producing text, and it uses that way whether the pattern behind a sentence is solid or threadbare. A well-supported fact and a confident invention come out of the same generator, in the same register, at the same speed.

So the wording carries no information about how sound the content is. There is no wobble in the voice, no hesitation before the invented citation, nothing in the prose that marks the difference. When such a sentence is false, people call it a hallucination.

Post-training can reduce this. A model can be trained to hedge, to say it does not know, to refuse. But the hedging is itself a learned pattern, chosen because raters liked it in similar situations, not because the model checked anything. Training it to say "I am not certain" more often does not connect those words to whether it is in fact certain.

That is what makes it dangerous rather than merely annoying. Every other unreliable source you deal with gives you something to go on. A person who is guessing usually sounds like it. A badly made website looks like one. Here the presentation is uniformly excellent, so the ordinary way of judging what to trust does not work, and it does not work precisely on the occasions when it matters most.

The final chapter of this module looks at what the weights actually contain, and at why outside sources, tools, and human checking still matter.

# citations(1)↓
  1. [1]arxiv.org