Loading slide
The human example answers were written with confidence. Clear, authoritative, well-structured. That is what "good" looked like to the people writing them, and what raters kept preferring. So that is what the model learned to sound like: confident, fluent, certain.
Here is the trap. The model learned to sound confident, full stop. Not confident when it has a basis and hesitant when it does not. Confident always. It picked up the style of sure-footed human writing, and sure-footed human writing almost never says "I am not certain." So the model rarely says it either. It produces a smooth, authoritative answer whether or not anything real sits underneath.
This is what people mean when they say a model hallucinates. It is not lying, because lying needs a sense of the truth to push against, and the model has none. It is doing precisely what it was trained to do: produce the most plausible-sounding continuation. Sometimes the most plausible-sounding continuation happens to be false. There is no inner gauge tracking what it really knows versus what it is inventing, so a fabricated citation rolls out with the same easy confidence as a true one.
This matters enough that we will return to it in this module's final chapter, where we look at what the weights actually contain. For now, hold the shape of it: the very training that made the model helpful and fluent is the same training that made it sound just as sure when it is wrong as when it is right. The polish is not evidence. It was the goal.