Loading slide
Loading contents...
The first is the stranger of the two, and it is easiest to see by comparison.
Ask a database for something it does not hold. It goes to where the record would be, finds nothing there, and tells you so. That is not cleverness. It is a consequence of the filing: a thing with an address either has something at it or does not, and looking is a real operation with two possible results.
Now ask the model something it never learned. There is nowhere to go and look. No index to consult, no shelf to find empty. The question arrives, the weights do what they always do, and text comes out.
So the model cannot come up empty, because coming up empty is not one of the things it can do. Where you were hoping for I have no record of that, the machinery has no step at which such an answer could be produced.
Think of a doctor who has practised in one country her whole career, seeing one population of patients. Within that world she may be superb. Ask about a disease common somewhere she has never worked, and she may answer with total confidence and be wrong. Not carelessness. That pattern was never in her experience, and the gap does not announce itself from the inside. She cannot feel the edge of what she has seen.
The model is in that position about everything at once. Its weights hold what training gave it a chance to learn, thinly in some places and not at all in others, and nothing in it marks where those places begin.
A trained model can be made to say it does not know. Post-training can teach that, and you have met how: raters preferred the hedge, so the hedge became a likely continuation in situations that look like this one. What it cannot do is check first. The words I am not sure are produced the same way every other phrase is produced, which means they can turn up when the model is right and fail to turn up when it is inventing.