Loading slide
Loading contents...
A model's knowledge is not filed. It is spread across billions of numbers, none of which means anything alone, and each of which takes part in thousands of unrelated things. That is a distributed representation, and it is what learning from examples produces rather than a decision anyone made.
Three things follow, and they are the reason this chapter exists.
Nothing can be cleanly removed, because no part of the model belongs to one fact alone. Nothing can be simply read, because nothing in there was written to be read, though researchers can now recover some of it with considerable effort. And the model cannot look up whether it knows something, because there is no index to look in. It answers either way, in the same voice.
And two things it cannot notice about itself, both for the same reason. Where its knowledge stops, and when. A thing that is not there produces no signal saying so.
That closes the machine itself. Across this module it grew from a research design into something enormous, learned to answer rather than merely continue, and turned out to write one token at a time with no plan behind it and no filing cabinet inside it.
Which leaves the part nobody has examined yet. The model is frozen, knows nothing of you, and forgets each conversation completely. So when a chat remembers what you said five minutes ago, something else is doing that work.
Everything the machine knows is a pattern smeared across billions of numbers, and no single one of them holds anything.
That is why nothing in it can be found, removed, or checked, including by the machine itself.