Loading slide
Loading contents...
When a model answers you, it feels like it worked the answer out first and then wrote it down. That is how a person would do it.
There is no first.
Plenty happens before each token. The numbers move up through every layer of the tower, and that is what decides which token comes out. But those numbers are not a sentence. There is no draft in there, in English or any other language, waiting to be typed up.
So where can a thought be kept, if the model is going to need it two tokens later?
Only one place. The text.
Every token the model writes becomes part of what it reads before choosing the next one. A number it works out in the third sentence is still there in the tenth, because it is on the page. Nothing else survives.
That is why a model asked to answer immediately and a model asked to work through the steps can give different answers to the same question. It is not that one tried harder. The second one gave itself somewhere to put the intermediate results.