A reply is produced one token at a time. The model reads everything so far, scores every token in its vocabulary, one is drawn, and that token joins the text before the next round begins.
Nothing is drafted ahead. Nothing already written is revised. An early word narrows what can follow, and the model lives with it.
The draw is not always the top scorer. A setting called temperature decides how sharply the scores favour the leader, which is why the same question can produce two different answers, and why turning it too far produces nonsense.
There is no inner monologue behind any of this. What the model writes is the only place a thought can be kept, which is why writing the steps down changes what it can do, and why those steps are a claim to check rather than a report of what happened.
All of which leaves one thing unexplained. Every token that gets chosen is chosen because of a score, and those scores come from somewhere. What does the model actually know, where is that knowledge kept, and why can it be so confidently wrong about it?
That is the last chapter of this module.
The whole chapter, simply
The machine picks one small piece of text, adds it to what it has already written, and picks again.
There is no plan behind the answer, so a long confident reply is not proof that anything was worked out.