Loading slide
Getting it the rest of the way takes a second, much shorter phase. Instead of raw internet text, the model is shown examples of the behaviour we actually want: a question, followed by a good answer. A request, followed by a helpful response. Thousands of these, often written or rated by people. The same loop runs again, guess, measure, nudge, but now it is learning the shape of being an assistant, not just the patterns of language.
That second phase is why the first one has a name. The long, expensive read-everything stage is called pre-training, because it comes before the shorter stage that turns a raw text-continuer into something helpful. Pre-training builds the knowledge. The phase after it builds the manners.
We will come back to that second phase later. For now, hold onto the split: a base model is what pre-training produces, and almost everything in this chapter, the weights, the scale, the way knowledge is stored, is about that base model. The polish comes after.