Loading slide
So the knowledge is in there, spread across the weights. The natural question is what it lets the model actually do. And the answer, at this stage, is narrower than you might expect.
After training on that ocean of internet text, the model has become very good at one specific thing: continuing text.
Show it the beginning of a Wikipedia article and it will write a plausible continuation. Show it the start of a news story and it will carry it forward. Show it a recipe and it will keep going.
This is called a base model. And a base model is not an assistant.
Ask it a question and it won't answer you. It will continue the text as if the question were the beginning of a document it had seen before. It might write more questions. It might write a quiz. It might write a forum thread with people debating the answer. It has no idea you wanted a direct response. The conversation on the right is what that actually looks like.
So the model knows an enormous amount and is still strangely unhelpful. Which raises an obvious question: how does it become the assistant you actually talk to?