Smaller, and better liked
So pre-training supplies the knowledge and post-training supplies the manners. That makes it sound like the second one is decoration.
In 2022 OpenAI measured it. They took models of several sizes, ran them through fine-tuning and preference training, and called the results InstructGPT. Then they asked people to compare those answers against answers from the original GPT-3, which had none of that treatment.
A post-trained model with 1.3 billion parameters was preferred to a base model with 175 billion.
A hundred and thirty times smaller. Less of everything the scaling curve measures, less text, less computing time, fewer parameters by two orders of magnitude. And people liked its answers better.
This does not overturn the case for scale. The small model could only be shaped into something good because pre-training had put the knowledge there first, and running the same treatment on a bigger base gives you something better still.
What it settles is what post-training does. Pre-training decides what a model knows. Post-training decides how much of that you can get at.