Loading slide
In 2020, OpenAI released GPT-3. At the time it was the largest language model ever trained: 175 billion parameters, 175 billion of those little weights. To picture the leap, the models that came just before it were a fraction of the size. This was not the next step up a staircase. It was a jump to a different floor.
The reactions split the room. People who got early access reported that, from nothing but a short prompt and with no special training for the task, it could write convincing prose, answer obscure questions, translate between languages, write code, and imitate the style of a particular author. The emergent abilities from the last slide, all showing up at once, in public.
It could also be confidently, strangely wrong. It would invent citations to papers that never existed. It would contradict itself a paragraph apart. It would answer a question about the present with total authority using facts that were already out of date. We will see later in this module that this is not a glitch bolted on by accident. It falls straight out of what the model is.
GPT-3 was the first time a wide audience met a model that felt genuinely surprising, not a narrow tool built for one job, but something with a broad and unpredictable range. Two years later, in 2022, ChatGPT put a friendlier version of the same idea in front of hundreds of millions of people, and the public arrival of AI began.