Loading slide

Loading contents...

[█████░░░░░░░░░░░░░][████████░░░░░░░░░░░░░░░░░░░░]10 / 37
<back>next

One, two, three

The argument about scale was not settled by argument. One lab built the same thing three times, each much larger than the last, and published what happened.

2018. The first GPT, the writing half of the transformer, trained on books. 117 million parameters. It was good at the tasks it was fine-tuned for and unremarkable otherwise.

2019. GPT-2, the same design, thirteen times larger at 1.5 billion parameters, trained on web pages rather than books. It wrote paragraphs that held together, and OpenAI did something unusual with it. They withheld the full model, saying it could be misused to generate misleading text at volume. Smaller versions went out first, and the complete one followed nine months later. Whether that was caution or publicity was argued about at the time.

2020. GPT-3. Another hundredfold, to 175 billion parameters.

Nothing fundamental changed across those three. Same architecture, same training task, more of everything. That was the finding.

# citations(3)↓
  1. [1]cdn.openai.com
  2. [2]openai.com
  3. [3]arxiv.org