Loading slide

Loading contents...

# module complete

Next: 08: The Model in the Machine

<back>start module 08:report

Google set out to translate, and built something more general than the job it was for.

Start with the tower. On every floor a word does two things in turn. It reaches out and gathers from the other words, then it goes away and is worked over alone. Repeat that on floor after floor and a word arrives generic and leaves fitted to the sentence it is in. Attention is only the first of those two moves, and the second holds most of the model's numbers.

Comparing every word at once threw away word order, so order is stamped onto each word before it enters. The sentence is no longer a sequence to the machine. It is a heap of words that each know where they sat.

Then the two questions that turned one design into an industry. What may a word see? Look both ways and you get a reader, good at understanding a sentence whole. Look only backward and you get a writer, able to go on producing the next word indefinitely.

And where do the answers come from? From the text itself. Cover a word, and the sentence has just set a question and supplied the answer. That is what made a large fraction of everything ever written into training material, and it is why one expensive model can be trained once and cheaply adapted many times.

The next module asks what changed as the models, the datasets, and the training budgets all grew, including which of the resulting claims are still disputed.

The whole chapter, simply

One machine reads by letting every word gather from every other word, then think on its own, over and over.

Text can mark its own homework, so the machine could learn from far more writing than people could ever label.

# citations(3)↓
  1. [1]arxiv.org
  2. [2]cdn.openai.com
  3. [3]arxiv.org