Loading slide

Tokens

  1. 01Tokens
  2. 02Memory card
  3. 03The model can't read
  4. 04Not words. Tokens.
  5. 05The emoji trick
  6. 06Try it yourself
  7. 07How many r's in strawberry?
  8. 08What the model actually sees
  9. 09Some languages cost more
  10. 10Why "cat" and " cat" differ
  11. 11Why not millions of tokens?
  12. 12So why not just use bytes?
  13. 13Context windows are measured in tokens
  14. 14Reinforce your understanding
  15. 15Question: Why not just use words?
  16. 16Question: Weird splits
  17. 17Quiz: answer
  18. 18What to carry out of this chapter
  19. 19Want to go deeper?
3 / 18
BackNext

The model can't read

Here is something that surprises most people.

A language model has never read a single word. Not one. It cannot. Everything a computer touches has to become numbers first. That has been true since the very first module, and language is no exception.

The last chapter was about what kind of numbers could carry meaning. But hiding underneath that was a more basic question, and we skipped over it. Before you can hand out numbers at all, you have to decide what gets one. What are the pieces? Whole words? Single letters? Something else?

Neither obvious answer works. Give every letter its own number and words dissolve: "dog" becomes d, o, g, three pieces that have nothing to do with barking. Give every whole word its own number and the list never ends: every name, every typo, every word someone invents next year would need a new entry.

The answer sits in between, and it has a name: tokenization.

Citations(2)↓
  1. 1. huggingface.co
  2. 2. arxiv.org