Loading slide
Loading contents...
This module followed the move from symbols to learned positions.
A token ID is only an address. The stored embedding at that address is a learned list of numbers. Training places tokens used in similar ways near one another.
What geometry cannot do on its own is read a sentence. Every position is settled before the sentence arrives, so bank sets off from the same spot whether the river or the money is meant. The words around it have to lean in and adjust it, which means working out which of them matter.
The next module builds the mechanism that lets surrounding words make that adjustment.
One number cannot show all the ways two words can be alike.
A longer list of numbers gives each word a place near words used in similar ways.
But that place cannot yet change to fit one particular sentence.