Where the weights live
That second step, the desk work, is worth one more moment, because of how much of the transformer it turns out to be.
Remember what a weight is. One learned number, sitting on one connection, deciding how much of a signal gets through. A trained model is nothing but an enormous pile of them.
Attention is the famous half. The desk work is the bigger one. Roughly two thirds of all the weights in a typical transformer sit in the desk work, not in attention at all.
Which matters for a claim that is coming. When you hear that a model's knowledge is spread across billions of numbers, this is largely where those numbers physically are. Attention decides which words get to talk to each other. The desk work is where whatever the model has absorbed about the world actually sits.