Loading slide
You just recalled what a weight is: a dial for how much to listen to something. Attention is that dial, turned word by word. That is the mechanism, so let's see it concretely.
The scores from the last slide, how much each word matters to "it," become a set of weights, one per word. A big weight on "trophy," a big weight on "big," tiny weights on "the" and "because." Now the machine does the obvious thing with weights: it takes a blend.
Think of mixing paint. Every word in the sentence carries its own splash of colour (its
That blend is attention. The fancy word "attention" hides a plain action: a weighted average of all the words, where the weights say who matters right now. High weight, large pour. Low weight, barely a drop.
And notice what this buys us, looking back at the last chapter. "Trophy" pours into "it" in a single step, directly, no matter how many words sit between them. There is no chain to fade along, because the relevant word is poured straight into the blend. The distance problem is simply gone.
One question remains, and it is the interesting one. We have been assuming the machine already knows that "trophy" matters to "it." But where do those weights actually come from? How does it score one word against another in the first place? That is the next slide, and it is the cleverest part.