Loading slide

Loading contents...

[████████░░░░░░░░░░][████████████░░░░░░░░░░░░░░░░]8 / 18
<back>next

One weight, many jobs

The delete test worked from one direction. Pick a fact, follow it down to the numbers holding it, find there are millions. Now turn it round and follow one number upward.

Take a single weight, somewhere in the middle of the model. What is it for?

The honest answer is that it is not for anything. It is not the croissant weight. It has some small say in croissants, and some small say in French geography, and some small say in the grammar of possessives, and in medical dosages, and in the tone of a polite refusal, and in thousands of other things that have nothing to do with each other. It contributes a little to all of them and is dedicated to none.

So the relationship runs both ways at once. Every concept is spread across many weights, and every weight takes part in many concepts. That second half is the one people miss, and it is why the delete test failed the way it did. There was nothing to cut out because there was no piece that belonged to Paris alone.

Which raises an obvious question. Why would a machine store things so awkwardly, when a filing cabinet is right there?

Because it has no choice. A model needs to represent far more distinct things than it has numbers to spare, and the only way to fit more ideas than you have places is to let them share the places. Researchers call this superposition, borrowing the word from physics, where it describes several waves occupying the same water at the same time.

Sharing has a cost, and the cost is exactly the mess you have been reading about. Concepts bleed into each other. Nothing can be isolated, inspected, or removed on its own. That is the price of packing more meaning into a fixed set of numbers than the numbers could otherwise hold, and the model pays it because the alternative is knowing less.

# citations(1)↓
  1. [1]arxiv.org