Loading slide

A wine it has never tasted

Now a bottle that was never in the training set. A cabernet, strongly plummy, with a touch of apricot.

It has no memory to search and no answer to look up. It does what it has always done. Take the four numbers, multiply each by its weight, add them up.

The guess comes out at 34 dollars 66. The wine costs 34 dollars 80.

Under a dollar out, on a wine it has never met.

Where does that 34 dollars 80 come from? The same place every price in this chapter came from. Every bottle here, including this one, was priced by one fixed rule: 14 dollars for a unit of cabernet, 6 for viognier, 16 for plum, 40 for apricot. Nobody ever showed the model that rule. It only ever saw bottles and prices.

Which is what makes the four weights worth a second look. They finished at about 14.2, 6.5, 15.8 and 39.2. Set those beside 14, 6, 16 and 40. The model did not approximate one wine's price. It reconstructed the rule that priced all of them, from nothing but eight wrong guesses and their corrections.

Now look at what is not there. Nowhere in this model is there a record of the eight bottles, their prices, or anything that happened during training. All of that was thrown away.

What survived is four numbers. Not a shrunken copy of the training set, and not a list with the details rubbed off. Four numbers that hold what all eight bottles had in common, and nothing about any bottle in particular.

It is easy to assume knowing must mean storing. It does not. You can recognise a friend's face in a crowd without holding a photograph of it anywhere. Nothing in your head is a picture of them. What you have instead is a pattern, built from every time you have seen that face, that fires when the right features arrive together.

Animal brains appear to work something like this. A memory is not filed in one cell that could be found and read out. It is spread across the strengths of many connections at once, which is why the same brain can recognise a face it has not seen in twenty years, in bad light, at an angle it never saw before. Storing knowledge this way, as a pattern across many connections rather than an entry in one place, is called a distributed representation, and it is what What the Weights Contain is about.

This model is doing the same trick with four connections instead of billions. Knowledge lives in the strengths, not in a stored list.

# citations(1)↓
  1. [1]mitpress.mit.edu

Loading contents...

[█████████████░░░░░][████████████████████░░░░░░░░]30 / 41
<back>next