Loading slide

Loading contents...

[██████░░░░░░░░░░░░][█████████░░░░░░░░░░░░░░░░░░░]8 / 25
<back>next

The octopus opened the ievable

Take those percentages at face value and something strange follows.

If ievable really holds a tenth of one percent, that is roughly one in a thousand. Ask a model to continue The octopus opened the ... a thousand times and, once, it should write:

The octopus opened the ievable.

That does not happen. So either the machine has some way of ruling nonsense out, or the numbers are wrong somewhere.

The numbers are wrong, in the direction of being too kind. They were drawn to be readable rather than measured from a real model, and two things make the real bottom of the list far thinner.

The first is that four tokens were standing in for the whole vocabulary. Divide 100 percent between four candidates and each gets a generous share. Divide it between tens of thousands, with almost all of them scoring far below the plausible ones, and the bottom of the list is crushed to figures like one in eight million rather than one in a thousand. That long thin bottom end has a name worth knowing, the tail of the distribution.

The second is that most services do not draw from the full list at all. They cut the tail off first, keeping the top handful of candidates and discarding the rest, so a token with one chance in eight million is given no chance at all.

So the absurd sentence stays hypothetical, and the octopus keeps opening jars.

But look at what saved it. Not the model. Sharing the odds across a vast vocabulary is arithmetic, and cutting off the bottom of the list is a rule someone applied from outside. At no point did anything recognise that ievable would be gibberish in that sentence.

Had the draw landed there, the model would have written it without hesitating, and gone on to the next token as though nothing had happened.

# citations(1)↓
  1. [1]arxiv.org