Loading slide

Loading contents...

[█████░░░░░░░░░░░░░][████████░░░░░░░░░░░░░░░░░░░░]5 / 18
<back>next

Try it yourself

Five things the model knows, each one a bar. Paris is the one you are trying to remove. The others are neighbours, and they are ordered by how much they lean on the same weights: France most, a Texas town of the same name least.

Turn the erase strength up and watch all five.

Gently, and nothing much happens anywhere, Paris included. Push harder and Paris starts to go, but France goes with it, then the Seine, then the rest. Push far enough to empty the Paris bar and you have flattened the row.

There is no setting that empties one bar and leaves the others full. Not because the control is too crude, but because the thing you are reaching for was never separate from the things beside it.

The bars are drawn to be read rather than measured from a real model. What is true to life is the shape: damage falling off with overlap, and no strength that separates a concept from its neighbours.

So the delete test fails, and it fails for a reason worth stating plainly. There is no clean cut, because there was never a clean thing to cut.

This is one of the live problems in the field. When a model has learned something that should not be there, private information, a copyrighted text, something harmful, there is no delete key. Removing one specific thing from a trained model without damaging the rest remains difficult and imperfect.

So the usual answer is not removal at all. It is a filter around the model that catches the output, or further training that steers away from the subject. Both are closer to teaching a habit of not mentioning something than to taking it out.

# citations(2)↓
  1. [1]arxiv.org
  2. [2]arxiv.org