Loading slide
Loading contents...
A language model runs on known mathematical operations. Its developers specify the architecture and the process used to train it.
Researchers can inspect its activity. In specific cases, they can trace mechanisms that helped produce an answer.
One recent circuit-tracing method still captures only a fraction of the computation, even for short prompts. Other work can find circuits for particular tasks, but those explanations do not yet carry cleanly across the whole model.
The gap is specific. We know which calculations the machine performs. We do not yet know how all of those calculations combine into each behaviour.
That tells us nothing by itself about whether the machine thinks. It tells us where the explanation currently runs out.