AlphaGo’s Move 37 renews debate over how AI reasons
Former AlphaGo team member Thore Graepel argues that search over possible futures gave the system a form of deliberation that today’s large language models lack.

AlphaGo’s unlikely Move 37 against Go champion Lee Sedol in 2016 was not, Thore Graepel argues, a flash of machine intuition. In an essay published by MIT Technology Review, the former AlphaGo team member says the move emerged from a combination of neural network evaluations and search through possible future moves.
AlphaGo defeated Lee 4–1 in their five-game match. Its policy network estimated which moves a strong human might make, while search machinery explored a game tree of possible variations. Graepel says that process allowed the system to choose a move its network considered highly unlikely for a human expert.
He contrasts that approach with large language models, which generate text by predicting one token after another. Although models can produce chains of thought, Graepel argues that these steps remain part of the same prediction process rather than a separate, independently auditable reasoning mechanism.
Graepel says current models generally lack a persistent, inspectable record of their hypotheses, confidence and evidence. He also argues that their knowledge is not clearly separated from the processes used to manipulate it, and that a chain of thought may not faithfully show how a model reached its answer.
He proposes systems that keep an explicit record of what they know, doubt or have ruled out, then revise those beliefs when evidence supports a change. Such systems could choose calculations, questions or experiments based on whether they reduce uncertainty, he writes.
Graepel acknowledges that real-world problems are less constrained than board games: the situation may be only partly known, and actions can have uncertain consequences. His proposal is an argument for a different approach to machine reasoning, not evidence of consensus among AI researchers.

