A decade ago, artificial intelligence demonstrated what looked like genuine creativity and intuition. The move AlphaGo played against Go champion Lee Sedol—Move 37—was so unorthodox that some commentators initially suspected a programming glitch. Lee Sedol himself later said he believed AlphaGo was creative.
But according to Thore Graepel, who helped build the system, this was not a flash of machine intuition. It was the product of a reasoning process that today's leading AI models do not possess.
Graepel grounds his argument in the psychological framework popularized by Daniel Kahneman: System 1 is fast, intuitive, and effortless, while System 2 is slow, deliberate, and step-by-step. AlphaGo, Graepel explains, was the first striking machine analogue of this split. Its policy network provided intuitive hunches, but its search machinery constructed and analyzed game trees of possible futures, explicitly testing those hunches. Intuition alone would not have chosen Move 37, and brute-force search would have been overwhelmed by the complexity of Go.
In contrast, a large language model operates purely through System 1. It picks the next token, over and over. This process is remarkably fluent, making for a powerful pattern-completion engine. However, Graepel warns that fluency is not the same as usefulness or truth. The field's response to this limitation has been chain-of-thought prompting, where the model generates intermediate steps before arriving at an answer. This has produced real gains, particularly in mathematics and coding.
Nevertheless, Graepel argues that chain-of-thought is not the same as AlphaGo's search. AlphaGo's search was grounded in the rules of Go; it evaluated the actual consequences of moves. An LLM generating a chain of reasoning is still just generating plausible-sounding text, step by step. It has no underlying mechanism for testing its own conclusions against a ground truth model of the world. It is a simulation of deliberation, not deliberation itself.
If the goal is to build AI that produces trustworthy results and truly novel insights in fields like science and medicine, Graepel concludes that systems must be equipped with genuine reasoning capabilities: the ability to look ahead, weigh evidence, and test hypotheses in a structured, deliberative way, rather than simply betting on likely patterns.