AlphaGo Creator Argues LLMs Are All Intuition, No Reasoning

Thore Graepel warns that large language models lack the search-based deliberation that made AlphaGo's gameplay creative and reliable

edit
By LineZotpaper
Published
Read Time3 min
The artificial intelligence that displayed what looked like genuine creativity against a Go grandmaster a decade ago relied on a form of reasoning that today's most celebrated AI systems fundamentally lack, according to Thore Graepel, a researcher who helped build AlphaGo. In an essay for MIT Technology Review, Graepel argues that large language models (LLMs) operate as powerful but narrowly 'intuitive' engines—System 1 in Daniel Kahneman's famous framework—while lacking the deliberative System 2 search that gave AlphaGo its creative spark. Without genuine reasoning capabilities, he warns, these models cannot be fully trusted for fields like science and medicine.

A decade ago, artificial intelligence demonstrated what looked like genuine creativity and intuition. The move AlphaGo played against Go champion Lee Sedol—Move 37—was so unorthodox that some commentators initially suspected a programming glitch. Lee Sedol himself later said he believed AlphaGo was creative.

But according to Thore Graepel, who helped build the system, this was not a flash of machine intuition. It was the product of a reasoning process that today's leading AI models do not possess.

Graepel grounds his argument in the psychological framework popularized by Daniel Kahneman: System 1 is fast, intuitive, and effortless, while System 2 is slow, deliberate, and step-by-step. AlphaGo, Graepel explains, was the first striking machine analogue of this split. Its policy network provided intuitive hunches, but its search machinery constructed and analyzed game trees of possible futures, explicitly testing those hunches. Intuition alone would not have chosen Move 37, and brute-force search would have been overwhelmed by the complexity of Go.

In contrast, a large language model operates purely through System 1. It picks the next token, over and over. This process is remarkably fluent, making for a powerful pattern-completion engine. However, Graepel warns that fluency is not the same as usefulness or truth. The field's response to this limitation has been chain-of-thought prompting, where the model generates intermediate steps before arriving at an answer. This has produced real gains, particularly in mathematics and coding.

Nevertheless, Graepel argues that chain-of-thought is not the same as AlphaGo's search. AlphaGo's search was grounded in the rules of Go; it evaluated the actual consequences of moves. An LLM generating a chain of reasoning is still just generating plausible-sounding text, step by step. It has no underlying mechanism for testing its own conclusions against a ground truth model of the world. It is a simulation of deliberation, not deliberation itself.

If the goal is to build AI that produces trustworthy results and truly novel insights in fields like science and medicine, Graepel concludes that systems must be equipped with genuine reasoning capabilities: the ability to look ahead, weigh evidence, and test hypotheses in a structured, deliberative way, rather than simply betting on likely patterns.

§

Analysis

Why This Matters

  • The essay challenges industry claims about 'reasoning' models, drawing a clear engineering line between associative prediction and grounded deliberation.
  • It provides a concrete framework (System 1 vs System 2) for evaluating AI capabilities, which informs safety and reliability expectations.
  • For critical applications like medical diagnosis or scientific research, the distinction determines whether an AI is a useful assistant or a potential liability.

Background

The concept of fast and slow thinking was popularized by Daniel Kahneman in his 2011 book Thinking, Fast and Slow. In AI, Deep Blue (1997) used brute-force search over chess positions. A decade later, Go remained unsolvable by search alone until DeepMind's AlphaGo combined learned intuition with search, succeeding in 2016. The recent wave of AI excitement, following the release of ChatGPT in late 2022, is built on large language models that predict text tokens without an explicit internal search mechanism, leading to debates over whether they can 'reason' or merely mimic reasoning.

Key Perspectives

Thore Graepel (Researcher): LLMs are purely instinctive pattern-matchers. They cannot check their own work internally against a ground truth. True reasoning requires an architecture that explicitly constructs and evaluates counterfactual futures. Proponents of LLM Reasoning: The prevailing view in the AI industry holds that chain-of-thought prompting grants models a form of temporal search, and that performance on reasoning benchmarks proves the capability is emerging or that the distinction between the two architectures is narrowing. Critics/Skeptics: Graepel's view aligns with a growing chorus of criticism that likens LLMs to sophisticated parrots. The concern is that industry hype glosses over fundamental architectural limitations, leading to unsafe deployment in high-stakes environments without proper human oversight.

What to Watch

  • Whether major AI labs release models that explicitly integrate search, world models, or other deliberative mechanisms alongside the token predictor.
  • Adoption of chain-of-thought in safety-critical settings and emerging failure modes in those deployments.
  • Academic research into neuro-symbolic AI or System 2 architectures that attempt to bridge the gap Graepel identifies.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.