Why Today's AI Models Lack the Reasoning Power That Made AlphaGo Creative
A core architect of AlphaGo explains why large language models' chain-of-thought tricks fall short of genuine deliberation and what AI needs to solve hard scientific problems.
The Move That Changed Everything
In March 2016, a machine placed a stone on the fifth line of a Go board during a match in Seoul. Professional commentators initially suspected a bug. The move appeared to violate intuition so dramatically that it seemed impossible for any rational player, human or machine, to consider it. Yet AlphaGo went on to win that game and the overall match 4-1 against Lee Sedol, widely regarded as one of the finest Go players in history. Lee himself reconsidered his assumptions about what machines could do. "I thought AlphaGo was based on probability calculation and that it was merely a machine," he reflected. "But when I saw this move, I changed my mind."
That single move has since become a touchstone in discussions about machine creativity. But the common interpretation misses the point. The breakthrough was not a flash of artificial intuition. It was the product of a reasoning system that large language models do not possess, and the absence of that capability now constrains what contemporary AI can achieve in science, medicine, and other high-stakes domains.
Two Systems, One Decision
AlphaGo's architecture split the problem of playing Go into two complementary parts. A policy network, trained on thousands of expert games, learned to approximate what strong human players would do in a given position. This component functioned as a fast, associative heuristic, the kind of pattern recognition that feels effortless. On its own, that network assigned move 37 a probability of roughly one in 10,000 - essentially dismissing it as implausible.
The second component was a search mechanism. It explicitly constructed a game tree, branching out into thousands of possible futures, each representing a different sequence of moves and counter-moves. This machinery evaluated the long-term consequences of candidate moves, updating its judgements as it explored deeper. It was this deliberative process that identified move 37 as strategically sound, overriding the policy network's initial scepticism.
The parallel to human cognition is striking. Behavioural scientists distinguish between fast, gut-level judgement and slow, step-by-step deliberation. AlphaGo implemented both. Its neural networks supplied hunches about which moves looked promising and which board positions looked favourable. Its search supplied the deliberation, stress-testing those hunches against sequences of future play. Neither half could have succeeded alone. Intuition without search would have rejected move 37. Search without intuition would have drowned in the combinatorial explosion of possibilities.
What Chain of Thought Cannot Do
Large language models operate differently. They generate text one token at a time, each choice informed by the patterns absorbed during training. This is system-one thinking at scale: rapid, associative, and remarkably fluent across almost every topic represented in their training corpora. Shortly after ChatGPT's release, it became clear that fluency alone was insufficient for many tasks. The industry's response was chain-of-thought prompting, in which models generate intermediate reasoning steps before committing to a final answer. Performance on mathematics and coding problems improved measurably.
Yet this is not the same as equipping a model with a separate reasoning engine. Chain-of-thought outputs are still produced by the same next-token prediction mechanism, simply iterated for longer. At Opentechwire, we have tracked how frontier labs frame these capabilities, and the terminology often obscures a fundamental limitation: the intermediate steps are not the product of an independent deliberative process. They emerge from the same statistical machinery that generates everything else.
Three shortcomings distinguish this from reasoning in the sense a scientist would recognise. First, these models maintain no explicit, persistent epistemic state. There is no structured representation of hypotheses under consideration, confidence levels attached to competing explanations, or unresolved questions held in working memory. A reasoning system should be able to revise its beliefs systematically as new evidence arrives. Current models lack that ledger.
Second, knowledge and the manipulation of knowledge are entangled in the weights of the neural network. There is no clean separation between what the system believes and how it uses those beliefs to draw inferences. In AlphaGo, the game tree was an explicit data structure, separate from the neural networks that evaluated positions. In a large language model, everything is implicit, distributed across billions of parameters.
Third, research has shown that chain-of-thought outputs are sometimes confabulated. Models reach a conclusion through one route but report another, constructing a plausible-sounding narrative after the fact. This matters in high-stakes applications. When a diagnostic system makes an error, clinicians need to know whether the reasoning was flawed, the evidence was unreliable, or the underlying assumptions were incorrect. A post-hoc rationalisation does not provide that accountability.
An Architecture for Open-World Reasoning
AlphaGo's game tree is a record of what the system knows about a position. It contains all the variations the program has considered, each annotated with evaluations from its neural networks. As reasoning progresses, the tree is updated, and the accumulated information is synthesised to select a move. This structure is inspectable. An analyst can trace exactly which futures were explored and why certain branches were pruned.
A general reasoning system should maintain an analogous epistemic state: a structured representation of what is settled, what remains uncertain, what has been ruled out, and which questions are still open. Reasoning then becomes a sequence of moves that update this state - deducing consequences, decomposing problems, deciding which experiment to run or calculation to perform next. Each move should be evaluated by an independent component that assesses whether it genuinely reduces uncertainty and whether belief updates are supported by evidence.
Open-world reasoning is harder than playing Go. The current state of affairs is only partially observable. The set of available actions is large and context-dependent. The consequences of actions are stochastic or unknown. But the neural models we have today, including large language models, can contribute to such a system. They can propose ways to tackle a problem given the current state of knowledge. They can interact with tools and external databases. They can help assess whether a claim is supported by available evidence.
The critical difference is that these capabilities would be embedded in an architecture that enforces epistemic discipline. An independent evaluation mechanism would audit each reasoning step, updating beliefs only when the change is justified. Over time, the system could accumulate certified knowledge and improve its reasoning policy by learning from past experiences. This is the scientific method implemented as a computational process, designed to produce knowledge that can withstand scrutiny.
Why This Matters Now
Thore Graepel, who was a core member of the AlphaGo team, recently left Google DeepMind to pursue this vision. His departure signals a broader tension within the AI research community. Scaling up system-one capabilities - making intuition sharper by training on more data - has delivered remarkable results. But it does not transform intuition into deliberation.
Move 37 mattered because a machine held a position, weighed possible futures, and chose a move its pattern-matching instincts would have rejected. Society needs analogous breakthroughs in drug discovery, materials science, climate modelling, and medical diagnosis. These are domains where the problem structure is not given in advance, where the rules are unknown or contested, and where the cost of error is high.
We will not get those breakthroughs from systems that generate convincing narratives after reaching a conclusion. We need systems whose conclusions arise from an auditable sequence of evidence, inference, and belief revision. The architecture exists in outline. What remains is to build it and to insist that the industry's focus on scale does not crowd out the harder work of engineering genuine reasoning into AI.
The current generation of models can complete patterns with extraordinary fluency. They cannot yet hold a hypothesis, test it systematically, and revise their beliefs in light of what they find. Until they can, the creative moves we need most will remain out of reach.



