Search versus prediction
A researcher who helped build AlphaGo has argued that today's large language models should not be credited with the same kind of reasoning that powered the Go-playing system's 2016 defeat of champion Lee Sedol. The distinction matters, the author says, because dependable scientific and medical applications require systems that can test possible actions and evaluate their consequences rather than simply produce plausible sequences.
AlphaGo combined two complementary mechanisms. A policy network proposed moves that resembled promising human play, while a search system explored branches of future moves and countermoves. That second component allowed the program to assess consequences beyond immediate plausibility. The architecture was particularly important in Go, where exhaustive calculation is impossible because the number of potential positions grows enormously.
The famous 37th move in the second game against Lee illustrates the difference. The policy component rated the move as extremely unlikely to be chosen by a strong human. AlphaGo selected it because its search process looked further into possible continuations and judged the unusual placement favourably. The system won the game and the five-game match, four games to one.
Large language models operate differently. They repeatedly predict a likely next token based on the preceding context. Additional techniques can improve performance, and models can produce text that resembles a step-by-step solution, but the article's central claim is that fluent output is not equivalent to constructing and checking an explicit tree of possible futures.
Why the distinction matters
The comparison resembles the difference between quick intuition and deliberate analysis. AlphaGo's networks supplied promising options, while search provided a structured way to challenge those suggestions. Neither piece was sufficient by itself: unrestricted search would be too expensive, while an intuitive proposal mechanism alone would have missed the decisive move.
Language-model developers increasingly encourage systems to generate intermediate steps, call tools and verify answers. Those methods can make outputs more useful, but the researcher warns against assuming that apparent deliberation proves an underlying capability comparable to AlphaGo's search. A model can reproduce the language of reasoning without reliably testing every inference.
This is an argument about architecture and reliability, not a claim that language models have no practical value. Their ability to summarise information, write software and propose solutions can be substantial. The concern is that high-stakes users may place too much confidence in a polished answer when the system lacks a dependable mechanism for exploring alternatives and rejecting attractive mistakes.
The AlphaGo example therefore remains relevant a decade later. Its achievement came from combining learned intuition with explicit evaluation. Reaching similarly trustworthy results in open-ended fields may require systems that integrate both capabilities rather than relying on text prediction alone.



