A public exchange among AI researchers kicked off after former Google DeepMind researcher Thore Graepel used an X post tied to a new MIT Technology Review piece to argue that today’s most advanced AI systems still lack something essential. Graepel wrote that large language models are “remarkably capable,” but that “generating longer chains of thought is not the same as genuine reasoning.”
That post laid out Graepel’s core complaint: current systems, in his view, usually do not keep an explicit, inspectable record of what they know, what remains uncertain, what evidence supports a conclusion, or whether real progress has been made. He argued that if AI is going to produce “trustworthy and genuinely novel insights in science, medicine and beyond,” its conclusions need to come from “an auditable process of evidence, inference and belief revision.”
Graepel tied that argument to AlphaGo. In the same post, he wrote that the system’s famous Move 37 was not the result of intuition alone, but of being able to “search possible futures, test its instincts and reason about what would happen next.” He also said that view helped drive his recent departure from Google DeepMind, writing that he left because he believes AI needs “a fresh approach to machine reasoning,” drawing on “some of the architectural lessons from AlphaGo.”
The argument turned into a debate about search, tools, and structure
The discussion quickly moved from Graepel’s critique to a broader question: what should count as reasoning in AI, and can LLMs get there by adding the right machinery?
Yann LeCun responded in clear terms: “Good piece. True reasoning must involve a search. LLMs don't possess this capability.” That lined up with Graepel’s AlphaGo comparison, which emphasized exploring possible futures rather than just producing fluent explanations.
David Duvenaud pushed on a different point. In a reply, he asked whether “we might simply be able to provide similar tools to LLMs to achieve true reasoning” by Graepel’s definition, and whether that was “more or less” what Graepel planned to do.
Graepel’s answer was that the issue is not just output length or raw verbal fluency, but disciplined structure. In one reply, he argued that reasoning improves when the process is carefully organized rather than left as a “stream of consciousness,” adding that otherwise humans fall into familiar cognitive biases. In a later follow-up, he said, “Definitely the tools, but more importantly, the rigorous application of the scientific method at scale!”
Gary Marcus also endorsed Graepel’s quoted claim that trustworthy AI requires “an auditable process of evidence, inference and belief revision,” adding, “It is madness to believe otherwise.”
What the exchange actually establishes is narrower than a settled consensus. Graepel publicly argued for auditable, search-like reasoning; LeCun agreed that reasoning requires search; Duvenaud questioned whether similar capabilities could simply be added to LLMs; and Marcus backed Graepel’s framing. The open question, based on these posts, is whether that kind of structured reasoning demands a fundamentally different architecture or can be layered onto today’s language models.
assignment_id: cd0511c4-75e4-42a3-8a75-9ae6bdcec1e3
brief_hash: ba3299256338b1416fba5625d98c59cd71491efbb1d9229e2c4144c3628c3f05
draft_status: drafted
claim-to-source support:
- Graepel argued advanced AI needs auditable reasoning beyond longer chains of thought: b765ef6b-fa0e-4dd2-b472-65f154433874
- Graepel linked the argument to AlphaGo Move 37 and search over futures: b765ef6b-fa0e-4dd2-b472-65f154433874
- Graepel said he recently left Google DeepMind over the need for a fresh approach: b765ef6b-fa0e-4dd2-b472-65f154433874
- Graepel said trustworthy novel insights require evidence, inference, and belief revision: b765ef6b-fa0e-4dd2-b472-65f154433874
- Graepel said reasoning should be carefully structured, not stream-of-consciousness: 6544844e-6c17-4ba1-a76c-2e54cf72a807
- Graepel later emphasized tools plus the scientific method at scale: 33cd07a7-bc24-41e0-99c0-d6238e4bda15
- Yann LeCun said true reasoning requires search and that LLMs lack it: 43657716-da14-416d-89c4-558bea4432d6
- David Duvenaud asked whether similar tools could be given to LLMs: 37c78241-f2b3-4e57-b397-cf49a3c34eba
- Gary Marcus endorsed Graepel’s quoted claim: 435da973-9a4a-44b2-b673-130ec2ac7716