A reflection on “LLMs don't reason”

Thore Graepel's opinion piece “Don't be fooled—LLMs don't reason” (Graepel was a core member of DeepMind's AlphaGo team) made me reflect on my own attempts at using LLMs for automated hypothesis testing.
One of Graepel's main criticisms of LLMs, as they stand, for domains where traceability and explainability are essential (he cites medicine, engineering and scientific research) is that they keep no explicit, persistent and inspectable record of the hypotheses they are considering, their confidence in them, the evidence they are weighing, or the questions that remain open.
AlphaGo, by contrast, explicitly built a game tree: a data structure recording the variations it had considered, each annotated with its neural networks' evaluations and updated as the search progressed.
Graepel's point matches my recent experience with LLM-backed agents for automated hypothesis testing. When you statistically test hypotheses against a dataset, you soon realise that the set of rejected and non-rejected hypotheses forms a body of knowledge. The agent has to build on that knowledge, draw entailments from it, and incrementally refine its understanding of a complex dataset.
That is exactly where the lack of “reasoning”, in Graepel's sense of the word, shows most clearly. The agents chose easy or irrelevant entailments. They took long, expensive routes to reach trivial conclusions. They re-tested hypotheses that had already been rejected, or drew conclusions that contradicted earlier results.
Has anyone found a good way to give agents a persistent epistemic state that actually holds up?
Starting points on agentic hypothesis testing
- POPPER (Stanford, ICML 2025): agentic hypothesis validation through sequential falsification, with Type-I error control.
- DiscoveryBench (Allen AI, ICLR 2025): a benchmark for LLM-driven, data-driven hypothesis search and verification.
- Kosmos (FutureHouse / Edison Scientific): an AI scientist built around a structured “world model” that accumulates findings across cycles.
Originally posted on LinkedIn.