AI Limits: Induction, Deduction and Why Models Can't Generalize
deepmind reasoning
| Source: Mastodon | Original article
LLMs excel at induction and deduction but falter at abduction, highlighting fundamental reasoning limits, according to DeepMind researcher Tom Zahavy.
DeepMind researcher Tom Zahavy has released a position paper – “LLMs Can’t Jump” – that argues large language models (LLMs) are fundamentally unable to perform abductive reasoning, the creative leap that underpins scientific breakthroughs such as Einstein’s equivalence principle. Presented at ICML 2026, the paper distinguishes three modes of inference: induction (pattern matching from data), deduction (logical inference from premises) and abduction (hypothesis generation). Zahavy contends that while transformers have mastered the first two, they lack the embodied simulation required to generate novel hypotheses, a step he calls the “jump”.
The claim matters because it challenges the prevailing assumption that scaling compute and data will eventually yield general scientific intelligence. If LLMs cannot invent new axioms, their role in discovery‑driven fields – from physics to drug design – may remain limited to analysis and synthesis of existing knowledge. The argument also reverberates through ongoing debates on AI safety and alignment, where the capacity to generate unforeseen strategies is a double‑edged sword.
The paper has already sparked a small but active discourse. Researchers are probing whether multimodal memory, embodied agents, or hybrid symbolic‑neural architectures can supply the missing abductive faculty. Upcoming workshops at the next NeurIPS and the European Conference on Artificial Intelligence are expected to feature rebuttals and experimental attempts to bridge the gap. As we reported on 30 September 2026, scaling alone has not guaranteed broader reasoning abilities; Zahavy’s thesis reinforces that view and points to a new frontier – engineering models that can “jump” rather than merely extrapolate. Watching how the community responds will reveal whether the limitation is technical, conceptual, or both.
Sources
Back to AIPULSEN