Can LLMs really not jump, or are we defining the jump incorrectly?
Tom Zahavy of Google DeepMind argues in LLMs can’t jump that models are strong at induction and deduction but structurally inadequate at abduction, which requires proposing a new explanatory framework. His central example is Einstein’s local equivalence between gravity and acceleration [1].
Two developments challenge that categorical claim:
- Co-Scientist, published in Nature, independently reconstructed a then-unpublished hypothesis about bacterial gene transfer. Selected hypotheses were tested in vitro [2].
- On July 22, GPT-5.6 produced a seven-node counterexample to the Dinitz-Garg-Goemans conjecture concerning cost-preserving rounding [3], [4]. The result had not been peer reviewed when announced and should not be treated as an established theorem. The counterexample is finite and directly checkable.
Do these examples refute Zahavy’s thesis? I do not think so.
In both cases, people set the goal and the evaluation framework. There is still no publicly available, verified example of a model constructing a theory on the scale of General Relativity using only the knowledge and observations available at the time.
But a categorical claim carries a burden of proof. If a system produces a previously unknown, verifiable, and meaningful structure, what objective criterion makes it “only search” rather than a jump?
The paper does not operationalize this distinction, and the field does not appear to have an agreed criterion. Even if a model solved P versus NP with a verified proof, the result could still be classified as “only deduction” because it worked within fixed axioms.
Hallucinations also need a more careful interpretation. Most are unsupported or incorrect outputs, not scientific hypotheses. Still, models can generate many candidate explanations. The missing component may be a reliable epistemic selection mechanism that rejects candidates through evidence, consistency checks, and experiments.
The question is therefore not only whether LLMs can jump. Should scientific novelty be defined by the resulting discovery, the mechanism that produced it, or both?
References
- Tom Zahavy, LLMs can’t jump, PhilSci-Archive, January 2026.
- Juraj Gottweis et al., “Accelerating scientific discovery with Co-Scientist”, Nature, 2026.
- Dmitry Rybin, counterexample announcement, July 22, 2026.
- Vera Traub, Laura Vargas Koch, and Rico Zenklusen, “Single-Source Unsplittable Flows in Planar Graphs”, arXiv:2308.02651.