What happened?
Mathematicians Timothy Gowers and Peter Sarnak credit large language models with serious mathematical skill, but say they see a hard limit when it comes to producing genuinely new ideas.
Gowers's argument is this: today's models are good at combining known methods and at trying many search paths. What is missing is the intuition to pick the few routes that are productive in a vast search space. Sarnak reaches the same place: AI can derive results from existing theory, but fails to develop the abstractions that underpin major proofs when it starts from an elementary question.
Why the distinction matters
The distinction here measures something different from what benchmark scores show. Scoring highly on a mathematics test means known techniques can be applied in the right order. Building a new theory requires deciding which technique is worth trying — and that decision is made before the trying.
It is a practical problem about the size of the search space. With an unbounded number of valid steps available, trying all of them is impossible; the mathematician's work is eliminating the ones that will not be tried. That models cannot do this elimination is not a gap that more compute closes.
A similar conclusion from DeepMind
DeepMind researcher Tom Zahavy arrives at a similar point. In a paper titled "LLMs Can't Jump", he names the bottleneck manipulative abduction: the ability to invent new foundational assumptions with no linguistic precedent.
The naming itself says something. A model putting forward a concept absent from its training data is, by definition, not predictable, because that concept leaves no trace in language. Zahavy suggests world models could offer a way forward.
The assessments in brief
- Gowers: models combine known methods and try many paths; the intuition to pick the productive one is missing.
- Sarnak: results can be derived from existing theory, but abstractions starting from an elementary question cannot be developed.
- Zahavy: the bottleneck is inventing new foundational assumptions with no linguistic precedent.
- Possible direction: world models.
Why the measurement itself is hard
Showing that a model "cannot produce new ideas" is a different job from computing a benchmark score. A benchmark asks questions whose answers are known in advance; a new idea, by definition, appears somewhere the answer is not known.
That is why these assessments rest on observation from inside the field rather than on a table of numbers. What Gowers and Sarnak are saying is this: ask the model to apply a known technique and you get a result, but ask which technique is worth trying and the answer stays shallow. In mathematics the second question is the work; the first is computation.
The broader debate
These assessments feed a wider question in the field: are models genuinely becoming more versatile, or are they "just" getting better on benchmarks and in familiar problem spaces?
That the question comes from mathematics is no accident. Mathematics is one of the few areas where the right answer can be checked by a machine; a proof is either valid or it is not. That makes it far clearer than other fields where exactly a model gets stuck.
What is not settled
These are not an experimental measurement but the assessment of leading figures in the field. Gowers's and Sarnak's observations rest on their own working practice; no breakdown was shared of how many models were tried on which tasks. Zahavy's work is a published paper and sets out its claim within its own framework.