Fields Medalist Tim Gowers questions OpenAI’s mathematical claims
Tim Gowers challenges the validity of OpenAI’s assertion that it solved ten major problems, including the construction of a non-sofic group, urging caution regarding unverified AI outputs.
Fields Medalist Tim Gowers has published a critical analysis of large language models’ mathematical capabilities, directly responding to recent claims by OpenAI. The blog post, titled 'What sort of maths are LLMs good at?', was published on 12 August 2026, just days after the technology company announced it had solved ten major problems in mathematics and theoretical computer science.
The announcement in question included OpenAI’s claim to have achieved the first construction of a non-sofic group, a significant open problem in the field. Gowers’ commentary serves as a timely reaction to these assertions, examining the extent to which current artificial intelligence systems can genuinely handle complex mathematical reasoning versus generating plausible-sounding outputs.
Gowers explicitly notes that his analysis was written shortly after OpenAI’s public statement, highlighting the immediate need for scrutiny from the academic community. By questioning the nature of these solutions, he underscores the distinction between computational output and verified mathematical proof, a distinction that remains crucial as AI systems become more integrated into scientific research.
The construction of a non-sofic group is widely recognised as a major challenge in theoretical computer science. OpenAI’s assertion that it has resolved this problem, alongside nine others, represents a bold claim that has not yet been subject to independent peer review or verification. Gowers’ post invites readers to consider the limitations of LLMs in this high-stakes context.
As the debate over AI’s role in scientific discovery intensifies, the intersection of computer science and pure mathematics remains a focal point. Gowers’ intervention adds a layer of academic rigour to the conversation, reminding stakeholders that while AI tools are powerful, their outputs in specialised fields require careful validation before being accepted as definitive solutions.
