Tech

Where AI still falls short: The puzzles that stump frontier models

While machine learning has advanced rapidly, specific cognitive tasks such as spatial reasoning and complex logic grids remain difficult for current AI models to master.

Editorial persona
Mara Ellison
Science and Space Editor
Published
Draft
Source: MIT Technology Review · View original source
AI models flub these intelligence tests. Can you fare any better?
Artificial Intelligence

Puzzles and games have served as a central benchmark for artificial intelligence development since the 1950s, when Arthur Samuel’s checkers algorithm helped popularise the term “machine learning”. As models have grown more sophisticated, their ability to solve certain types of problems has improved dramatically. However, a new interactive feature from MIT Technology Review highlights that significant gaps remain in how machines process specific cognitive tasks.

The feature presents seven distinct tests designed to expose these limitations. One area where humans maintain a clear advantage is spatial reasoning. In mental rotation problems, which require determining if different images represent the same object from varying angles, language models often fail. Despite their ability to analyse visual inputs, current models struggle to manipulate three-dimensional objects in the way that spatial thinkers, such as architects and mechanical engineers, can.

Another challenge for AI is its reliance on memorised data. In a 2024 study involving “Knights and Knaves” logic puzzles, researchers from Google and the University of Illinois Urbana-Champaign found that models could be tripped up by subtle variations of classic riddles. When a puzzle closely resembles one seen during training, models may overlook key differences and respond with memorised answers rather than deriving a new solution. A similar phenomenon occurs in tests like SimpleBench, where top-tier models stumble on simplified questions that humans can solve by spotting the trick.

Visual pattern recognition also poses difficulties, particularly in two dimensions. The ARC-AGI benchmark, a famous puzzle-based test, requires models to infer abstract rules from examples. Research indicates that models perform better when grids are presented as strings of numbers rather than images. Even when models answer correctly, they often use complex, non-generalisable rules, whereas humans tend to draw on simple visual concepts.

Conversely, there are areas where AI outperforms humans. Psychologists have designed problem suites that exploit human cognitive biases and intuitive mathematical errors. In these “lightning round” questions, humans often provide knee-jerk answers, while models respond more deliberatively. This suggests that machine cognition differs fundamentally from human intuition in specific contexts.

Scale also plays a critical role in model performance. A study by Apple researchers found that large language models can solve simple versions of the Tower of Hanoi and river-crossing puzzles, but their performance degrades significantly when the number of variables reaches six or higher. Similarly, researchers from the University of Washington, Stanford University, and the Allen Institute for AI observed that models struggle with logic grid puzzles as complexity increases, raising questions about whether these failures represent unique reasoning limits or simply the accumulation of errors.

The MIT Technology Review piece invites readers to attempt these puzzles to compare their own performance against current frontier models. While AI is improving quickly, these tests provide a useful window into the technology’s strengths and weaknesses, illustrating where human cognition still holds the edge.

Continue reading

More from Tech

Read next: AI model proposes solution to 370-year-old royal cipher
Read next: Why IMAX 15/70 Cameras Are So Loud
Read next: Insight Partners keeps diversified strategy as AI capital crowds into OpenAI and Anthropic