Large language models can generate essays, pass medical licensing exams and write code, yet they routinely fail puzzles that most humans solve in seconds. That gap between impressive capability and basic reasoning failure has become a focus for researchers seeking to measure genuine machine intelligence. The Columbia University team observed that the best models initially solved only 18% of New York Times Connections puzzles. But by early 2025 some models achieved near-perfect accuracy on the same task. Despite that rapid progress, deeper tests reveal enduring weaknesses in spatial reasoning, adapted logic puzzles and abstract visual problems.

What You Need to Know

Researchers are using puzzles like mental rotation, Knights and Knaves variations and the ARC-AGI benchmark to probe where LLMs fall short. These tests expose a tendency for models to rely on memorized patterns rather than reasoning from scratch. The results suggest that even as AI passes traditional metrics, it still lacks the flexible cognition humans apply to novel problems.

Puzzle Tests Expose AI Reasoning Gaps

For the University of Illinois Urbana-Champaign team, the classic Knights and Knaves puzzles provided a revealing test. In these problems, some characters always tell the truth and others always lie. When researchers introduced slight variations, LLMs often reverted to memorized solutions and missed the trick. The same phenomenon appears in SimpleBench, where questions resemble more complicated training examples. Humans spot the difference quickly, but top-tier models trip.

Spatial reasoning remains a notable weakness. Mental rotation problems that require manipulating 3D objects in the mind are trivial for architects and engineers but stump language models. Even when LLMs can analyze visual inputs, they cannot perform the transformations that human spatial thinkers handle instinctively. These failures show that current models lack a genuine understanding of physical space.

  • Spatial reasoning: Mental rotation and 3D manipulation tasks consistently defeat LLMs, highlighting a gap in world model understanding.
  • Adapted logic puzzles: Variations on Knights and Knaves and SimpleBench show models rely on memorization rather than adaptation.
  • Abstract visual patterns: ARC-AGI problems force models to infer general rules; humans use simple visual concepts while models often craft byzantine, non-generalizable rules.

Why This Matters

These persistent failures carry real-world implications. Industries deploying AI in robotics, autonomous navigation and medical imaging rely on systems that can reason about space and adapt to new situations. A model that cannot solve a mental rotation problem may struggle to interpret a camera image from an unfamiliar angle or to predict how a physical object will behave. For researchers, the results also raise questions about how to define and measure intelligence. Benchmarks like ARC-AGI push models toward non-human strategies that pass tests without genuine reasoning. If AI systems become highly capable at narrow tasks while failing basic cognition tests, trust in their reliability will erode. Regulators and developers must consider whether current evaluation methods capture the true limits of machine intelligence before deploying these tools in safety-critical environments.