Large language models can generate essays, pass medical licensing exams and write code, yet they routinely fail puzzles that most humans solve in seconds. That gap between impressive capability and basic reasoning failure has become a focus for researchers seeking to measure genuine machine intelligence. The Columbia University team observed that the best models initially solved only 18% of New York Times Connections puzzles. But by early 2025 some models achieved near-perfect accuracy on the same task. Despite that rapid progress, deeper tests reveal enduring weaknesses in spatial reasoning, adapted logic puzzles and abstract visual problems.
Puzzle Tests Expose AI Reasoning Gaps
For the University of Illinois Urbana-Champaign team, the classic Knights and Knaves puzzles provided a revealing test. In these problems, some characters always tell the truth and others always lie. When researchers introduced slight variations, LLMs often reverted to memorized solutions and missed the trick. The same phenomenon appears in SimpleBench, where questions resemble more complicated training examples. Humans spot the difference quickly, but top-tier models trip.
Spatial reasoning remains a notable weakness. Mental rotation problems that require manipulating 3D objects in the mind are trivial for architects and engineers but stump language models. Even when LLMs can analyze visual inputs, they cannot perform the transformations that human spatial thinkers handle instinctively. These failures show that current models lack a genuine understanding of physical space.
Why This Matters
These persistent failures carry real-world implications. Industries deploying AI in robotics, autonomous navigation and medical imaging rely on systems that can reason about space and adapt to new situations. A model that cannot solve a mental rotation problem may struggle to interpret a camera image from an unfamiliar angle or to predict how a physical object will behave. For researchers, the results also raise questions about how to define and measure intelligence. Benchmarks like ARC-AGI push models toward non-human strategies that pass tests without genuine reasoning. If AI systems become highly capable at narrow tasks while failing basic cognition tests, trust in their reliability will erode. Regulators and developers must consider whether current evaluation methods capture the true limits of machine intelligence before deploying these tools in safety-critical environments.



