AI models flub these intelligence tests. Can you fare any better?
In late‑2024 a Columbia University research team measured the performance of leading LLMs on New York Times Connections, finding they correctly solved only 18 percent of the grid‑based riddles. By early 2025, newer model releases from major AI firms were able to approach near‑perfect scores on the same test, illustrating a steep improvement curve. Parallel work by Google and the University of Illinois Urbana‑Champaign demonstrated that when presented with “Knights and Knaves” logic puzzles—variations of a well‑known truth‑lie scenario—top‑tier models repeatedly fell for subtle wording changes, revealing an over‑reliance on memorized patterns rather than genuine inference. The article also highlights that visual‑spatial challenges, such as mental‑rotation IQ items and the ARC‑AGI abstract‑grid benchmark, remain a persistent weakness, with models only succeeding when the visual data is pre‑converted into symbolic strings.
These findings sit against a broader narrative of AI systems rapidly eclipsing human performance on language‑heavy benchmarks while still lacking robust world modeling. The Columbia and Google‑UIUC experiments underscore that scaling alone does not guarantee transferable reasoning; architectural advances that integrate true 3‑D perception or causal logic are still nascent. Competitors are racing to embed multimodal “world models” that can manipulate objects virtually, but the persistent puzzle failures suggest a lag between raw data memorization and the kind of flexible problem‑solving humans use in everyday cognition. As AI products move into domains like design assistance, autonomous robotics, and education, the gap highlighted by these puzzles could become a differentiator for firms that succeed in building genuine spatial and abstract reasoning capabilities.
Looking ahead, the next frontier will be systematic evaluation suites that combine visual, linguistic, and logical elements, forcing models to demonstrate cross‑modal reasoning under novel constraints. Researchers should monitor whether upcoming multimodal architectures—such as next‑generation vision‑language transformers—can close the performance gap on tasks like ARC‑AGI without resorting to brittle token‑level tricks. Industry adopters must also assess risk: deployments that assume human‑level reasoning may falter when confronted with edge‑case puzzles, leading to errors in safety‑critical or creative workflows. Continued benchmarking against human‑designed puzzles will be essential to gauge genuine progress beyond headline metrics.
Key Takeaways
Columbia’s late‑2024 test showed top LLMs solved only 18 % of NYT Connections puzzles, a figure that rose to near‑perfect only after a year of rapid model updates.
Google and UIUC found that even the most advanced models misinterpret subtle variations in classic Knights‑and‑Knaves logic problems, exposing over‑reliance on memorized patterns.
Visual‑spatial reasoning, including mental‑rotation IQ items and ARC‑AGI grid puzzles, remains a blind spot for current multimodal models unless inputs are converted to symbolic representations.
Future AI competitiveness will hinge on building true world‑modeling abilities that can handle novel, cross‑modal puzzles without resorting to dataset‑specific shortcuts.
About the Source
This analysis is based on reporting by MIT Technology Review. Here is a short excerpt for context:
Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur…Read the original at MIT Technology Review