Puzzles have served as a benchmark for artificial intelligence since the mid-20th century, tracing back to early algorithms designed for checkers. While modern large language models (LLMs) have demonstrated significant improvements—such as moving from a 18% success rate on New York Times Connections puzzles in 2024 to near-perfect performance by early 2025—they continue to encounter specific cognitive hurdles.
Researchers are currently utilizing a variety of tests to highlight the differences between machine and human problem-solving. Spatial reasoning remains a notable weakness; despite their ability to process visual data, LLMs often struggle with mental rotation tasks that require manipulating 3D objects. Similarly, abstract reasoning benchmarks like ARC-AGI reveal that while models can achieve high scores, they often rely on non-generalizable, complex rules rather than the intuitive visual strategies employed by humans.
Memory can also act as a liability for AI. Studies, such as those involving "Knights and Knaves" logic puzzles and the "SimpleBench" suite, suggest that models may fail when a problem closely resembles training data, causing them to overlook subtle variations that a human would easily detect. Furthermore, as the complexity of tasks increases—such as in multi-step river-crossing scenarios or logic grid puzzles—models frequently falter once the number of variables exceeds a certain threshold.
Conversely, psychologists have identified areas where humans are more prone to error than AI. Some math-based riddles are designed to trigger intuitive, "knee-jerk" human mistakes, whereas models may respond with more deliberative logic. By analyzing where these technologies succeed and fail, developers hope to gain a clearer understanding of the fundamental differences between machine-learned patterns and human intelligence.
Source: MIT Technology Review
No comments yet. Be the first to share your thoughts.