Top AI models fail a simple looking test that people solve easily

Researchers built puzzles that need careful, repeated looking: counting scattered fields, tracing a tangled rope, matching shapes. The best AI model solved only 10.6% of them, while three human testers averaged 96.1%.

Why it matters

Today's AI image tools describe photos well, so people assume they can truly look at pictures. This study shows they struggle to count, trace, or compare carefully across an image. Those are skills needed for reading medical scans, inspecting factory parts, or checking satellite photos. Until that gap closes, it is risky to trust these tools for jobs where missing one detail matters. Note the test used specially built images, not everyday photos, so real-world behaviour may differ.

Who's behind it: Jiarui Zhang and colleagues, University of Southern California.

Summary by the Lemma AI · how we grade

Read the original paper (arxiv.org)