Moonshot AI: New Benchmark Exposes Poor Visual AI Perception

TL;DR. Moonshot AI introduces PerceptionBench, a new benchmark revealing that even leading multimodal AI models struggle with basic visual perception tasks. - PerceptionBench isolates ten sub-skills for visual processing, distinct from logical reasoning or external knowledge. - No frontier model achieved 60% accuracy, with GPT-5.6 Sol narrowly leading competitors. - The findings suggest many perceived reasoning errors are actually rooted in flawed image interpretation.

Sources

Back to QLANKR News