Moonshot AI: New Benchmark Exposes Poor Visual AI Perception
TL;DR. Moonshot AI introduces PerceptionBench, a new benchmark revealing that even leading multimodal AI models struggle with basic visual perception tasks. - PerceptionBench isolates ten sub-skills for visual processing, distinct from logical reasoning or external knowledge. - No frontier model achieved 60% accuracy, with GPT-5.6 Sol narrowly leading competitors. - The findings suggest many perceived reasoning errors are actually rooted in flawed image interpretation.
- Moonshot AI's PerceptionBench evaluates multimodal models' visual perception independently of reasoning or knowledge.
- The benchmark isolates ten atomic visual sub-skills, built from real-world model errors.
- No tested frontier model, including GPT-5.6 Sol, Kimi K3, and Claude Fable 5, reached 60% accuracy.
- The research indicates that many errors attributed to logical reasoning are actually fundamental visual perception failures.