New benchmark: AI models still cannot see properly
Moonshot AI's PerceptionBench isolates the visual perception of multimodal models from reasoning. None of the 16 frontier models tested reached 60 percent accuracy.
Research
Moonshot AI's PerceptionBench isolates the visual perception of multimodal models from reasoning. None of the 16 frontier models tested reached 60 percent accuracy.