DeepSeek adds vision for agents: every image costs at most 384 tokens
V4-Flash-Vision-Exp adds image processing while preserving text performance. On the company's own benchmarks it comes close to Opus 4.8.
Models
V4-Flash-Vision-Exp adds image processing while preserving text performance. On the company's own benchmarks it comes close to Opus 4.8.
Moonshot AI's PerceptionBench isolates the visual perception of multimodal models from reasoning. None of the 16 frontier models tested reached 60 percent accuracy.