DeepSeek: A Multi-modal Flash Beats Claude Opus 4.8 on Agent Tasks From DeepSeek
DeepSeek introduces the V4-Flash-Vision-Exp: a multi-modal model that matches Anthropic’s Claude Opus 4.8 on agentic tasks (meaning it solves problems sequentially on its own) and can handle up to 600 images per request.
The V4-Flash-Vision-Exp is a model that parses through images, screenshots, and charts. It performs closely with Anthropic’s Claude Opus 4.8 on agentic tasks and is able to handle up to 600 images per request — all for the same cost as the default V4-Flash (which is text-only).
Last year, multi-modal models like this would be considered high-priced and inaccessible. With this latest release from DeepSeek, we’re continuing to see an increase in affordability for this technology. This will continue to put more heat on OpenAI and Anthropic.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles