DeepSeek unveils an experimental multimodal version of its V4 Flash model, saying it nears the performance of Anthropic's Opus 4.8 on multimodal agentic tests
DeepSeek unveiled an experimental AI model that can understand visual prompts, saying the tool nears the performance of an advanced model by US rival AnthropicPBC.
Adding visual-prompt understanding shifts V4 Flash from a text-and-coding performance contest toward multimodal agentic work. DeepSeek’s comparison with Anthropic’s Opus 4.8 makes Anthropic the immediate benchmark for that broader capability claim.
First-order effects
DeepSeek can test V4 Flash with workloads that combine images and instructions, while positioning its experimental version against Anthropic’s Opus 4.8 on multimodal agentic tests.
Anthropic faces a new direct performance comparison from DeepSeek in a category beyond the coding strengths attributed to V4 before launch.
Second-order effects
Buyers evaluating DeepSeek and Anthropic gain another model-selection criterion: multimodal agentic performance, rather than text or coding results alone.
DeepSeek’s prior V4 Flash preview now has a clearer upgrade path, pressuring rival model providers to demonstrate comparable visual-input performance in agentic evaluations.
Third-order effects
If multimodal agentic tests become a recurring comparison point, frontier-model competition will increasingly turn on end-to-end task performance across visual and text inputs rather than on single-modality benchmarks.
AI procurement is likely to become more workload-specific, with enterprises comparing models by the useful tasks they complete instead of treating a general benchmark score as sufficient.
The trend: Frontier AI competition is broadening from coding and general benchmarks to multimodal agents evaluated on complete tasks.
Multimodal API support 🔌 🔹 Set model='deepseek-v4-flash-vision- exp' 🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing 🔹 Supports Chat Completions, Messages & Responses 🔹 Supports mixed text + image input; images can be provided via base64,
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major
I will move my low cost AI task to this new flash version. In just half a month, it got event better at SWE + added vision support. Opus4.8 level SWE @ DeepSeek flash price. Great stuff.
wtf is happening today: DeepSeek has launched DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model built for agents that need to see. And its performance on visual-agent benchmarks moves close to or even outperforms Opus 4.8. Again: this is the Flash model, the
Files API is now live. 📁 🔹 Free to use 🔹 Upload an image once, then reference it by file_id to save request bandwidth 🔹 Reuse the same image across requests—no need to upload it again Learn more: https://api-docs.deepseek.com/ ... 4/n
Multimodality unlocks more agent use cases. 👀 V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows. 2/n
DeepSeek V4 Flash Vision Exp is live on OpenRouter! @deepseek_ai's new model supports Image input at V4 Flash pricing, matches V4 Flash 0731 on text agent benchmarks, and beats Opus-4.8 on Agents' Last Exam and ZeroBench. Use it now: https://openrouter.ai/...
This is important too Two more notes: - they don't call models “exp” or “preview” as a joke. 3.2-exp was half a generation behind 3.2. Vision-Full will go crazy - nothing said on open weights so far