DeepSeek releases an experimental multimodal AI model that can understand visual prompts, alongside text. The company says the model’s “agent performance” is close to Anthropic’s Opus 4.8. It also publishes a benchmarking table indicating the model performs better than Opus 4.8 on three of eleven benchmarks.
The Next Web and Bloomberg both describe the model as an extension of DeepSeek’s text-only V4-Flash, with added capability to read images and screenshots. Both outlets report that the model is made available via DeepSeek’s API platform, positioning it for developers to test the system on tasks that include images rather than text alone.
While both sources agree on the core announcement—an experimental vision-capable model, its availability through an API, and its claimed proximity to Anthropic’s Opus 4.8—the outlets mainly differ in emphasis. The Next Web provides more detail on the model name and the specific benchmark comparison, while Bloomberg focuses on the rivalry context and the claim of near-matching performance.