DeepSeek launches V4-Flash-Vision-Exp, native image input for visual understanding
On August 25, DeepSeek launched the multimodal V4-Flash-Vision-Exp model with native image input; Harness updated to v0.1.1-rc.1.
On August 25, DeepSeek launched the DeepSeek-V4-Flash-Vision-Exp visual understanding model with native image input. The model can identify people, read screenshots, and parse charts. With the addition, the V4 matrix now spans Flash, Pro, and Vision-Exp product lines.
DeepSeek's official benchmarks show V4-Flash-Vision-Exp achieving a significant jump over V4-Flash on vision-enabled Agent Benchmarks, with multimodal capability approaching Claude Opus-4.8. Third-party testing, however, notes that fine-grained person recognition still lags: in one comparison, Doubao correctly identified three similar-looking actors in one frame while V4-Flash-Vision-Exp treated them as the same person.
DeepSeek Harness was updated to v0.1.1-rc.1 alongside the launch, adding official image request support. User tests note the model is free to use on release, continuing the V4 series' dense update cadence. The model is API-only and its weights have not been open-sourced; Harness in the same release also improved subagent invocation for Claude Code and Codex.