DeepSeek-V4-Flash-Vision-Exp
DeepSeek's experimental multimodal variant of V4-Flash for document and chart understanding
DeepSeek released DeepSeek-V4-Flash-Vision-Exp on August 21, 2026 as an experimental image-input variant of DeepSeek-V4-Flash, the sparse mixture-of-experts model with 284 billion total parameters and 13 billion active per token. It matches the text-only V4-Flash on agentic and reasoning tasks while adding image understanding, aimed at document and chart reading, visual question answering, and agent workflows that mix text and images. Images can be passed as inline base64 data, URLs, or through DeepSeek’s Files API, and the model tokenizes each image at 384 tokens, well below what GPT and Claude models typically use per image.
On DeepSeek’s internal multimodal agent benchmarks, the company reports V4-Flash-Vision-Exp closing much of the gap to Anthropic’s Opus-4.8 on tasks that require reasoning over visual input, a notable jump from the text-only V4-Flash. It’s priced at $0.22 per million input tokens and $0.66 per million output tokens, with a separate lower rate for cached input, and remains available only through the DeepSeek API while the company gathers feedback ahead of a non-experimental release.