DeepSeek-VL

DeepSeek's first vision-language model, preceding DeepSeek-VL2

Free Multimodal
Visit Product Page →
4096 context tokens
7B parameters
DeepSeek License license
Mar 2024 released

DeepSeek-VL is DeepSeek’s first vision-language model, released on March 11, 2024 in 1.3B and 7B parameter versions. It pairs a hybrid vision encoder, tuned to handle high-resolution images up to 1024x1024 pixels without a heavy compute cost, with a language model trained on a mix of web screenshots, OCR data, charts, and documents rather than only captioned photos. That training mix was meant to make it useful for practical tasks like reading a chart or a slide, not just describing generic images, and DeepSeek folded vision-language training in from the start of pretraining so the model kept strong text-only performance alongside its visual grounding.

It was a modest but credible early entry in DeepSeek’s push beyond pure language models, and it set up the architecture DeepSeek expanded on later that year with the larger, mixture-of-experts DeepSeek-VL2.