Doubao Vision

ByteDance's multimodal Doubao tier for image and document understanding

Freemium Multimodal
Visit Product Page →
128000 context tokens
Undisclosed parameters
Proprietary license
Dec 2024 released

Doubao Vision is ByteDance’s multimodal tier for image and document understanding, launched on December 18, 2024 at Volcano Engine’s Force conference. It reads and reasons over visual content: identifying objects, analyzing charts, working through logic problems posed as images, reading code from screenshots, and answering questions tied to a picture rather than only describing it. ByteDance built it as part of the broader Doubao model family that powers its consumer assistant app and its Volcano Engine cloud API.

As with other Doubao releases, price was central to the pitch. ByteDance priced the model at roughly 0.003 yuan per 1,000 tokens, which it billed as about 85% cheaper than the industry average, enough to process hundreds of standard-resolution images for a single yuan. It’s a closed, proprietary model available only through ByteDance’s own app and API.