Doubao Vision
ByteDance's multimodal Doubao tier for image and document understanding
Doubao Vision is ByteDance’s multimodal tier for image and document understanding, launched on December 18, 2024 at Volcano Engine’s Force conference. It reads and reasons over visual content: identifying objects, analyzing charts, working through logic problems posed as images, reading code from screenshots, and answering questions tied to a picture rather than only describing it. ByteDance built it as part of the broader Doubao model family that powers its consumer assistant app and its Volcano Engine cloud API.
As with other Doubao releases, price was central to the pitch. ByteDance priced the model at roughly 0.003 yuan per 1,000 tokens, which it billed as about 85% cheaper than the industry average, enough to process hundreds of standard-resolution images for a single yuan. It’s a closed, proprietary model available only through ByteDance’s own app and API.