Granite Vision

IBM's compact vision-language model for document and chart understanding

Free Multimodal
Visit Product Page →
16384 context tokens
2B parameters
Apache 2.0 license
Feb 2025 released

Granite Vision is IBM’s compact vision-language model, released as part of the Granite 3.2 update in February 2025 and built specifically for visual document understanding rather than general image chat. At 2 billion parameters, it reads tables, charts, infographics, plots, and scanned documents and turns them into structured text, targeting benchmarks like DocVQA and ChartQA where IBM says it matches open models several times its size. The small footprint is deliberate: IBM designed it for everyday enterprise workloads, such as parsing invoices or extracting data from reports, where running a large multimodal model isn’t practical or affordable. It supports a 16K token context window, is open-weighted under Apache 2.0, and is distributed through Hugging Face and IBM watsonx, with later versions (3.3 and the Granite 4.0 vision variants) building on the same document-focused design.