Granite Vision
IBM's compact vision-language model for document and chart understanding
Granite Vision is IBM’s compact vision-language model, released as part of the Granite 3.2 update in February 2025 and built specifically for visual document understanding rather than general image chat. At 2 billion parameters, it reads tables, charts, infographics, plots, and scanned documents and turns them into structured text, targeting benchmarks like DocVQA and ChartQA where IBM says it matches open models several times its size. The small footprint is deliberate: IBM designed it for everyday enterprise workloads, such as parsing invoices or extracting data from reports, where running a large multimodal model isn’t practical or affordable. It supports a 16K token context window, is open-weighted under Apache 2.0, and is distributed through Hugging Face and IBM watsonx, with later versions (3.3 and the Granite 4.0 vision variants) building on the same document-focused design.