HyperCLOVA X Vision
Naver's multimodal extension of HyperCLOVA X for image understanding
HyperCLOVA X Vision is Naver’s multimodal extension of its HyperCLOVA X language model, adding image understanding on top of the text model’s Korean-language strength. Naver introduced it in August 2024, and the company has said it answered 83.8 percent of Korean college entrance exam questions that included image inputs, ahead of the 77.8 percent Naver reported for GPT-4o on the same test. That benchmark matters for Naver’s pitch: rather than compete on generic multimodal leaderboards, the model is built to handle Korean-specific documents, exam questions, and visual context that Western models tend to get wrong. It’s available through Naver Cloud’s API and the CLOVA X app, aimed at Korean enterprises that need document QA, chart reading, or image-grounded chat in their own language. Naver has continued to expand the lineup with smaller open variants under the HyperCLOVA X SEED family and an omnimodal version that adds audio.