Imagen 4
Google's text-to-image model with photorealistic output
Imagen 4 is Google DeepMind’s text-to-image model, announced at Google I/O in May 2025 and made generally available through the Gemini API and Vertex AI. It generates images at resolutions up to 2K and renders legible text and typography inside images, an area where earlier Imagen versions and many competing models struggled. Google ships it as a family of three variants: a fast tier built for quick drafts at roughly ten times the speed of Imagen 3, a standard tier, and an Ultra tier for maximum fidelity and prompt adherence. Pricing is per image and tiered by variant, with the fast option priced lowest and Ultra priced highest.
The model sits alongside Gemini and Veo in Google’s generative media lineup and competes directly with OpenAI’s GPT Image models and Midjourney. It’s positioned less around raw novelty and more around reliability: consistent adherence to detailed prompts, sharper hands and faces, and fewer of the rendering artifacts that plagued earlier diffusion models. Developers access it through Google AI Studio and Vertex AI, and it also powers image generation inside the Gemini consumer app.