Gemini 2.5 Flash

Google's low-latency, cost-efficient Gemini tier

Paid Language
Visit Product Page →
1000000 context tokens
Undisclosed parameters
Proprietary license
Apr 2025 released

Gemini 2.5 Flash is Google’s cost-efficient tier of the 2.5 generation, released as a preview in April 2025 and moved to general availability the following month. It’s the first Flash model built with visible thinking built in from the start: it can reason step by step through a problem before answering, and developers can set a “thinking budget” to trade off answer quality against latency and cost. It takes text, image, and audio input and handles the same million-token context window as its Pro sibling.

Google prices it at $0.30 per million input tokens and $2.50 per million output tokens, with thinking tokens billed at the output rate, and offers a further stripped-down Flash-Lite variant for even cheaper, simpler workloads. It scores around 83% on MMLU-Pro and performs well on coding and math benchmarks relative to its price, making it Google’s default recommendation for high-volume applications that don’t need Pro-level reasoning depth.