DeepSeek-V4.1-Flash

DeepSeek's cheaper, vision-capable follow-up to V4-Flash with tiered peak pricing

Paid Multimodal
Visit Product Page →
1048576 context tokens
552B (MoE) parameters
Proprietary license
Sep 2026 released

DeepSeek shipped V4.1-Flash on September 10, 2026, a 552-billion-parameter mixture-of-experts model with native vision support and a 1,048,576 token context window that extends to 393,216 tokens of output. Pricing runs on a peak/off-peak split: off-peak rates are $0.003 per million input tokens on a cache hit, $0.15 on a cache miss, and $0.60 per million output tokens, with peak-hour rates (weekday mornings UTC) exactly double.

The tiered pricing is DeepSeek leaning into its usual playbook: undercut Western API pricing hard during low-demand windows to pull in batch and non-latency-sensitive workloads, while still charging enough at peak to manage capacity.