DeepSeek-V4.1-Flash
DeepSeek's cheaper, vision-capable follow-up to V4-Flash with tiered peak pricing
Paid
Multimodal
1048576
context tokens
552B (MoE)
parameters
Proprietary
license
Sep 2026
released
DeepSeek shipped V4.1-Flash on September 10, 2026, a 552-billion-parameter mixture-of-experts model with native vision support and a 1,048,576 token context window that extends to 393,216 tokens of output. Pricing runs on a peak/off-peak split: off-peak rates are $0.003 per million input tokens on a cache hit, $0.15 on a cache miss, and $0.60 per million output tokens, with peak-hour rates (weekday mornings UTC) exactly double.
The tiered pricing is DeepSeek leaning into its usual playbook: undercut Western API pricing hard during low-demand windows to pull in batch and non-latency-sensitive workloads, while still charging enough at peak to manage capacity.