DeepSeek V4.1 Flash Launches With Native Vision and Lower AI Inference Costs

DeepSeek's new 552B MoE model is live on its API and is set to replace more of the V4 lineup.

Saganote
Saganote ·
3 Min Read

TL;DR: DeepSeek has released V4.1 Flash, a 552B multimodal MoE model with native vision, lower KV-cache requirements, and a new API endpoint designed for cheaper, faster inference.

DeepSeek V4.1 Flash is now officially available, giving developers a new multimodal model built around a different architecture from the company's earlier V4 releases. DeepSeek says the model is the smallest in its new architecture family while targeting faster inference, higher throughput, and lower serving costs. The company announced the release on September 10, with the model now live through the official DeepSeek announcement.

DeepSeek V4.1 Flash uses a 552B MoE architecture

DeepSeek V4.1 Flash uses a 552 billion-parameter Mixture-of-Experts design with a Causal Encoder-Decoder architecture. Only 8 billion parameters are active during input processing and 16 billion during output generation. The setup is intended to reduce the amount of computation needed at each stage without shrinking the model's overall parameter count.

The model also supports native visual understanding, so it can process images alongside text. DeepSeek says the new architecture combines new pretraining methods with larger-scale reinforcement learning post-training, and its published tests place V4.1 Flash ahead of DeepSeek V4 Pro across performance, cost, speed, and total runtime.

Smaller KV cache targets lower agent costs

One of the biggest changes is the model's memory footprint. DeepSeek says V4.1 Flash needs one-quarter of the HBM and one-eighth of the SSD storage required for the previous generation's KV cache. KV cache stores information from earlier tokens so a model can process long or repeated contexts more efficiently.

That matters particularly for agent workloads, where long-running sessions and cache hits can make memory and inference costs a large part of the total bill. DeepSeek says reducing the persistent cache footprint is designed to make those workloads cheaper to serve.

V4.1 Flash is live on the DeepSeek API

The new model is available through the DeepSeek API under the model name deepseek-flash. Older deepseek-v4-flash and deepseek-v4-flash-vision-exp identifiers are being temporarily routed to V4.1 Flash for compatibility.

DeepSeek is also phasing out V4 Pro. Starting at 04:00 UTC on September 14, requests sent to deepseek-v4-pro will route to V4.1 Flash and use V4.1 Flash pricing until V4.1 Pro launches.

For developers following the company's recent model changes, this release builds directly on the earlier DeepSeek V4-Flash API public beta and agent benchmarks. The new release moves beyond that earlier Flash generation with native multimodal support and a new architecture.

API pricing drops with the new model

DeepSeek has also changed its API pricing for V4.1 Flash. The company says off-peak rates are half the peak rates, with the new pricing taking effect at 04:00 UTC on September 10. Neowin reports the published rates as $0.003 per million input tokens for cache hits, $0.15 for cache misses, and $0.60 for output during off-peak periods. Peak rates are $0.006, $0.30, and $1.20 respectively.

That puts the release into the broader pricing picture covered in Saganote's DeepSeek Pricing 2026 guide, where Flash and Pro pricing can be compared with the rest of DeepSeek's lineup.

Open weights and wider deployment

DeepSeek says it will support the open-source community in adapting V4.1 Flash for inference and plans to explore additional deployment options. The model is available through Hugging Face, and the published model documentation lists support for contexts of up to one million tokens.

The release also arrives as DeepSeek prepares for a larger expansion of its business. Saganote recently covered the company's funding plans and potential 2027 IPO, adding another piece of context to the company's push toward larger-scale AI infrastructure and deployment.

Model routing changes are coming

Developers using `deepseek-v4-pro` should account for the September 14 routing change. DeepSeek says those requests will move to V4.1 Flash at V4.1 Flash rates until V4.1 Pro launches.

DeepSeek V4.1 Flash is therefore more than a new model name. It introduces native multimodal input, a redesigned architecture, a smaller KV-cache footprint, lower API pricing, and a planned transition away from V4 Pro. The next major change will be how DeepSeek positions V4.1 Pro once it arrives.


Share this
Saganote

About Author

Saganote

Saganote is an independent technology publication covering artificial intelligence, cybersecurity, startups, software, consumer technology, and innovation. Our editorial team researches, writes, and reviews original news, analysis, and explainers to provide accurate, timely, and well-sourced coverage of the technology industry.