Models
DeepSeek releases open-weight V4.1-Flash with native vision and 1M-token context
DeepSeek released V4.1-Flash on Sept. 10, an MIT-licensed mixture-of-experts model with 552 billion parameters, built-in image understanding and a one-million-token context window. The company also cut API prices and later reversed a plan to route requests for its V4 Pro model to the new release.
DeepSeek's change log describes V4.1-Flash as the smallest model in a new architecture family. The Hugging Face model card says it pairs a causal encoder with a decoder, activates about 8 billion parameters per token while reading a prompt and 16 billion while generating output, and stores its main key-value cache in a four-bit format to cut memory use. SiliconANGLE reported that the total parameter count is nearly double that of the V4-Flash model it replaces.
The API documentation lists off-peak prices of $0.15 per million uncached input tokens and $0.60 per million output tokens, with rates doubling during weekday peak hours. Older V4 Flash model names now route to the new model. DeepSeek initially said requests for V4 Pro would be answered by V4.1-Flash from Sept. 14, but the change log now states that V4 Pro service continues with billing unchanged.
DeepSeek's published results put the model at 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, ahead of V4 Pro and, by the company's account, slightly above Claude Opus 5. Those scores are self-reported.
Source details
- Source
- DeepSeek API Docs
Source reporting
Read the original reporting and research behind this briefing.