DeepSeek V4.1 Flash: Native Multimodal Model Sets New Speed Records
deepseek multimodal
| Source: Mastodon | Original article
DeepSeek's V4.1 Flash multimodal model, launched in a surprise beta on Sep 8, 2026, set a new speed record of 420 tokens per second without compromising accuracy.
DeepSeek has quietly opened a beta for its latest large‑language model, V4.1 Flash, marking the company’s first truly native multimodal offering. Launched on 8 September 2026, the model is now accessible through the DeepSeek API under the identifier deepseek‑flash. Unlike earlier versions that added image handling as an afterthought, V4.1 Flash supports text and image inputs and outputs directly from the factory, and it does so while claiming a record‑breaking 420 tokens per second throughput without a measurable drop in accuracy.
The speed boost stems from a re‑trained Mixture‑of‑Experts architecture that expands the backbone to 552 billion parameters and stretches the context window to one million tokens. DeepSeek reports a 2‑4× increase in throughput compared with the prior V4 Flash model, while keeping the same “Flash” pricing tier. A key efficiency lever is a compressed key‑value cache, which trims memory usage and lowers inference cost. The launch also retires the older V4‑Flash and V4‑Flash‑Vision‑Exp variants.
Why it matters is twofold. First, the combination of native multimodality, massive context length and unprecedented speed narrows the performance gap between DeepSeek and rivals such as Anthropic’s Claude or the emerging topological‑network models from Paris‑based Arlequin AI. Second, the cost‑effective cache design could make high‑throughput multimodal services more affordable for developers, potentially accelerating adoption in areas like real‑time visual assistants and large‑scale document analysis.
As we detailed in our deep‑dive on DeepSeek V4.1 Flash on 10 September, the model’s architecture and pricing were already attracting attention. The next steps to watch are independent benchmark results, user‑feedback from the beta, and whether DeepSeek will extend the same efficiency gains to future releases or open the model to broader commercial licensing.
Sources
Back to AIPULSEN