DeepSeek unveils DeepSeek-V4.1-Flash, its smallest 552B‑parameter model with 1 M‑token context
deepseek startup
| Source: Techmeme | Original article
Chinese AI startup DeepSeek launched DeepSeek‑V4.1‑Flash, its smallest model featuring a new Causal Encoder‑Decoder design, a 552‑billion‑parameter backbone and a 1‑million‑token context window.
DeepSeek, the Chinese AI startup, unveiled its latest offering on Thursday: DeepSeek‑V4.1‑Flash, the smallest model in the company’s new Causal Encoder‑Decoder family. The model packs a 552 billion‑parameter Mixture‑of‑Experts (MoE) backbone but activates only 8 billion parameters for input and 16 billion for output, keeping the active footprint modest while retaining the scale of a flagship system. Weighing about 510 GB, the model supports contexts of up to one million tokens and is available immediately via the DeepSeek API with native multimodal capabilities. DeepSeek released the weights under an MIT licence, positioning the model as both a research resource and a commercial service.
The launch matters because it demonstrates a shift toward “efficient scale” – large‑parameter backbones that can be throttled to lower compute and cost for inference. A one‑million‑token window opens new possibilities for long‑form reasoning, document analysis and multimodal tasks that previously required stitching together multiple prompts. By offering an open‑source licence, DeepSeek also invites the broader community to experiment, potentially accelerating innovation around MoE architectures and long‑context models.
What to watch next includes performance benchmarks that will reveal whether the reduced active parameter count translates into the promised speed and cost advantages. Industry observers will also track pricing on the DeepSeek API and any uptake by developers building large‑context or multimodal applications. Finally, the release may spur rival firms to announce comparable “flash” variants, intensifying competition in the race to deliver high‑capability models that are affordable to serve at scale.
Sources
Back to AIPULSEN