DeepSeek V4.1 Flash API priced 3‑6× lower per answer than V4 Pro
deepseek
| Source: Mastodon | Original article
DeepSeek's V4.1 Flash API delivers V4 Pro‑level accuracy while costing 3‑6× less per answer and runs 2.9× faster, though multi‑step tasks may fail and unanswered queries consume 16K tokens.
DeepSeek has taken its V4‑1 Flash model from beta to general availability, announcing a pricing structure that undercuts its flagship V4‑Pro while delivering identical accuracy on a suite of benchmark tasks. The company says the new API charges $0.30 per million input tokens and $1.20 per million output tokens at peak rates, with off‑peak prices halved. In a four‑day test covering eleven answerable tasks, the model answered all 33 prompts correctly – the same score achieved by V4‑Pro – but at three to six times lower cost per correct answer and 2.9 × faster throughput. The tests also note that “no‑answer” queries still consume about 16 K tokens.
V4‑1 Flash is a 552‑billion‑parameter mixture‑of‑experts system that activates roughly 8 B parameters while reading a prompt and 16 B during generation, and it supports a 1 M‑token context window across the V4 family. The model’s design, a causal encoder‑decoder, is intended to balance speed and efficiency for developers building chat‑based applications, coding assistants, and other high‑throughput workloads.
The launch matters because it dramatically lowers the cost barrier for using a top‑tier large language model, potentially reshaping pricing dynamics in a market where providers such as Anthropic and OpenAI have faced criticism for expensive API rates. By offering comparable performance at a fraction of the price, DeepSeek could accelerate adoption of LLM‑driven services in both enterprise and consumer‑facing products.
What to watch next includes uptake metrics from early adopters, any adjustments to the off‑peak discount structure, and whether competitors respond with comparable low‑cost tiers. Follow‑up performance benchmarks and real‑world case studies will also reveal whether the speed and cost advantages translate into broader market impact. As we reported in our DeepSeek V4 Explained guide earlier this month, the V4 family’s mix of scale and pricing is already disruptive; V4‑1 Flash may cement that trend.
Sources
Back to AIPULSEN