
- DeepSeek V4-Flash costs $0.14/M input tokens — cheaper than GPT-5.4 Nano and 36x less than GPT-5.5’s $5/M
- V4-Pro is the largest open-weights model ever at 1.6 trillion parameters, with performance trailing frontier models by only 3-6 months
- Both models use 1M-token context with 90% less compute and 93% less KV cache than the previous generation
- Open MIT license means any company can self-host frontier-class AI, fundamentally reshaping enterprise AI economics
On April 24, Chinese AI lab DeepSeek dropped two preview models that sent shockwaves through the AI industry — not because they’re the absolute best, but because they’re terrifyingly close to the best at a fraction of the price. DeepSeek-V4-Pro and V4-Flash represent the most aggressive price-performance play in AI history, and they’re fully open-weights under an MIT license.
The Numbers That Change Everything
V4-Pro: The Largest Open-Weights Model Ever
DeepSeek-V4-Pro packs 1.6 trillion total parameters with 49 billion active parameters in a Mixture-of-Experts architecture. That makes it the largest open-weights model ever released — bigger than Kimi K2.6 (1.1T), GLM-5.1 (754B), and more than twice the size of DeepSeek’s own V3.2 (685B). It supports 1 million tokens of context and costs just $1.74 per million input tokens and $3.48 per million output tokens.
For context, OpenAI’s GPT-5.5 charges $5/$30 for the same input/output volumes, and Anthropic’s Claude Opus 4.7 charges $5/$25. DeepSeek V4-Pro delivers comparable intelligence at roughly one-sixth the cost.
V4-Flash: Sub-Dollar Frontier AI
The Flash variant is where the economics get truly disruptive. At $0.14 per million input tokens and $0.28 per million output tokens, it undercuts even OpenAI’s GPT-5.4 Nano ($0.20/$1.25). With 284 billion total parameters and just 13 billion active, it scores 47 on Artificial Analysis’s Intelligence Index — nearly double the median score of 28 for comparable open-weight models — while running at 84 tokens per second.
Business Insight — At these prices, the cost barrier between closed and open AI models has effectively collapsed. A startup spending $10,000/month on GPT-5.5 API calls could achieve near-equivalent results with DeepSeek V4-Pro for under $1,700 — or switch to V4-Flash for under $100.
The Efficiency Breakthrough Behind the Price
90% Less Compute at 1M Context
The pricing isn’t a loss-leader strategy — it’s backed by genuine architectural innovation. According to DeepSeek’s technical paper, V4-Pro achieves only 27% of the single-token FLOPs and 10% of the KV cache size compared to V3.2 when processing 1 million tokens. V4-Flash pushes this even further: just 10% of the FLOPs and 7% of the KV cache.
This means DeepSeek can serve inference at dramatically lower hardware costs. The efficiency gains come from architectural improvements to the Mixture-of-Experts routing and attention mechanisms, allowing far fewer parameters to be activated per token while maintaining output quality.
Huawei Ascend 950 Integration
Adding another strategic dimension, Fortune reports that DeepSeek expects to lower V4-Pro prices further later this year as Huawei scales up production of its new Ascend 950 AI processors. This signals deepening integration between DeepSeek and China’s domestic chip ecosystem — a move that could insulate the company from ongoing U.S. export controls on NVIDIA GPUs.
Business Insight — DeepSeek’s efficiency-first architecture is a direct challenge to the “scale is all you need” thesis that has driven billions in GPU investments. If near-frontier performance can be achieved with 10% of the compute, the ROI calculus for massive data center buildouts changes significantly.
Performance: Close Enough to Matter
Benchmarks Tell the Story
DeepSeek’s own benchmarks show V4-Pro-Max outperforming GPT-5.2 and Gemini 3.0 Pro on standard reasoning tasks, while falling “marginally short” of GPT-5.4 and Gemini 3.1 Pro. The company candidly acknowledges a developmental lag of approximately 3 to 6 months behind the absolute frontier. In coding benchmarks, both V4 models deliver performance “comparable to GPT-5.4.”
VentureBeat’s analysis frames it as “near state-of-the-art intelligence at 1/6th the cost,” while MIT Technology Review identified three reasons why the release matters: it proves open-weights models can stay competitive, it compresses premium AI economics into a lower band, and it forces enterprises to revisit cost-benefit calculations around closed models.
The Open-Weights Advantage
Both models ship under the MIT license, meaning any organization can download, modify, fine-tune, and self-host them. NVIDIA has already published integration guides for running V4 on Blackwell GPUs, and the Unsloth team is expected to release quantized versions that could run V4-Flash on consumer hardware — potentially even a 128GB MacBook Pro.
Business Insight — For enterprises weighing build vs. buy decisions, DeepSeek V4 eliminates the quality gap that previously justified premium API pricing. The real question is no longer “can open-source match closed models?” but “how long until open-source consistently leads?”
What This Means for the AI Market
DeepSeek V4 arrives at a critical inflection point. OpenAI just launched GPT-5.5, Google committed $40 billion to Anthropic, and Microsoft invested $10 billion in Japan’s AI infrastructure. The prevailing narrative has been that frontier AI requires frontier spending.
DeepSeek’s release challenges that narrative directly. As The Register put it, the new models offer “big inference cost savings” that could reshape how companies budget for AI. For developers and enterprises already running on DeepSeek’s API or self-hosting open-weights models, V4 represents a generational leap in what’s achievable without writing seven-figure checks to OpenAI or Anthropic.
The competitive pressure is real and immediate. If DeepSeek can deliver 85-95% of frontier performance at 15-20% of the cost — and keep improving on domestic Huawei chips — the premium pricing power of closed-model providers faces a structural challenge that no amount of investment can simply outspend.
Related
- DeepSeek V4 Arrives With Near State-of-the-Art Intelligence at 1/6th the Cost — VentureBeat
- Three Reasons Why DeepSeek’s New Model Matters — MIT Technology Review
- DeepSeek Unveils V4 Model With Rock-Bottom Prices and Huawei Integration — Fortune
- Build With DeepSeek V4 Using NVIDIA Blackwell — NVIDIA Blog
- DeepSeek V4 — Almost on the Frontier, a Fraction of the Price — Simon Willison
Sources
- TechCrunch — DeepSeek previews new AI model that ‘closes the gap’ with frontier models
- Simon Willison — DeepSeek V4: almost on the frontier, a fraction of the price
- VentureBeat — DeepSeek V4 arrives with near state-of-the-art intelligence at 1/6th the cost
- Fortune — DeepSeek unveils V4 model with rock-bottom prices and Huawei chip integration
- The Register — DeepSeek’s new models offer big inference cost savings
AI Biz Insider · AI Business EN · aibizinsider.com

댓글 남기기