DeepSeek shipped V4-Pro to general availability yesterday.
Then it raised API prices by up to 1,100%.
Peak hours now cost 5x to 11x more than yesterday. Off-peak is half that. The cheap Chinese inference era just ended.
Here's what changed:
V4 Flash output tokens: $0.28/M → $1.32/M at peak
V4 Pro cached input: $0.0036/M → $0.044/M at peak
V4 Pro output: $0.87/M → $3.96/M at peak
DeepSeek processed 8 trillion tokens on a single day in August. Five trillion were free usage. Infrastructure has to be paid for eventually.
The "AI will stay cheap" assumption just collapsed.
Enterprises that built cost models on DeepSeek's rock-bottom pricing are now looking at 2x to 11x cost increases depending on workload timing. The assumption that inference costs would keep falling — that drove every AI ROI calculation in the last 12 months — just got invalidated by the company that forced everyone else to cut prices.
Two realities emerge:
1. The open-weight option matters more now. V4-Pro is MIT licensed. You can run it yourself on your own infrastructure. The cost shifts from API calls you don't control to compute you do.
2. Peak/off-peak pricing forces operational decisions. You now need to schedule batch jobs between 10:00 and 13:00 UTC, or pay double. Every AI workflow needs a cost-aware scheduler.
Audit your AI inference spend today. If you're building on DeepSeek's API, model the new pricing against your actual usage patterns. If you haven't stress-tested your AI budget against a 5x cost increase, you're not budgeting — you're hoping.
DeepSeek just raised API prices 1,100%. Your AI budget assumptions just died.
AI-Assisted Content — Produced with AI assistance and human editorial review.
Learn more
0 Comments