In a five-day window, three of China's top AI labs made moves that look contradictory on the surface. DeepSeek raised API prices by as much as 1,100% on certain tiers. Alibaba, the same week, open-weighted a 2.4-trillion-parameter flagship it had never released before. Zhipu AI shipped GLM-5.3, boosting coding benchmarks by roughly 6x on the exact same base model — no retraining. Together, these moves signal a shift from competing on price alone to competing on pricing power. This article covers the timeline, rate sheet, Qwen license, GLM-5.3 scores, strategy breakdown, head-to-head prices, unverified claims, a six-step checklist, and FAQ. Cross-read DeepSeek V4 GA, Qwen3.8-Max launch, and GPT-5.6 Luna cuts. Node pricing is on the pricing page.
[ SECTION_01 ] // TIMELINE What happened in two weeks, and where teams get stuck
Three questions usually block a decision: which billing line is the 1,100% headline; whether open weights mean free commercial use or a geo ban; and whether DeepSeek is still the cheapest frontier-class API. Pin the dates first.
| Date | Event |
|---|---|
| Jul 16, 2026 | Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny |
| Jul 30, 2026 | OpenAI cuts GPT-5.6 Luna, its cheapest tier, by 80% |
| Aug 2–3, 2026 | Alibaba previews, then launches, Qwen3.8-Max as a hosted API |
| Aug 6–7, 2026 | OpenAI makes Luna the free default with unlimited text chats |
| Aug 10, 2026 | Meta releases Muse Glimmer (30B, Apache 2.0) and teases open weights for flagship Muse Spark 1.2 |
| Aug 12, 2026 | Alibaba publishes Qwen3.8-2.4T-A95B open weights on Hugging Face / ModelScope; xAI ships Grok 4.6 |
| Aug 13, 2026 | DeepSeek-V4-Pro goes GA and announces a price increase effective Aug 17; Google ships discounted Gemini 3.7 Flash |
| Aug 14, 2026 | Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base |
| Aug 17, 2026, 00:00 Beijing time | DeepSeek's new pricing takes effect |
Zoom out and the picture sharpens. While Chinese labs raised prices and opened flagship weights, US labs cut prices and went free at the consumer layer. That is two sides of the same pricing fight. Chinese financial media has started calling the domestic cadence 「一周三更」— three major updates a week — with MiniMax H3 in the same rotation as DeepSeek, Alibaba, and Zhipu.
- Friction 1 — headline math: 「11x」, 「over 1,100%」, and 「350%」 are all true. They describe different lines: cache-hit input, output, and cache-miss input.
- Friction 2 — license folklore: Claims that Alibaba banned downloads from the US, EU, UK, and South Korea are false. The published text has no territorial clause. The real gates are revenue and product type.
- Friction 3 — official is no longer cheapest: At peak hours, DeepSeek's official API now prices above some resellers (GMI Cloud, Novita, and others currently list V4 Pro below the official peak rate).
[ SECTION_02 ] // MATRIX The numbers: DeepSeek rates, Qwen specs, and a price matrix
New DeepSeek rates took effect at 00:00 Beijing time on Aug 17. Peak hours are 9am–12pm and 2pm–6pm Beijing time (01:00–04:00 and 06:00–10:00 UTC).
| Billing item | Old | New off-peak | New peak | Peak increase |
|---|---|---|---|---|
| V4-Flash cache hit (input) | ¥0.02 | ¥0.05 | ¥0.10 | ~400% |
| V4-Flash cache miss (input) | ¥1.0 | ¥1.5 | ¥3.0 | 200% |
| V4-Flash output | ¥2.0 | ¥4.5 | ¥9.0 | 350% |
| V4-Pro cache hit (input) | ¥0.025 | ¥0.15 | ¥0.30 | ~1,100% |
| V4-Pro cache miss (input) | ¥3.0 | ¥4.5 | ¥9.0 | 200% |
| V4-Pro output | ¥6.0 | ¥13.5 | ¥27.0 | 350% |
The 1,100% figure applies to peak-hour cache-hit input — the tier that started closest to free. Output, which dominates most real bills, rose 350%. Independent cost modeling found a realistic heavy-usage workload (roughly 84M tokens/month, mostly off-peak, half cache hits) sees a bill increase closer to 1.8x.
| Spec | Detail |
|---|---|
| Parameters | 2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared) |
| Context window | 262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max defaults to 1M |
| Release cadence | Preview Aug 2 → API live Aug 3 → open weights Aug 12 |
| API pricing (international) | $2/M input, $6/M output |
| License | Not Apache 2.0 — a custom Qwen3.8-Max License |
| Why it matters | First time Alibaba has open-weighted a Max-tier flagship; Qwen3.5 / 3.6 / 3.7 Max stayed API-only |
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6% | 28.3% | +23.7 pts |
| DeepSWE v1.1 | 46.2% | 66.9% | +20.7 pts |
| Agents' Last Exam (CLI) | 23.8% | 28.5% | +4.7 pts |
| CyberGym | 77.2% | 84.5% | +7.3 pts |
| AutomationBench | 26.2% | 48.2% | +22.0 pts |
GLM-5.3 still trails GPT-5.6 Sol (34.6%) and Claude Fable 5 (33.7%) on Terminal-Bench 3.0. It is a top open-weight result, not an outright frontier win.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Open weights? |
|---|---|---|---|
| DeepSeek V4-Pro (peak) | ¥9.0 (~$1.26) | ¥27.0 (~$3.78) | No |
| DeepSeek V4-Pro (off-peak) | ¥4.5 (~$0.63) | ¥13.5 (~$1.89) | No |
| Qwen3.8-Max (international API) | $2.00 | $6.00 | Yes (custom license) |
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 | No |
| Claude Opus 5 (implied, per Alibaba's comparison ratio) | ~$5.00 | ~$25.00 | No |
Even after the hike, DeepSeek V4-Pro off-peak is still well below Claude Opus 5. It is no longer the outright cheapest option. Qwen3.8-Max international pricing and OpenAI Luna both undercut DeepSeek off-peak. 「Chinese model = cheapest model」 held for most of 2025 and early 2026. It is not a safe assumption anymore.
[ SECTION_03 ] // DEEP_DIVE Three strategies, two price wars, and what remains unverified
1. DeepSeek: from flat-rate cheap to time-of-day pricing — a capacity problem, not a strategy pivot. The easy misread is that China's cheapest model finally caved to margin pressure. The structure reads more like a company making compute constraints visible on the price sheet. Flat, always-cheap pricing worked as acquisition while GPU capacity kept pace. Once usage grew exponentially and capacity did not, something had to become explicit. 「Encouraging more flexible workload scheduling」 is corporate-speak for peak-hour compute is now scarce. At peak, the official API is now higher than several resellers. The assumption that the official API is always the cheapest way to run DeepSeek has been broken for the first time.
2. Alibaba: open weights buy ecosystem goodwill; a custom license protects the revenue ceiling. Publishing the 2.4T checkpoint and attaching a custom license — not Apache 2.0 — happened together. Any Model-as-a-Service or AI Work Assistant business earning over $50 million in any 12-month period must negotiate a separate commercial license. Products with 100M+ monthly active users or $20M+ in monthly revenue must prominently display the model name. The bet: win developer mindshare internationally, keep leverage over firms that can build a competing inference business. That is a different bet than Meta's Muse Glimmer, which ships under unrestricted Apache 2.0.
One rumor is worth killing explicitly. Claims that the license bans downloads from the US, EU, UK, and South Korea are false. The published LICENSE file contains no geographic clause. In a release cycle this fast, the LICENSE file — not the announcement thread — is the primary source.
3. GLM-5.3: no new base model, just a bigger post-training bet. Same 743B-parameter base as GLM-5.2, no retraining, and a roughly 6x jump on Terminal-Bench 3.0 (4.6% → 28.3%) from scaling reinforcement-learning environments. As pretraining scaling laws show diminishing returns, post-training RL scale is becoming an independent lever with a much lower cost floor than retraining a new foundation model. Mid-tier labs without OpenAI-scale compute can still close the gap on agentic and coding benchmarks.
What is disputed or unverified:
- The 1,100% headline is accurate and incomplete. It applies only to peak-hour cache-hit input. Output — the cost that dominates most real bills — rose 350%.
- Claims that Qwen3.8-Max runs on Alibaba's in-house Zhenwu M890 chips and 「Pangu AL128」 supernodes appear in Chinese financial outlets. They have not been independently confirmed by Alibaba technical documentation or third-party benchmarks. Treat as vendor-adjacent, unverified reporting.
- GLM-5.3's reported discovery of a serious vulnerability in Cursor comes from VentureBeat and Zhipu's own disclosure. Technical details have not been made public. Read it as vendor-sourced, not independently audited.
- Reports that China's Ministry of Commerce may prepare retaliatory export controls on AI or semiconductor technology are speculative and sourced to unconfirmed media reports, not an official announcement.
Why this matters: two price wars in parallel. Over the past month, China's top labs have shipped at a pace domestic media calls three updates a week. US labs ran the opposite play at the consumer layer: OpenAI cut Luna 80% on Jul 30, then made it free and unlimited on Aug 6–7; Google shipped a coding-focused model at half the price of its three-week-old predecessor on Aug 13. Chinese labs open-weight flagships and introduce tiered, higher pricing on the compute-constrained top end. US labs race toward free and cheap at the consumer end.
There is also a geopolitical layer. Moonshot's Kimi K3 already drew US security scrutiny. Some analysts read Alibaba's choice of this window to open-weight a 2.4T flagship as a move to lock in international mindshare before any potential regulatory tightening. That is an informed interpretation, not a confirmed fact. See also Kimi K3 open-weight explainer.
[ SECTION_04 ] // CHECKLIST Six steps to verify the price war before you change spend
- Split the headline into billing lines. Write cache-hit input (~1,100%), output (350%), and cache-miss input (200%) as separate rows. Do not budget from a single multiple.
- Compare official peak rates with resellers. Check GMI Cloud, Novita, and other listings against DeepSeek's peak sheet. Official is no longer the default cheapest path.
- Read the Qwen LICENSE file, not the launch thread. Confirm there is no geo ban. Check whether your product hits Model-as-a-Service / AI Work Assistant plus $50M in any 12 months, or 100M MAU / $20M monthly revenue.
- Label GLM-5.3 scores as vendor-run. Terminal-Bench 3.0 at 28.3% has no published independent rerun. When you cite Sol 34.6% and Fable 5 33.7%, keep the source visible.
- Reprice the bill with off-peak and cache-hit rate. Move batch jobs, evals, and overnight agent loops off peak. For heavy users, model ~1.8x, not 11x.
- Self-hosting open weights needs a stable host. Long-running evals, agent orchestration, and license-compliance logs need a 24/7 reproducible environment. Laptop sleep and drifting VMs break the audit trail. Spin a trial node from the pricing and order pages and record versions, license text, and logs.
China open-weight + pricing cluster — triage sketch (verify against primary sources)
deepseek V4-Pro peak cache-hit input ~1100% · output +350% · effective Aug 17 00:00 CST
peak Beijing 09:00-12:00 and 14:00-18:00 · UTC 01:00-04:00 and 06:00-10:00
qwen 2.4T-A95B weights live Aug 12 · custom license · no geo ban
glm53 same 743B base · TB3.0 4.6% → 28.3% · vendor-run, no 3rd-party rerun
rumor Zhenwu M890 / Cursor vuln / MOFCOM export controls = unverified
[ SECTION_05 ] // FACTS_FAQ Citable figures, FAQ, and a host for the eval loop
- DeepSeek V4-Pro peak cache-hit input: ¥0.025 → ¥0.30, about 1,100%; output ¥6.0 → ¥27.0, 350%.
- Qwen3.8-2.4T-A95B: 2.4T total / 95B active; 262,144 tokens native, ~1.01M extendable; international API $2 / $6.
- Qwen license gates: MaaS / AI Work Assistant over $50M in any 12 months needs a separate license; 100M MAU or $20M monthly revenue needs prominent naming; no geographic ban.
- GLM-5.3: same 743B base as 5.2; Terminal-Bench 3.0 4.6% → 28.3%; still behind Sol 34.6% and Fable 5 33.7%.
- Heavy-user bill: third-party modeling ~1.8x, far below the 11x headline.
- FX reference: about ¥7.15 / $1, approximate only.
FAQ:
- Q: Is DeepSeek still cheaper than GPT-5.6 or Claude? A: Off-peak still undercuts Claude Opus 5. Luna ($0.20 / $1.20) and Qwen3.8-Max international ($2 / $6) now undercut DeepSeek off-peak on at least one dimension.
- Q: Can I use Qwen3.8-Max open weights commercially for free? A: Personal and internal use are fine. The $50M MaaS / AI Work Assistant gate is the catch.
- Q: Is Qwen3.8-Max banned for US, EU, or UK users? A: No. The published license has no geographic restriction.
- Q: What changed between GLM-5.3 and GLM-5.2? A: Same 743B base. Gains come from scaled post-training RL, not a new foundation model.
- Q: Will Meta open-source Muse Spark 1.2, not just Glimmer? A: Not yet. Glimmer is a 30B distill. Spark 1.2 remains a stated intention.
Primary sources used while writing. Verify the latest official pricing and license terms before republishing. Details flagged above as unverified (domestic chip claims, the Cursor vulnerability report, and export-control rumors) have not been independently confirmed. Figures reflect public information as of 17 Aug 2026.
DeepSeek official pricing page (treat the live page as source of truth):
https://api-docs.deepseek.com/quick_start/pricing
Alibaba Qwen official model repositories on Hugging Face:
ModelScope model repositories:
Zhipu GLM-5.3 official technical page:
Real limits of the usual workarounds: Budgeting from a single 1,100% multiple overstates off-peak heavy-user bills and understates peak batch shock. Treating Qwen weights as Apache 2.0 and standing up a Model-as-a-Service product can trip the custom license. Running 2.4T-class evals and long agent loops on a laptop that sleeps, or on a drifting VM, breaks reproducibility and the license audit trail.
For teams that need dedicated Apple Silicon, 24/7 uptime, and day/week/month elasticity to run open-weight evals, agent orchestration, and long inference on a stable macOS host, NOVAKVM Mac Mini bare-metal rental is usually the better production path. Keep pricing narratives on the vendors' rate cards; keep execution on a resident bare-metal node. Compare tiers on the NOVAKVM pricing page, start a trial on the order page, and use the help center for remote sessions.