Is DeepSeek Building Its Own AI Chip?
Inside the July 2026 Reuters Report

If you are tracking whether custom AI inference silicon can finally break the Nvidia tax, the global picture shifted again in July 2026: OpenAI and Broadcom taped out Jalapeño—a purpose-built LLM inference ASIC aimed at roughly 50% lower serving cost—while on July 7 Reuters reported that DeepSeek has quietly started its own inference chip program despite already deep integration with Huawei Ascend. This article leads with the economics and TCO framework Western readers care about, then covers Reuters sourcing credibility, Liang Wenfeng (DeepSeek CEO) quotes, Alibaba T-Head's eight-year Zhenwu roadmap, the July 2026 global vendor table, five structural drivers, inference-vs-training trade-offs, a six-step enterprise playbook, five FAQs, and sourcing risks. Compute options are on the pricing page.

Core questions answered (research cutoff 2026-07-09)
Question Answer
Is DeepSeek really building custom silicon? Most likely yes, still early. Reuters (2026-07-07) says the project started roughly a year ago; DeepSeek is talking to design houses, foundries, and memory vendors and hiring chip engineers off public job boards. No official announcement yet.
Did Liang Wenfeng announce it? No. He framed export bans and compute hunger as strategic pressure in 2023–2024 interviews—that is motivation, not a product launch.
What about Alibaba? Production-scale, not rumor. T-Head Zhenwu 810E is shipping; 560,000+ units delivered; annualized revenue in the tens of billions of RMB.
Latest status? DeepSeek: early R&D plus ~$7.4B funding round with chip-related use of proceeds. Alibaba: mass production. Globally: every major AI company is pursuing custom inference silicon.
Why does anyone build chips? Economics first. Inference is the recurring "rent" of AI commercialization. Custom ASICs at scale can cut 30–65% TCO versus GPUs. Supply-chain control and hardware-software co-design follow.

What Reuters reported (July 2026)

  • DeepSeek is developing a proprietary AI chip focused on inference, not training.
  • The effort reportedly began around mid-2025 ("about a year ago" in the report) and remains early stage.
  • The company is in discussions with chip design firms, wafer foundries, and memory suppliers.
  • Recruiting for chip design engineers has accelerated in recent months, largely off public job boards via direct outreach.
  • Success would reduce dual dependence on Nvidia and Huawei Ascend—even though DeepSeek already adapted V4 to Ascend and used Ascend for part of V4-Flash training.

Pain points: why readers misread this story

  • Founder quotes ≠ product announcements: Liang Wenfeng discussed compute constraints; Reuters described corporate actions. Do not merge them into "Liang officially announced a chip."
  • Partnership and self-build are not opposites: Mid-2026 analysis that DeepSeek would lean on Ascend short term is compatible with a long-term ASIC bet—cooperation is live; self-build is early.
  • Training vs inference: The rumored chip is an inference ASIC, not a direct challenge to Nvidia's training stronghold.
  • Indirect signals matter: A June 2026 external funding round of roughly 51 billion RMB (~$7.4B) included custom AI chips and domestic compute centers; the UE8M0 FP8 format is widely read as hardware-software co-design for domestic accelerators.
DeepSeek chip credibility assessment
Dimension Assessment
Source tier High. Reuters standard "three people familiar with the matter" phrasing.
Official confirmation None. As of research cutoff, no press release or social post from DeepSeek.
Circumstantial evidence Strong. $7.4B funding allocation, quiet engineer hiring, UE8M0 FP8 co-design signals.
Recommended wording Write "according to Reuters and other outlets, DeepSeek has started an inference ASIC project." Do not write "Liang Wenfeng officially announced custom silicon." Flag early stage and lack of confirmation.
DeepSeek compute and silicon timeline
Date Event
2023–2024Liang Wenfeng (DeepSeek CEO) in Waves interviews: export bans are the top challenge; relentless compute hunger
2025-01DeepSeek R1 released, trained on Nvidia H800 (export-restricted since late 2023)
Mid-2025Custom chip project reportedly initiated
2026-04DeepSeek V4 adapted to Huawei Ascend; V4-Flash partial training on Ascend
2026-06First external funding round ~$7.4B, proceeds include custom chips
2026-07-07Reuters exclusive: DeepSeek developing custom inference chip
2026-07The Information: Zhipu AI also evaluating custom silicon

Key Liang Wenfeng quotes (Waves 2023/2024 interviews, compute-related)

  • "Our real challenge has never been capital—it is the export ban on advanced chips." — July 2024
  • Domestic vs overseas training efficiency gaps stack to roughly 4× more compute needed for equivalent results.
  • "Many domestic chips fail because they lack a supporting developer community… China still needs people standing at the technology frontier."
  • "For researchers, hunger for compute is endless… we deliberately deploy as much compute as we can."

Relation to the chip rumor: Liang Wenfeng never publicly announced "DeepSeek will build chips." Reuters reported hiring and supplier talks—corporate behavior, not a founder launch event. Keep long-term founder framing separate from official project disclosure.

Alibaba: eight years of T-Head, not a headline surprise

  • September 2018 Cloud Summit: C-SKY and Damo Academy chip teams merged into T-Head Semiconductor; Jack Ma personally named the unit, signaling a long-term chip commitment; Zhang Jianfeng called silicon a group-level strategic priority.
  • Do not write "Jack Ma recently said Alibaba will build chips." Accurate framing: Ma set strategy in 2018; Joe Tsai explained export-control pressure in 2024; CEO Eddie Wu disclosed 2026 production metrics.
Jack Ma vs Joe Tsai vs Eddie Wu: public chip-related statements
Leader Role Chip-related public statements
Jack Ma 2018 strategic decision-maker Named T-Head, elevated chips to group strategy; reduced public appearances after stepping down as chairman in 2019
Joe Tsai Current chairman 2024 podcast: U.S. chip export limits "clearly affect" Alibaba Cloud; China AI ~two years behind the U.S.; long-term belief China will develop advanced semiconductors; export controls contributed to shelving an Alibaba Cloud spin-off
Eddie Wu Current CEO FY2026 earnings call: T-Head AI chips cumulative delivery 470,000+ units, annualized revenue in the tens of billions of RMB; T-Head IPO not ruled out
Alibaba T-Head Zhenwu product line
Model Timing Highlights
Hanguang 8002019Early AI inference accelerator
Zhenwu 810EJan 2026Training + inference; 96GB HBM2e; performance between Nvidia A800 and H20; in mass production
Zhenwu M8902026144GB memory, 800GB/s die-to-die link, ~3× 810E performance
Zhenwu V900Planned 2027 Q3216GB memory, 1200GB/s interconnect
Zhenwu J900Planned 2028 Q3Next-gen parallel compute architecture

2026 commercial metrics: cumulative shipments 560,000+; annualized revenue tens of billions of RMB; customers include Alibaba Cloud and China Unicom, with 400+ enterprises on Zhenwu clusters; T-Head registered capital raised to 1 billion RMB (June 2026); Alibaba pledged 380 billion RMB over three years to cloud and AI infrastructure. WSJ reported newer chips aim for CUDA compatibility to lower migration friction; manufacturing shifted from early TSMC flows toward domestic foundry (industry consensus points to SMIC 7nm-class nodes).

AI companies building chips is a global pattern, not a China-only story. TrendForce (2026): hyperscaler custom AI chip shipment growth hit 44.6%, well above general-purpose GPU growth at 16.1%—custom silicon is outpacing GPUs on growth for the first time in a meaningful way.

Major AI silicon programs as of July 2026
Company Program Stage Focus Key numbers / events
DeepSeekUnnamed inference ASICEarly R&DInference$7.4B funding; quiet hiring; no official confirmation
Alibaba (T-Head)Zhenwu 810E / M890Mass productionTrain + infer560,000+ units shipped; annualized revenue tens of billions of RMB
HuaweiAscend 950 seriesProductionTrain + inferDeepSeek V4 adaptation; surging orders
OpenAIJalapeño (with Broadcom)Tape-out complete, pending deployInference9-month design-to-tape-out; Azure deployment targeted end of 2026
GoogleTPU v6/v7Large-scale commercialTrain + inferGemini end-to-end on TPU
AmazonTrainium3 / InferentiaCommercialTrain + inferAnthropic large-scale Trainium adoption
MicrosoftMaia 100DeployingInferencePowers Azure / OpenAI workloads
MetaMTIAInternal deployInferenceRecommendation systems; prior redesign
AnthropicSamsung custom chip talksExploratoryTBDThe Information, July 2026
Zhipu AICustom chip evaluationEarlyInferenceThe Information, July 2026

Global milestones: 2026-06-24 OpenAI + Broadcom Jalapeño announcement; 2026-07-02 Anthropic reportedly in Samsung 2nm custom-chip talks; 2026-07-07 Reuters DeepSeek inference ASIC; same day The Information on Zhipu evaluating custom silicon.

One-line answer: The fight moved from "who has the best model" to "who has the cheapest, most controllable compute."

Five drivers (importance order)

  1. Economics: inference is AI's recurring rent — Training is a capex spike; inference scales linearly with daily active users. At ChatGPT-scale usage, inference spend exceeds training. Morgan Stanley once estimated ~$852M hardware for a 24,000-GPU Blackwell cluster vs ~$99M for an equivalent Google TPU footprint. SemiAnalysis and Bernstein peg custom ASIC TCO advantage at 40–65% versus GPUs; hyperscalers can cut per-token cost 30–40%. Nvidia datacenter GPU gross margins exceed 70%—custom silicon converts a permanent GPU tax into upfront R&D.
  2. Supply-chain security and geopolitics — U.S. export controls on advanced AI chips to China; Chinese policy pushing domestic compute. Security here means predictable supply, not slogans alone.
  3. Hardware-software co-design — DeepSeek UE8M0 FP8 and MLA; OpenAI Jalapeño tuned for ChatGPT serving; Google TPU bound to TensorFlow/JAX. General GPUs trade efficiency for flexibility; ASICs trade flexibility for known-workload efficiency.
  4. Competitive leverage — Even partial self-supply strengthens Nvidia negotiation and supports a full-stack narrative.
  5. Energy and sustainability — Inference ASICs optimize performance-per-watt by stripping unused GPU logic.
Training vs inference: why most programs start with inference ASICs
Dimension Training Inference
WorkloadDynamic, experimental, frequent architecture changesStatic model, predictable request patterns
Software stackDeep CUDA moat (cuDNN, NCCL, Nsight)Hand-tuned kernels for fixed models feasible
Chip needsPeak FLOPS plus programmabilityThroughput, latency, cost per token
Economic scaleLarge one-time cluster spend24/7 continuous, often larger aggregate spend
ExamplesNvidia H100/B200 dominanceTPU (partial), Trainium, Maia, Jalapeño, rumored DeepSeek ASIC

Bottom line: training remains Nvidia territory; inference is the custom ASIC battleground. Security narratives and cost savings both matter—the economics case stands on its own.

  1. Grade your sources: Prioritize Reuters, WSJ, OpenAI official posts, and Alibaba earnings calls. For DeepSeek silicon, label "reportedly / not officially confirmed"—never "proven."
  2. Split training vs inference demand: Model R&D is training-heavy; online products are inference-heavy. Procurement strategy diverges completely.
  3. Map current dependencies: List Nvidia GPU, Huawei Ascend, and cloud API share; score export-control and domestic-replacement policy impact.
  4. Model inference TCO: Build a three-year view with per-token cost, cluster capex, power, and cooling. Benchmark against the 30–65% ASIC TCO band (SemiAnalysis/Bernstein framing).
  5. Score supply-chain risk: Rate single-vendor reliance (Nvidia), geopolitical exposure, and CUDA lock-in on three axes.
  6. Draft a hybrid roadmap: Keep general GPUs or cloud rental for training; phase custom ASICs or domestic accelerators into inference; isolate bare-metal dev nodes at the edge.
  7. Quarterly refresh: This topic can move every 2–4 weeks—stamp last updated and bump dateModified when upstream news shifts.

  • DeepSeek funding: June 2026 external round ~51 billion RMB (~$7.4B), proceeds include custom AI chips and domestic compute centers.
  • Alibaba T-Head shipments: 560,000+ cumulative units (H1 2026); annualized revenue tens of billions of RMB; Zhenwu 810E performance between Nvidia A800 and H20.
  • Custom silicon growth: TrendForce (2026) hyperscaler custom AI chip shipment growth 44.6% vs general GPU 16.1%.
  • TCO advantage band: Multi-year hyperscale inference deployments: custom ASIC 40–65% TCO edge over general GPUs (SemiAnalysis, Bernstein).
  • Nvidia gross margin: Datacenter GPU margins above 70% (Breakingviews/Reuters citations).

Editorial risk reminders

  • Do not write "confirmed" before DeepSeek issues an official statement—use "reportedly" or "sources said."
  • Keep training and inference separate—the tables above exist for that reason.
  • Jack Ma timing: emphasize the 2018 strategic decision, not "Ma recently said."
  • Refresh cadence—last updated: 2026-07-10.

Risks and uncertainty: Early programs fail (Meta MTIA was redesigned); model architecture shifts can obsolete ASICs; unconfirmed DeepSeek reporting may contain sourcing gaps.

Q: Is the DeepSeek custom chip report credible?
Reuters on July 7, 2026 cited three people familiar with the matter—a high-credibility bar. DeepSeek has not officially confirmed. The project is early stage and targets inference, not training.

Q: Did Liang Wenfeng publicly announce DeepSeek would build chips?
No. In 2024 he said export bans on advanced chips are the biggest challenge and stressed deploying compute, but he never announced a custom silicon program.

Q: Who at Alibaba is driving the chip strategy?
Jack Ma named T-Head in 2018 as a group-level bet. Joe Tsai later framed export controls as a forcing function. CEO Eddie Wu disclosed 2026 mass-production metrics. Alibaba chipmaking is mature business, not a fresh rumor.

Q: Why build inference chips first instead of training chips?
Inference workloads are stable, large-scale, and continuous—ideal for ASIC tuning. Training needs CUDA depth and flexibility where Nvidia still leads.

Q: Is custom silicon mainly about national security or saving money?
Both. Near term, cutting inference cost and supply-chain risk is most urgent; geopolitics accelerated an economic case that already existed.

DeepSeek had not officially confirmed a chip program as of this writing. This article draws on Reuters, WSJ, OpenAI official materials, Waves interviews, and public earnings—not investment advice.

Primary references—re-open links after upstream updates to verify:

OpenAI Official: Jalapeño inference chip with Broadcom

Caixin Global: Alibaba Zhenwu 810E analysis

SCMP: Joe Tsai on chip export restrictions

Running large-model development and agent workflows on a personal laptop or shared office network creates compute contention, inconsistent environments, and weak audit trails. Laptop sleep kills long sessions; shared networks make consistent compute tiers hard to enforce.

If you need isolated 24/7 dev terminals, dedicated Apple Silicon, and controlled egress policies, layering iOS CI and AI agents onto dedicated Mac Mini bare-metal nodes beats gambling on compute-starved mixed environments: NOVAKVM offers multi-region Mac Mini M4 / M4 Pro flexible terms with default SSH—suited to iOS CI/CD and AI agent automation. See the pricing page, order on the order page, and deploy via the help center.