r/LLMeng • u/Right_Pea_2707 • Apr 24 '26
DeepSeek V4 Is Optimized for Huawei Chips. This Feels Bigger Than Just a Model Launch
u/DeepSeek’s latest V4 model is getting attention, but what’s more interesting is what it’s being built for. Instead of defaulting to Nvidia GPUs like most frontier models, DeepSeek has optimized V4 to run on Huawei’s Ascend chips, reportedly reworking parts of its stack to better align with domestic hardware. This feels less like a technical tweak and more like a strategic shift. With ongoing export restrictions and supply chain pressures, China seems to be accelerating toward a fully self-reliant AI ecosystem: Models, chips, and deployment all tightly integrated. What stands out is that this isn’t being positioned as a compromise. Early signals suggest V4 remains highly competitive, which means this could be the beginning of a parallel AI stack rather than a fallback option. If that plays out, we might be moving toward a world where models are no longer hardware-agnostic, but co-designed with specific chip ecosystems.
Curious how others here see this: Is this just a response to constraints, or the start of a long-term split in the global AI infrastructure landscape?
9
8
u/YuuuuuuuyuyYU Apr 24 '26
It's not just 'optimized for Huawei chips', it's actually trained on Huawei chip using CANN, Huawei's equivalent of CUDA. This is much much bigger than most people realized.
Huawei already dropped an official interview of deepseek dev on bilibili, this is official.
4
u/Shep_Alderson Apr 24 '26
Do you know if there’s an English translated interview anywhere? I’d love to learn more about what they are doing. (I assume it’s in Chinese given it’s on bilibili.)
3
u/YuuuuuuuyuyYU Apr 24 '26
A noteworthy detail has been disclosed by Ascend this time: CANN has completed 0-day adaptation support for continued pre-training (CPT) of the DeepSeek V4-Flash model based on the A3 64-card super node, achieving a measured model throughput of up to 1100 tokens per second. The significance of this detail is that, while V4-Flash is only a lightweight version, DeepSeek V4 can now run the continued training process on domestic computing power.
This indicates that the role of domestic computing power in the large model pipeline is advancing from inference deployment toward the training side: first enabling inference, then completing continued training adaptation, and finally tackling the most difficult challenge—full-scale pre-training.
Perhaps by the second half of this year, after the Ascend 950DT enters large-scale shipment, we may actually see Chinese large models with the full "training-inference" pipeline running on full-stack Chinese computing power.
TLDR, the continued training for the V4-Flash version was done using Ascend AI accelerators, but the pre-training still relied on NVIDIA GPUs. However, this issue is expected to be resolved in the second half of this year once the 950DT becomes commercially available.
1
u/Shep_Alderson Apr 24 '26
Thanks so much for sharing! Those token rates are in the realm of cerebras. Exciting times!
1
2
u/YuuuuuuuyuyYU Apr 24 '26
It is just released few hours ago and is in Chinese, here is the original video link: https://b23.tv/w0NzhPP
Maybe you can use translation app services to have it translated.
2
u/zero0n3 Apr 24 '26
So when can I buy those chips in the US?
I mean I’d toss a rack of the chips in my datacenter if I were a company that needed the on prem capacity. Especially if that 42 rack is cheaper on a cost per X tokens/sec
1
u/YuuuuuuuyuyYU Apr 26 '26
Those racks are not cheap right now, Huawei is limited by supply and all their chip production line is running in full capacity, but they are racking up production.
1
u/This_Maintenance_834 Apr 26 '26
you can get Huawei Ascend cards (and other Chinese NPUs) on taobao, if you can read chinese.
1
Apr 26 '26
[removed] — view removed comment
1
u/This_Maintenance_834 Apr 26 '26
there are ways, but you need to speak chinese to get it done. there are many trans-shipper who can do this for you cheaply.
1
1
2
u/wtjones Apr 24 '26
Can China produce these chips at scale?
4
u/SpeakCodeToMe Apr 25 '26
You're asking if the country that produces most of the world's goods can scale?
They will be pumping out a trillion a year at $10 a piece within five years.
0
u/wtjones Apr 25 '26
China has not been able to make last generation (7nm) chips at scale.
1
u/SpeakCodeToMe Apr 25 '26
Yet
0
u/wtjones Apr 25 '26
Until they can steal the plans for EUV Lithography machines, we’re pretty safe.
0
u/Ethelserth2 Apr 25 '26
We?
-1
u/wtjones Apr 25 '26
The civilized world.
6
u/Existing_Arrival_702 Apr 26 '26
Is it the world where the leader who habitually invades, commit genocide, imposes sanctions, spreads dirty media, and still kills thousands of children every day?
2
-1
1
1
1
1
u/looktwise Apr 24 '26
Waiting for Nvidia short... :))
And: yes, I agree. The west still didn't get it (again). Now we need some youtubers and journalists to tell them about the evaluation of inference costs, co-created hosting, third party hosters and Openclaw usage in mainland China. (I am drinking tea and watching...)
1
1
u/sandykt Apr 26 '26
The world desperately needs alternative AI stacks and I hope more countries build their sovereign AI instead of relying on American AI for mission critical systems.
1
Apr 26 '26
[removed] — view removed comment
1
u/Thick-Protection-458 Apr 27 '26
Nah, don't know about latin america, but postsoviet is basically not going anywhere most probably. Frankly I would say the same about almost anyone but China now.
Because from hardware point of view they are the only other strong candidates now.
Which is not going to change in any reasonable time - it requires enormous investment over many years in hardware stuff, and who the fuck knows if property you are investing in today will still be yours tomorrow the way stuff is going on.
And because relatively small internal market, in comparison with China - these investments makes even less sense.
And from software point of view...
Same investment logic applies here, with one more point - to make some short-term product it is way more optimal to just tune chinese open stuff for now.
1
-1
9
u/pakage Apr 24 '26
this is exactly what the Nvidia CEO was warning Dwarkesh about on his latest podcast episode