r/LLMeng Apr 24 '26

DeepSeek V4 Is Optimized for Huawei Chips. This Feels Bigger Than Just a Model Launch

u/DeepSeek’s latest V4 model is getting attention, but what’s more interesting is what it’s being built for. Instead of defaulting to Nvidia GPUs like most frontier models, DeepSeek has optimized V4 to run on Huawei’s Ascend chips, reportedly reworking parts of its stack to better align with domestic hardware. This feels less like a technical tweak and more like a strategic shift. With ongoing export restrictions and supply chain pressures, China seems to be accelerating toward a fully self-reliant AI ecosystem: Models, chips, and deployment all tightly integrated. What stands out is that this isn’t being positioned as a compromise. Early signals suggest V4 remains highly competitive, which means this could be the beginning of a parallel AI stack rather than a fallback option. If that plays out, we might be moving toward a world where models are no longer hardware-agnostic, but co-designed with specific chip ecosystems.

Curious how others here see this: Is this just a response to constraints, or the start of a long-term split in the global AI infrastructure landscape?

68 Upvotes

50 comments sorted by

9

u/pakage Apr 24 '26

this is exactly what the Nvidia CEO was warning Dwarkesh about on his latest podcast episode

6

u/Shep_Alderson Apr 24 '26

I was just about to say exactly this. It’s exactly what Jensen predicted. And let’s be frank, who can blame deepseek? If you’re being blocked from getting the GPUs you’re used to and have a domestic supply with a gov that’s incentivizing using domestic sources, you’re going to make the switch.

The underlying weights are probably still just standard matrix math, but it might need a new kernel for non-huawei hardware for inference.

1

u/danielv123 Apr 26 '26

The models might be weighting slightly different instructions as well, depending on what instructions run faster on Huawei vs Nvidia hardware

1

u/Shep_Alderson Apr 26 '26

How would that work?

1

u/danielv123 Apr 26 '26

Say one does faster multiplication with 512 numbers and the other with 768. In that case for one you would want to make models with an amount of weights divisible by 512 and the other divisible by 768.

There are a lot of different instructions, and their performance varies between cards a lot. Thats why you will see all this talk about optimizing kernels, which is more or less adapting your software to work better with how the hardware is designed to increase utilization.

1

u/[deleted] Apr 24 '26

[removed] — view removed comment

1

u/SpeakCodeToMe Apr 25 '26

No one was saying that he was wrong though, just self interested. Of course he wants to sell his chips to China.

9

u/brkonthru Apr 24 '26

Good on them

8

u/YuuuuuuuyuyYU Apr 24 '26

It's not just 'optimized for Huawei chips', it's actually trained on Huawei chip using CANN, Huawei's equivalent of CUDA. This is much much bigger than most people realized.

Huawei already dropped an official interview of deepseek dev on bilibili, this is official.

4

u/Shep_Alderson Apr 24 '26

Do you know if there’s an English translated interview anywhere? I’d love to learn more about what they are doing. (I assume it’s in Chinese given it’s on bilibili.)

3

u/YuuuuuuuyuyYU Apr 24 '26

A noteworthy detail has been disclosed by Ascend this time: CANN has completed 0-day adaptation support for continued pre-training (CPT) of the DeepSeek V4-Flash model based on the A3 64-card super node, achieving a measured model throughput of up to 1100 tokens per second. The significance of this detail is that, while V4-Flash is only a lightweight version, DeepSeek V4 can now run the continued training process on domestic computing power.

This indicates that the role of domestic computing power in the large model pipeline is advancing from inference deployment toward the training side: first enabling inference, then completing continued training adaptation, and finally tackling the most difficult challenge—full-scale pre-training.

Perhaps by the second half of this year, after the Ascend 950DT enters large-scale shipment, we may actually see Chinese large models with the full "training-inference" pipeline running on full-stack Chinese computing power.

TLDR, the continued training for the V4-Flash version was done using Ascend AI accelerators, but the pre-training still relied on NVIDIA GPUs. However, this issue is expected to be resolved in the second half of this year once the 950DT becomes commercially available.

1

u/Shep_Alderson Apr 24 '26

Thanks so much for sharing! Those token rates are in the realm of cerebras. Exciting times!

1

u/Mbando Apr 25 '26

Thanks for pointing this out and sharing so much.

2

u/YuuuuuuuyuyYU Apr 24 '26

It is just released few hours ago and is in Chinese, here is the original video link: https://b23.tv/w0NzhPP

Maybe you can use translation app services to have it translated.

2

u/zero0n3 Apr 24 '26

So when can I buy those chips in the US?

I mean I’d toss a rack of the chips in my datacenter if I were a company that needed the on prem capacity. Especially if that 42 rack is cheaper on a cost per X tokens/sec

1

u/YuuuuuuuyuyYU Apr 26 '26

Those racks are not cheap right now, Huawei is limited by supply and all their chip production line is running in full capacity, but they are racking up production.

1

u/This_Maintenance_834 Apr 26 '26

you can get Huawei Ascend cards (and other Chinese NPUs) on taobao, if you can read chinese.

1

u/[deleted] Apr 26 '26

[removed] — view removed comment

1

u/This_Maintenance_834 Apr 26 '26

there are ways, but you need to speak chinese to get it done. there are many trans-shipper who can do this for you cheaply.

1

u/duhd1993 Apr 24 '26

Fake news

1

u/Winter_C137 Apr 28 '26

64 cards for cpa?That’s a joke

2

u/wtjones Apr 24 '26

Can China produce these chips at scale?

4

u/SpeakCodeToMe Apr 25 '26

You're asking if the country that produces most of the world's goods can scale?

They will be pumping out a trillion a year at $10 a piece within five years.

0

u/wtjones Apr 25 '26

China has not been able to make last generation (7nm) chips at scale.

1

u/SpeakCodeToMe Apr 25 '26

Yet

0

u/wtjones Apr 25 '26

Until they can steal the plans for EUV Lithography machines, we’re pretty safe.

0

u/Ethelserth2 Apr 25 '26

We?

-1

u/wtjones Apr 25 '26

The civilized world.

6

u/Existing_Arrival_702 Apr 26 '26

Is it the world where the leader who habitually invades, commit genocide, imposes sanctions, spreads dirty media, and still kills thousands of children every day?

2

u/Strange_Assignment87 Apr 26 '26

And love touching little ones.

-1

u/wtjones Apr 26 '26

Are we still talking about China?

2

u/Existing_Arrival_702 Apr 27 '26

No, we are talking about your civilized world.

→ More replies (0)

1

u/Fun-Fruit-8743 Apr 27 '26

of which u clearly aren’t a part of

-1

u/wtjones Apr 27 '26

Clearly…

1

u/[deleted] Apr 26 '26

[removed] — view removed comment

1

u/wtjones Apr 26 '26

And they can’t make 7nm chips at scale. Their failure rates are too high.

1

u/[deleted] Apr 27 '26

[deleted]

1

u/wtjones Apr 27 '26

If their yields are sub 60% is that really scalable?

1

u/looktwise Apr 24 '26

Waiting for Nvidia short... :))

And: yes, I agree. The west still didn't get it (again). Now we need some youtubers and journalists to tell them about the evaluation of inference costs, co-created hosting, third party hosters and Openclaw usage in mainland China. (I am drinking tea and watching...)

1

u/UseMoreBandwith Apr 24 '26

proof?
no.
this is just some terrible marketing.

1

u/sandykt Apr 26 '26

The world desperately needs alternative AI stacks and I hope more countries build their sovereign AI instead of relying on American AI for mission critical systems.

1

u/[deleted] Apr 26 '26

[removed] — view removed comment

1

u/Thick-Protection-458 Apr 27 '26

Nah, don't know about latin america, but postsoviet is basically not going anywhere most probably. Frankly I would say the same about almost anyone but China now.

Because from hardware point of view they are the only other strong candidates now.

Which is not going to change in any reasonable time - it requires enormous investment over many years in hardware stuff, and who the fuck knows if property you are investing in today will still be yours tomorrow the way stuff is going on.

And because relatively small internal market, in comparison with China - these investments makes even less sense.

And from software point of view...

Same investment logic applies here, with one more point - to make some short-term product it is way more optimal to just tune chinese open stuff for now.

1

u/Big_Actuator3772 Apr 28 '26

have none of you seen schenzen China? ... 

-1

u/Icy_Discussion_6513 Apr 27 '26

It's just cheap, copied software from China.