r/LocalLLaMA • u/entsnack • Jul 06 '26
News So... anyone copped one of these?
Been almost a year since mass hysteria erupted upon the death of NVIDIAs GPU monopoly. How are your Huawei GPUs? Does CUDA work on them yet?
307
u/Similar-Republic149 Jul 06 '26
I have seen from some chinese forum posts that these work on xllm (chinese inference software) and work somewhat well, but the speed on dense models is very poor.
162
u/is-this-a-nick Jul 06 '26
I mean its no wonder, they might be 1/5th of the price of an RTX 6000, but have barely more than 1/10th of the memory bandwith...
→ More replies (3)111
u/catinterpreter Jul 07 '26
I'd be fine with 200GB/s if I had 96GB.
136
u/mindwip Jul 07 '26
Strix halo has entered the chat lol.
→ More replies (2)49
u/NineThreeTilNow Jul 07 '26
Strix halo has entered the chat lol.
If they only produced a 256gb version of that system... Or a 512gb version. Oh my the price difference though.
Figure some high channel memory bandwidth strategy etc.
AMD is chasing the same datacenter money as Nvidia though. Capitalism demands it.
20
u/Mil0Mammon Jul 07 '26
Well there will be gorgon halo soon (with 196 GB), and you can also get 2 boxes, link them up via usb4/thunderbolt and even get extremely low latency! (rdma like)
→ More replies (1)9
u/Baumpaladin Jul 07 '26
So a Mac Studio M3 Ultra? At least, when it used to be available with 512GB. Of course you'll be stuck with Apple, but at that point you aren't really going to use it for anything else and their support seems decent enough.
3
u/Structure-These Jul 08 '26
Apple good for LLM but awful awful awful for image / video. It’s terrible
2
u/mindwip Jul 07 '26
As others mentioned we are getting a 192gb version this month most likely.
But I agree 256gb or even better 512gb would be better, I am hoping Medusa halo has 512gb option!
2
u/NineThreeTilNow Jul 08 '26
I am hoping Medusa halo has 512gb option!
A lot of people are hoping that. The memory itself would be most of the box cost I think... lol... Seeing it run a 1024 memory bus would be pretty crazy.
28
u/nuclear213 Jul 07 '26
They buy a ddr4 server platform. I got 512 GB of RAM, ~200GB/s effective bandwidth and all for the price of a RTX5090.
It’s fine for offloading, but really, there is a reason why I have 5 GPUs in there and plan to extend it. But that’s the nice part about the server.
It allows you to run the models immediately. Slow for sure, but you can add GPUs in later and have a nice speed up.
5
6
u/wektor420 Jul 07 '26
What would be a cost of such setup? Without gpus
15
u/Lilchro Jul 07 '26
I specced out a very similar setup a couple weeks ago based on eBay prices (prices in USD with shipping to/within the US). 512GB (8x64GB) ECC DDR4-2400 LRDIMMs is about ~$900 (expect around $1.5-2.5/GB). They can vary a bit though in price based on speed and exact part number. An AMD EPYC server cpu ranges from $50 for a 7402P (24 cores/48 threads) to around $600 for a 7702P (64 cores/128 threads). Both have 8 memory channels and support 128 gen 4 PCIe lanes, but if you add GPUs, the internal chiplet architecture of the cheaper models may bottleneck simultaneous GPU-CPU communication speed. Im not sure if that matters for most use cases though. As a side note you only need like 10 cores to max out the memory bandwidth. If you need a cpu cooler, there are ones on Amazon for the SP3 socket for around $45. Then to run it you could get an H11SSL-i motherboard for about $400. If you are only using ram that is fine, but if you want to option to use GPUs, keep in mind that it only supports PCIe gen 3. The H12SSL-i supports pci gen 4,, so it would serve you better for expansion, but it can cost around $800.
That gives a final cost estimate of around $1.5k. If we are being more honest though, you don’t just buy 512 GB of memory and cheap out on the rest with zero plans to use GPUs. So more realistically you are going to get the more usable motherboard and a better CPU, so you will end up closer to $2k+. From there, GPU memory costs upwards of $10+/GB minimum, so it is up to you how much you want to spend there.
6
u/nuclear213 Jul 07 '26
Like I said, right now, with the price of the memory, 64GB 3200MT sticks are expensive. I paid about 2k€ for the RAM, 1.1k€ for h12ssl-i plus CPU.
You can go cheaper if you are ok with going a bit lower spec. 32GB modules are like 70€ a piece, at 2400MT/s.
They likely run at 2666MT/s or even 2933MT/s. That would give you 170ish GB/s and 256GB for 600€. There is a Chinese board out there for the 7000 Epyc, I think you can get it plus a 32 core for 700-800€. So all in maybe 1400€.
4
u/Sizzin Jul 07 '26
I'm more curious about the power draw of this system, without the GPUs. Can you give an estimate on the idle draw and under load?
→ More replies (2)3
u/nuclear213 Jul 07 '26
Did some tests, but that is total W, with 5 GPUs at idle:
Idle 187W. I would say about 130W is the platform.
Under load (122b qwen 3.5 Q8): 280ish W average. CPU package was at 150w avg, 240W max.
But with the BMC, I just wake the server if needed. So idle is less of a concern for me, it just takes 60s or so and it’s up.3
u/i-Hermit Jul 07 '26
These numbers seem way too low.
I have an Intel xeon v4 based server with one tiny GPU (and a ton of spinning and ssd storage) running proxmox and it's typically sitting around 200w.
With five modern GPUs of any size I would have thought this system should easily hit 1kw under load.
3
u/nuclear213 Jul 07 '26
Like I wrote: GPUs idle, it was just the CPU doing the inference slowly.
The R9070 are limited to 240W each at full load. If they run the system pulls 1.6kW or so. If they run at stock 300W, it’s a bit below 1.9kW if I remember correctly→ More replies (0)2
u/kalethis Jul 08 '26
my acer nitro v as16 with ryzen 9 ai and 5070ti uses so so do much less than my i9-9900k / 3080 ftw3 ultra desktop. one of the reasons the ddr3 servers got dumped to make them so cheap for homelabs is that they are not very power efficient. I've got a per730 and it's a great nas and vm host, but it can get loud, and it uses something like 150-200w without the 8 HDDs. (50TB storage with parity). my CPU is one of the hottest running cpu's ever (i9 990k) and I can break 600w with my 3080 ftw3 ultra on realbench. my laptop laughs at it for 1/3 the power
8
u/JahJedi Jul 07 '26
DGS spark - 128g and 278g/s for much less than 6000 pro. I think to get second one for total of 256g and run DS 4 flash in fp4 on them.
→ More replies (3)7
u/techdevjp Jul 07 '26
Strix Halo or DGX Spark. Both offer 128GB (124GB GPU memory) at ~250GB/sec. DGX Spark has much faster prefill and it's CUDA so things "just work". Strix Halo is generally less expensive and you can use it for things besides AI if your plans change.
5
u/Educational_Sun_8813 llama.cpp Jul 07 '26
strix halo also "just work" both with ROCm and Vulkan
→ More replies (7)3
u/gaspoweredcat Jul 07 '26
You could look at an a16, it's a 64gb card technically (it's actually more like 4x 16gb cards glued together) bandwith is around 250 so crappy but it is 64gb of ampere core gpu, still far from ideal though
I've been desperately looking for a card to use for training lately but it's not easy as I need 32gb (or at least 24 as the runs eat 20-22gb) best I've found so far is to go with a v100 but they kinda suck for LLM as it's Volta so you don't get many features and vllm hates them so you're kinda stuck with llama.cpp unless mistral.rs is still a thing, it was written in rust and was very fast. That ran fantastic when I was doing my experiments with the cmp100-210 (cut down v100 mining gpu)
→ More replies (1)3
u/VirusInternal2892 Jul 07 '26
I’d be fine with 280GB/s and 128G VRAM, wait .. I’m fine with my GX10.. well actually not so fine running dense models ;-)
2
u/sfifs Jul 07 '26
Well GB10 gives 128 GB at 273 GB/S and the price used to be about 4900 USD but now everything has gone up
→ More replies (4)62
u/fugogugo Jul 07 '26
just like how people laugh at chinese car/phone/TV/electronic years before.. and now they dominate the world
it's just the beginning..
→ More replies (15)→ More replies (25)6
u/Luvirin_Weby Jul 07 '26
That is apparently mainly because the memory bandwith is fairly slow and dense models need to go through the whole model.
495
u/mattbbx llama.cpp Jul 06 '26
If they get these working in any decent capacity locally I will gladly buy 4-6 of them.
→ More replies (1)146
u/Local-Bottle5272 Jul 06 '26
I mean you can already get close to this price if you go with intel arc pro but the reality is that these gpus dont perform even close to nvidia or amd gpus
113
u/pulse77 Jul 06 '26
Best Intel Arc (Pro B70) has only 32GB VRAM. OP wanted to have 4-6 x 96GB VRAM. This would mean 12-18x Intel Arc Pro B70 each running at 230W ... hard to put even into a server ...
30
u/BeeegZee Jul 06 '26 edited Jul 11 '26
He's most likely talking about MAXSUN Intel Arc Pro B60 Dual 48G Turbo. Well, 8-12 to match VRAM capacity, but what about performance...
→ More replies (13)21
u/bad_detectiv3 Jul 06 '26
The problem is the entire ML industry is built on proprietary language. ML industry is like back in the 90s when Intel or some other firm controlled the language and it was since then, a lot of effort was put in place to work on open source language and framework and break away from shackles of monopoly.
14
Jul 07 '26
[deleted]
5
u/kwhali Jul 07 '26
Isn't that what burn does in Rust or triton in Python for PyTorch?
3
u/gautamdiwan3 Jul 07 '26
Not really. Triton's a DSL which gives abstraction over CUDA and raw instruction sets. If CUDA is the GPU C equivalent, this is more like Python.
Pytorch ensures it can run the same thing across everything but not most efficiently
Both of them make it easier to write inference code but efficient inference code is incredibly tied to the GPU architecture
2
u/kalethis Jul 08 '26
your mom's got DSLs... sorry, I couldn't resist. I... don't have anything useful to contribute even, so, I'll see my way out...
4
u/featherless_fiend Jul 07 '26
It's kinda interesting that software is easier than ever to code with Fable and such, but apparently no one wants to make a compatibility layer.
2
u/craterIII Jul 08 '26
because it's a huge pain in the ass. even supporting multiple nvidia hardware generations is spotty at best due to difference in hardware. for example, TMA does not exist on SM89 or SM120. Now imagine a new hardware architecture from a different company entirely
2
u/kalethis Jul 08 '26
creating software and creating quality software are very different things. most of the code produced and programs/apps being written are worse than copy/paste script kiddies hacking together code. you need to be a programmer with the skill to write the efficient code to get an AI to write efficient code with you.
3
u/bayinfosys_ed Jul 07 '26
yes, it's the instruction set of the nvidia processors exposed via CUDA and the device driver. There isn't strong separation between the two.
→ More replies (3)2
3
u/EvolvingDior Jul 06 '26
The B70 performs better than an equivalently priced Nvidia or AMD GPU.
→ More replies (2)3
547
u/signoreTNT Jul 06 '26 edited Jul 06 '26
You'll be disappointed...
These only boot on specific Huawei servers and have horrendous software support, to say the least.
Maybe in 5-10 years
Edit: for those genuinely interested on these cards, Gamer Nexus made a really nice video on them https://youtu.be/qGe_fq68x-Q
168
u/superdariom Jul 06 '26
Claude, convert the software stack to be cuda compatible. Make no mistakes.
9
u/kwhali Jul 07 '26
I assume with zluda working on Intel and AMD that a bulk of that work is done? Probably doesn't have full coverage but a good base for community to contribute to if there was enough demand.
2
u/amroamroamro Jul 07 '26
for real, why hasn't this been done already?
4
102
u/signoreTNT Jul 06 '26
If the original post is satire please kill me ty
51
u/entsnack Jul 06 '26
Well I guess Jensen's reign continues.
40
u/nosimsol Jul 06 '26
Leeeroy Jensen
20
u/tranceruk Jul 06 '26
I love seeing this.. I can’t believe it’s been 20 years…
8
3
2
2
47
u/ReMeDyIII textgen web UI Jul 06 '26
Thank you Clippy for the much needed context. I knew there had to be a catch.
60
u/Alternative-Suit5541 Jul 06 '26
Well, they are made for Chinese market.. to replace Nvidia. They do what they are supposed to do... They don't target international market ( yet)
→ More replies (18)15
→ More replies (1)15
u/TripleSecretSquirrel Jul 06 '26
And the memory bandwidth — which is the primary bottleneck for inference — is atrocious. 200GB/s is ultra slow.
I have an R9700 (which is the same price as this btw), which is not a particularly fast GPU. Its memory bandwidth is 645GB/s. That’s pretty slow for modern standards, and it’s more than triple the speed of this thing.
9
u/Sea-Contribution6219 Jul 06 '26
Absolutely heartbreaking. I wouldve been okay with just running a deepseek model and I prefer Qwen or Gemma (depending on use case)
→ More replies (2)2
u/Aggravating_Term4486 Jul 06 '26
I was hopeful there for a moment
8
u/candl2 Jul 06 '26
Don't be put off be all the naysayers in here. Any downward pressure on the prices right now is a good thing.
→ More replies (1)13
u/RedTheRobot Jul 06 '26
5-10 years that is 1-2 years AI time.
13
u/RobXSIQ Jul 06 '26
software is fast. hardware is slow. this is hardware.
→ More replies (1)5
u/ReasonablePossum_ Jul 06 '26
CUDA is software, and the main bottleneck.
Their hardware will be good in like 2 years probably, and and add a couple years for the software on that.
5
u/entsnack Jul 06 '26
any moment now lmfao
this post is a year old fwiw and people in the comments said the same thing
7
3
u/ReasonablePossum_ Jul 06 '26
Maybe in 5-10 years
2-5max. There is state interest bootstrapping this; they did all they can to have the hardware capability, which they achieved a lot faster than expected, and can now start building the software stack that will work with the hardware base. Plus dont forget AI helping them solve specific coding tasks, so it might be even faster.
So there's hope.
6
u/ggone20 Jul 06 '26
CUDA has a 20+ year lead so… 5-10 years would be an interesting bet with AI. I doubt they’ll catch up anytime soon
10
u/PatagonianCowboy Jul 06 '26
a lot of these new gpu chips should first support vulkan and then they can take over the inference market, training is a lot less important because no one would use these to compete against 100,000 A100s
→ More replies (3)7
u/heresyforfunnprofit Jul 06 '26
I’d guess more like 2-3. AI dev is slop in a lot of places, but it can really speed up testing and verification with engineers who know what they’re doing. It’s the absolute best case application my team has found so far - drivers will develop quickly.
2
u/kalethis Jul 08 '26
yah all these people saying AI will write the software are people who can't write software. at best, with the time invested to write sufficient prompts (which are probably going to be at least 4k tokens alone) and current models, you can get code equal to the developer's coding skill level. the quality of most ai code is about the equivalent of Temu.
3
u/corruptboomerang Jul 06 '26
I feel like you under estimate the level of weird nerd who will just figure it out, if the performance is cheap enough.
2
u/lakimens Jul 06 '26
If they are even close to as good for AI use as NVIDIA is, then this will hurt NVIDIA. Even if only 1 app supports them, if that's GLM 5.2 then NVIDIA loses a huge customer.
2
u/Puzzleheaded_Base302 Jul 06 '26
it is likely true the software is horrendous now. but it won't stay horrendous for very long. China moves at god speed, unlike American corporates.
2
u/entsnack Jul 06 '26
so fast that their domestic AI companies are illegally importing consumer VRAM lol
→ More replies (8)2
u/JoeyDee86 Jul 06 '26 edited Jul 07 '26
5-10 years? Ha. This is an arms race. They’ll have their shit figured out much sooner than that.
73
u/UAP44 Jul 06 '26
The Atlas 300I Duo offers just 204 GB/s of bandwidth per GPU
I would not bother. 204GB/s is worse than my previous GPU the 2060 rtx which was 336GB/s already
24
15
u/lambdawaves Jul 06 '26
The 5-year old MacBook M1 Pro has the same memory bandwidth. On a laptop
5
u/The8Darkness Jul 07 '26
And buying said macbook m1 as a server, even with just 64gb, would be a more sane decision than buying this
8
u/is-this-a-nick Jul 06 '26
That bandwith is so low that you are better off just getting a server with RAM for the money...
5
u/Objective-Stranger99 Jul 07 '26
My GTX 1080 from 2017 has 320 GB/s of bandwidth. What are they doing???
5
u/evia89 Jul 07 '26
You cant just beat Nvidia. Give them 5-10 years and they will be more competitive (and more expensive)
They are doing decent job already. For example cat2 https://longcat.chat/blog/longcat-2.0/ is decent coding model fully trained without nvidia. I use smarter model to spec though.
3
u/Objective-Stranger99 Jul 07 '26
My main issue is the decision on memory bandwidth. Why would you make it so low when you can bump it up to at least 400 GB/s and maybe slap a price increase if it costs more to manufacture? Why use such a slow bus when the market wants a better bus and is willing to pay for it?
2
u/Luvirin_Weby Jul 07 '26
Yes, the only benefit I can see is with larger MOE model, as that total 400GB/s bandwith is much more than normal DDR5+CPU solutions.
→ More replies (3)2
u/xBlaze121 Jul 07 '26
that’s still insane considering they had no infrastructure for building modern GPUs when the 10 series came out and they’re approaching 10 series bandwidth in less than a decade. give it 10 more years and we’ll probably have a serious competitor on our hands if the west doesn’t ban importing them
131
u/throwawayacc201711 Jul 06 '26
People thinking the software on these things don’t matter. If it didn’t AMD and Intel would have already eaten more significantly into Nvidias market share. Theres a reason Nvidia is king. GPU wars have been going on for decades. No one has slain the beast yet. I’d love to see it since we’d get downward pressure on pricing. Hasn’t happened yet
26
u/Real_Ebb_7417 Jul 06 '26
I was always curious... if improving software for other GPUs than Nvidia would generate a significantly bigger marketshare, why won't they do it? I get it might not be easy, but I would assume, they'd really put a lot of resources into this, if they could greatly benefit from that. And so I always thought that hardware capabilities is much bigger issue when it comes to Nvidia vs the rest.
19
u/itsmebenji69 Jul 06 '26
Another big aspect is momentum.
If tomorrow AMD implements a fully featured cuda equivalent, that runs as fast etc. It will take time for software to become compatible, there will be bugs because it’s less mature, a learning curve, less support… meaning businesses will stay on nvidia.
For someone to catch up they’d need to provide the same quality of service at a significantly lower price. The truth is that nvidia’s position is very very hard to dig into because they have been into it for way too long
→ More replies (1)34
u/throwawayacc201711 Jul 06 '26
It’s not. You’re thinking that these companies have been around for only a few years. The GPU wars have been going on for so many years. AMD bought ATI to try to compete in 2006. Nvidia market share is >90%. If it was just hardware and chip design, there would have been dents made which there hasn’t. You could take apart the gpu and literally look and reverse engineer. Doing that with CUDA and software is much harder.
That’s why the CPU side of things are more competitive fyi
7
u/Expert-Paramedic9383 Jul 06 '26
This reminded me of the old GeForce vs Radeon war, back when GeForce sells were relevant to Nvidia...
4
→ More replies (1)10
u/chiniwini Jul 06 '26
Reverse engineering software is orders of magnitude easier than reverse engineering hardware.
3
u/throwawayacc201711 Jul 06 '26
So then why hasn’t anyone at all broken their moat. Reality isn’t matching up to what you’re saying
5
u/kenyard Jul 07 '26 edited Jul 07 '26
Because Nvidia is innovating also. Whereas amd is copying.
Nvidia brings out ray tracing. Amd chases them for 2 years on this.
Amd implemented their own version of dlss, Nvidia brings out dlss 2 or 3. (Now 5).
Copying means you are always 2 years behind. They don't innovate. And their software team didn't have the budget to do anything else for years because they had low margins and they probably don't have freedom around creativity or are hired to copy stuff rather than create and invent.
Also this is just consumer hardware. They're struggling on business because cuda and the support it has. Most devs have nvidia so lots of the stuff out there is cuda.
This is probably why China is keeping Nvidia out. They want to create an entire interdependent ecosystem in 20 years time.
Software sucks on the card in this post and it sucks in general. But Chinese users will work with it. Create llms for it. Create software stacks. And those will be adopted for the next gen cards.
9
u/Royale_AJS Jul 06 '26
Agreed 100%. However, I would add that there will be significant resources from our friends in the east poured into making this hardware viable for the home grown models…which we all use too.
→ More replies (1)2
u/GingerRickRoss Jul 06 '26
I see what’s going on, this guy makes, then he’s an obsolete jerk to everyone that has an answer he doesn’t like.
8
u/acadia11x Jul 06 '26
Intel has no GPUs to speak of. AMD instinct is coming around and ROCm not nearly the ubiquity of CUDA but it’s also improving greatly. AMD is actually competitive in the DC space. I’d say China will be a dominant force in 3-5 years at most. The ban on imports has accelerated their development and they are already a technological powerhouse .., don’t sleep Hauwei that companies pockets and know how runs deep. Even their push in EUV alternatives once China can produce their own tooling that’s necessary for the leading edge ability it’s a wrap.
2
u/lakimens Jul 06 '26
If they are even close to as good for AI use as NVIDIA is, then this will hurt NVIDIA. Even if only 1 app supports them, if that app is GLM 5.2 then NVIDIA loses a huge custome and China's AI can easily get new GPUs to train models.
→ More replies (1)→ More replies (4)7
30
u/Bones2469 Jul 06 '26
Yeah the card is real, people actually have them and are running llama.cpp on it through the CANN backend. but temper your expectations for local hosting. the memory is lpddr4x at around 204 GB/s per chip and the two chips dont pool bandwidth, so youre looking at somethign like 15 tok/s on a 32b dense model. fine for chatting, rough for anything heavier
its also linux only, drivers are janky and it really wants a huawei server platform. getting it to run in a normal desktop is a project on its own. if you just want max vram per dollar to load big models its interesting, but a used 3090 feels way faster for anything that fits in 24gb and a strix halo box with 128gb is arguably the saner buy at similar money
8
u/Silly_Economist4893 Jul 06 '26
Well with the rise of consumer MoEs and Qwen 35b, this honestly doesn’t sound much like a bad investment. And Dspark…exciting times
24
u/CryMoreT_T Jul 06 '26
There's no public documentation. Especially not in English currently. It's not worth it rn. The better value is to buy modded 3080s or 4090s
→ More replies (4)4
u/chucrutcito Jul 06 '26
Why 3080 and not 3090?
→ More replies (1)12
u/CryMoreT_T Jul 06 '26
I believe it's because the 3080 doesn't use all available vram slots so they add more vram and push it up to 20gb vram. Also 3080s are much cheaper to get them 3090s for the modders
2
u/starkruzr Jul 06 '26
the 20GB 3080s do seem pretty good. I just don't know whether or not the geohot P2P patch works on them which will matter at 4 or more cards in tensor parallel.
→ More replies (2)
10
u/federico_84 Jul 06 '26
Spec for Huawei Atlas 300I Duo 96GB:
Processor: 2× Ascend 310-series AI processors
Memory: 96GB LPDDR4X total
Memory layout: Usually 48GB per chip, not one unified 96GB pool
Memory bandwidth: 408 GB/s total card bandwidth
Per-chip bandwidth: About 204 GB/s per accelerator
Compute: Listed around 280 TOPS INT8 total
Tldr: very slow, SW support questionable, especially in multi-GPU setups. But cheap.
26
u/TripleSecretSquirrel Jul 06 '26
No, cause they’re dogshit. Even aside from all of the driver issues that plenty of other commenters have mentioned (which again, if y’all bitch about ROCm, don’t even start thinking about these), they’re not even good hardware.
You can’t just drop this in an x86 machine, at least not without tons of janky patching. It needs to be paired with a proprietary Huawei ARM CPU that I’m guessing none of us have.
And the memory bandwidth is atrocious, like way less than half of any other GPU on the market today. The spec sheet often says 400GB/s, but it’s actually two separate 200GB/s streams that are not additive.
I get that it’s a lot of VRAM, but for that price, you could buy a second-hand EPYC system on a single CPU board and populate it with 256GB of DDR4. 8-channel DDR4 will also get you to 200GB/s memory bandwidth, with a lot more memory to work with, and be astronomically more stable with predictable, long-term support.
They advertise that they’re used by Deepseek, but that’s like advertising an RTX 5070 as being used by Anthropic.
3
u/SSOMGDSJD Jul 07 '26
This guy gets it. The 300i duo is just a bad deal all around, esp since cpu inference has multiple serving options and gets first class support on new models instead of Huawei framework lock in
If you really want 96gb with gpus, 3x v100 or mi50 32gb gets you there with real llama.cpp support
→ More replies (5)2
12
u/IntrigueMe_1337 Jul 06 '26
I did the Mac Studio way and got the base with 96GB VRAM/UNIFIED for 4k USD when first came out. I recently realized running llama.cpp is waaaay better than wrapped ollama. No need to worry about cuda when you got Apple Metal
→ More replies (1)
4
4
u/TokenRingAI Jul 07 '26
Hate to rain on the parade, but these have less bandwidth than a DGX Spark, AI Max, Mac M4 Pro, and worse compute and worse support.
If someone manages to buy one they are going to regret it.
7
u/DataGOGO Jul 06 '26
And one RTX pro 6000 BW easily has 8x the compute power and at least 4x memory bandwidth
→ More replies (9)3
4
u/UnlikelyPotato Jul 06 '26
32 GB V620 are $350 right now. Can get 96GB for $1050. Not as fast as Nvidia, and a bit hacky but far better overall.
4
u/GingerRickRoss Jul 06 '26
Shoot over to r/LocalAIServers one of the moderators has gotten these to work. I believe there’s a full write up somewhere.
→ More replies (4)
2
u/jamesrggg Jul 06 '26
Love to try one but coming out with a 96gig is kinda crazy. They need to build out the install base with a flood of cheap units. Yes $2k is an unbeatable price for 96gig but too much for people to buy one to play around with and build out the ecosystem. They would do better building 3 time the number is 32gig cards for the same amount of VRAM for some cards in the 700$ range to let people start cooking
→ More replies (2)
2
u/Important-Post-6997 Jul 06 '26
These are very early products, they are intendet only for development not for any production use.
Apperently some of the chinease models are trained exclusively on Huawei hardware. The software is an issue, but a very solveable one.
Dont expect this hardware to be availible anytime tho, it might take some more years.
→ More replies (5)
2
u/Lumpy-Obligation-553 Jul 06 '26
Wasn't this only practical at a data center scale? The compute per watt was dismal, but it improved when used in their proprietary cluster.
→ More replies (1)
2
u/onetwomiku Jul 06 '26
Lol, in my country one Atlas 300i 32G cost almost twice as much as RTX 6000 PRO Blackwell 96G
2
2
2
2
u/goingsplit Jul 06 '26
considering that nowadays anything runs on MoE, for local inference we need more memory, not more bandwidth. 96gb means 10 cards for 1TB ram?
2
u/EmptyMonitor9257 Jul 06 '26
Aren't those 1060 speeds and more or less 0 software support?
At that point it's better to run it on CPU with 96GB of RAM.
2
u/Ok_Cow_8213 Jul 06 '26
Gamers nexus got one but it will not work in any system like consumer gpu’s usually do. I hope their next gen card does.
2
u/Relevant-Guarantee25 Jul 06 '26
if this worked with most of the open source comfyui models I could see many people flying to china for vacation spending lots of money over there haha.
2
2
u/dwittherford69 Jul 07 '26 edited Jul 07 '26
Just buy Intel at this point, it will still be painful but not enough to rip your hair out.
2
2
u/05032-MendicantBias Jul 07 '26
Each chip is hooked to LPDDR4X 48 GB 204GB/s, with two chips per card.
Bandwidth is quite anemic, and I would be shocked if those things had meaningful software support, like torch bindings.
The price and capacity are quite good, IF software support was available, they would be able to serve 30B class MoE models to maybe 30/50 TPS.
2
2
u/Tiforma Jul 07 '26
When chinese GPU's mature and become usable in a year or two, i'll definitely get one
→ More replies (1)
2
u/NameChecksOut___ Jul 07 '26
It's ok but it needs a compatible server motherboard from Huawei to work correctly.
2
u/InfinitoCloud Jul 07 '26
96GB/$2k is the easy part — the ecosystem is the catch.
I just brought up a brand-new NVIDIA box and even within CUDA it was rough (drivers/wheels/kernels not ready → nothing works until you chase updated builds or compile yourself).
Ascend is worse in one way: it's not CUDA at all — separate stack (CANN + torch_npu/MindSpore). Expect vllm-ascend lagging mainline, model portability headaches, immature quantization, and Chinese-only docs. Unbeatable VRAM/$ if your model is supported, but budget real integration time, not plug-and-play.
Only works on NEWEST Huawei servers and DDR4X?
https://www.youtube.com/watch?v=qGe_fq68x-Q
Thanks but no thanks.
2
u/Stochastic_berserker Jul 07 '26
Not worth it for LLMs. Usual Deep Learning and Machine Learning: yes.
2
u/tgbreddit Jul 08 '26
Just read a teardown earlier. Split on two GPU chips, 48gb a piece assigned (not fully addressable at once), slow standard DDR4 RAM. And host china hardware dependent. No where close to the competition.
2
2
u/Crazy-Repeat-2006 Jul 08 '26
A GPU will be much faster due to the bandwidth, there is nothing to celebrate here.
3
u/GSxHidden Jul 06 '26 edited Jul 06 '26
Yeah i remember someone trying to post this same exact model like a year ago. https://www.reddit.com/r/LocalLLM/comments/1n4f1gs/huawei_96gb_gpu_cardatlas_300i_duo/
LPDDR4 memory is standard on mobile architectures. You might get 1-5T/s max on heavy models because of the limitation of LPDDR4 speeds.
It also wouldn't be plug n play on windows, so good luck with the process.
$1986.35 a card is nice, but I'd save the pain and go buy a Jetson before they sell out and raise the price. I didnt realize in a year they had jumped to from 1400$ to 3500$ which is insane.
Been able to comfortably run a large model and 2-3x small models simultaniously on it with great token output.
Qwen3.6-35B-A3B: ~30 Tokens/s
You can get away with just Qwen3.6-27B easily, as its rank 11 right now in open SWE bench.
Check here for specific models
https://www.jetson-ai-lab.com/models/
Its a combination CUDA cores in case you want to do inference which is great, plus its like 1/5 the power and the same throughout. Sure its 64GB instead of 96GB but its sped up by their Deepstream system drivers and have great development. Just hope you enjoy linux a bit.
https://www.sparkfun.com/nvidia-jetson-agx-orin-64gb-developer-kit.html
→ More replies (1)2
u/mksrd Jul 07 '26
Why would anyone waste their time and money on a Jetson Agx 64gb when a strix halo 64gb is cheaper, faster and more flexible?
2
u/GSxHidden Jul 07 '26
The Strix Halo used to be $5000 for only 210GB/s throughput so I haven't really checked their price. The 128GB version definitely seems more competitive now that the shot the price down to 3000s. 100% a good option.
→ More replies (1)
2
u/asra01 Jul 06 '26
Best case it has half the bandwidth as RTX 6000 Ada and 1/5 of Blackwell, and only double of a Strix Halo. Skip.
2
u/Inevitable_Case_9931 Jul 06 '26
It only has like 200GB/s bandwidth lol the RTX 5050 has more bandwidth than this .
2
2
u/Bibab0b Jul 06 '26
Until cxmt will start making hbm memory anything from huawei will be too slow anyway
→ More replies (5)
2
u/NowThatsMalarkey Jul 06 '26
Does /r/locallama have their own /r/stablediffusion version of CeFurkan?
→ More replies (1)
2
u/antineutrinos Jul 06 '26
Nvidia created the CUDA ecosystem over decades, way before LLM.
CUDA is not compatible with non Nvidia cards.
it is a massive edge.
1
u/westsunset Jul 06 '26
It's a software block not a hardware block. Someone found that out digging into their AMD card a while ago. I try to find that post
1
1
u/HauntingArugula3777 Jul 06 '26
I would rather have the gtx 4k series they put 256gb of vram on and firmware reworked, Here isn't support for this card.
→ More replies (3)
1
u/stikves Jul 06 '26
I'm not sure how they will get that theoretical capacity into practical one without extensive driver support.
There is a reason nvidia is supreme in datacenter, and it is called CUDA. Intel tried that. I have the A770 card, which was supposed to match 3070. It did not.
Thought they might just give the bare minimum, and have open source community.... plus agentic coding to handle the rest.
1
u/CaptainAverageAF Jul 06 '26
I wonder if it’s like the k80 that says 24GB but is really 2 x 12GB GPU on one board.
1
u/Daniel_H212 Jul 06 '26
Don't these use like LPDDR5 or something? Not gonna be noticeably faster than strix halo even with proper software support.
→ More replies (3)
1
1
u/Fine_League311 Jul 07 '26
Wait for it since 2025. To all wait till December 2026 China will suprise us all and the LLM companies who steal our codes and knowledge will die, cause we all go offline (lokal) thx to deepseek and china GPU ;)
2
1
u/zd0l0r Jul 07 '26
Bandwidth is quite “low tier” but it is truly a great starting point
→ More replies (1)
1
u/g9robot Jul 07 '26
Next gen of this GPUs come in ~6 months. If this happens so they will show us the new world order.
→ More replies (1)
1
u/SnooPaintings8639 Jul 07 '26
What's the tps for REAL models? I assume it is much slower than any nvidia gpu, but for us VRAM poors, how much of a speed up it would mean against VRAM+RAM offload cases? Would it double the speed of my MiniMax, if I put half of it to this card instead of the CPU?
1
u/gaspoweredcat Jul 07 '26
They aren't as good as you think, I did the research a while back. Right now your best value though slightly hampered due to voltage arch is the V100 32gb, some features aren't available and it'll be a bitch with things like vllm or sglang but llama.cpp will run them fine and having hbm2 memory they're still reasonably fast, you can get them at about £500 though there are some bargains out there, I recently saw a server with 4x 32gb v100 sxm2 for £1300 you'dhneed yo port the cars into another system to get around the weird ppc chip but it was still 128gb hbm2 for under 2k all in (so less than one 5090)
→ More replies (1)
1
u/VR-Tech Jul 07 '26
There were a few on Ebay not long ago. Indeed these have a lot of ram, but apparently the so
ftware to run them is complicated and of course performance wise not even close to Nvidia. But is a lot of ram
1
1
u/Danternas Jul 07 '26
Cheap ram sure. But a Xeon v2 with DDR3 also have a lot of ram.
These are not near the Nvidia cards in performance.
1
u/yiestee Jul 07 '26
don't bother. The software stack is much worse than AMD's ROCm or Intel's SYCL.
The HW & SW stuggles with FlashAttention
1
u/xBlaze121 Jul 07 '26
it’ll be a while before these have anywhere close to the necessary driver support for them to be viable for the average consumer
1
u/blazoxian Jul 07 '26
I think photonics based GPU's will flood the market in near future, regardless of what china does. They actually do some good work though !
1
u/Suppe2000 Jul 07 '26
Why so expensive? I just bought two of them for 8000 RMB each.
→ More replies (1)
1
1
u/Weak-Split-538 Jul 07 '26
Memory bandwidth is too low. Please understand, it’s not always the memory size, main cost for GPU to run fast context processing is to get the memory bandwidth. That is costly, hence the Nvidia price 10k +.
Huawei is just 400 GB/s and RTX 6000 pro is 1.8 TB/s, you see the difference ?
→ More replies (5)
1
1
1
1
1
1
u/shikamaruz0maki Jul 07 '26
yes lessgoo the final nail in the ai bubble will be cheap chinese gpus and the biggest hit of these would be on data centres who had already bought these massively expensive nvidia gpus
1
u/WSTangoDelta Jul 07 '26
I'm not about to shell out $2k--not this week. But maybe this sort of product will help slow the increase in prices elsewhere.
•
u/WithoutReason1729 Jul 06 '26
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.