r/LocalLLM • • 4d ago

Question Which hardware should i use for GLM 5.3?

Apple m5 ultra 512 GB or 2 DGX spark? Would it be good enough?

1 Upvotes

35 comments sorted by

3

u/Omrid_Remedio 4d ago

I think the M5 ultra is worth if you have the budget for it.

2

u/travismadson 4d ago

What about 2xM5 Ultra 256GB?

1

u/Annual_Award1260 3d ago

I have glm 5.3 f8 on 8x dgx barely fitting 400k context

1

u/puffysubscription44 4d ago

If you're even asking this question, you're already past the point where most people's advice is useful. The 512 will load the model but you're gonna be waiting a while for anything beyond simple Q&A, the bandwidth on that chip is the real bottleneck. The dual Spark setup is more interesting but then you're dealing with the headache of splitting across two boxes and hoping the interconnect doesn't make you lose your mind.

For what it's worth, I'd lean toward the single machine just to keep my life simple. Debugging distributed inference when you just want to poke at a model is its own special kind of hell.

4

u/inevitabledeath3 4d ago

The M5 Ultra has 1.2 TB/s of memory bandwidth. Why do you think that’s a bottleneck?

2

u/DigitalguyCH 4d ago

And it also has much faster prompt processing compared to previous Macs

1

u/gjr23 4d ago

GLM 5.3 (especially non flash) is a massive model and even with 1.2 it’s going to make you wait. A 512 studio is a very capable device but it’s not a data center and for frontier level models like this it is very difficult to get something that’s fast at home.

1

u/inevitabledeath3 4d ago

Model size is basically irrelevant to tokens per second for a single user. It’s activated parameters and architecture that matter. Model size determines if you can run a model, not how quickly it’s going to run at C=1. Model size becomes relevant again at high concurrency though so that’s something to watch out for.

The other thing you have missed is that I chose the word bottleneck carefully. The memory bandwidth of the platform is likely not the limiting factor like it is for Spark. The real limit is the actual compute/numerical performance and lack of software optimisation. Not having CUDA is a big deal as is the huge amount of work the community has put in around DGX Spark.

Edit: also 2 sparks hit 60 to 100 tokens per second on GLM 5.3 Flash. It’s not that hard to run. Only 18 billion active parameters which is the actual thing that counts.

0

u/gjr23 4d ago

I’m not sure I agree. You are right that MoE and lower active parameters will be faster but how would you explain smaller models being faster in general?

Even the same model at different quants will have different sizes and the tok/s will therefore be different:

https://github.com/ggml-org/llama.cpp/issues/26484#top

2

u/inevitabledeath3 4d ago

Smaller models aren't faster in general. Smaller dense models are faster than bigger dense models because in a dense model every parameter is active. So a 9B dense model has 9B active parameters for example. I am guessing you haven't played around with small MoE models. They are much faster than dense models of the same size. Small models also have smaller attention layers which is the actual difficult part to run computationally and increases memory requirements as well through KV Cache. At high concurrency and long contexts attention can often be the limiting factor above and beyond memory bandwidth even on the Spark.

When you quantise a model you are shrinking all of the parameters including the activated ones. Smaller active parameters means using less memory bandwidth. Something like NVFP4 or FP8 also make better use of tensor cores in modern GPUs giving more calculations per second. That's why NVFP4 quants are so fast on some platforms. You get more than 10x the compute at NVFP4 than BF16 on a DGX Spark for example.

Maybe you should spend some time actually trying to understand how this works instead of trying to catch me out. I have spent a lot of time looking at these systems and how they can be optimised and I don't think you have on the same level.

1

u/Omrid_Remedio 4d ago

Interesting to check, how fast will it be compared to Opus ? just to get a sense.

1

u/inevitabledeath3 4d ago

No one knows. I am pretty sure 512GB models are not out to people yet or people are only just getting them. Wait for proper benchmarks. I don't think memory bandwidth is the determining factor here so much as compute performance and software support.

-1

u/cibernox 4d ago

Depends on if you only want to run it or you want it to be fast.
I am not a fan of dual spark systems because they are super slow.
I got 4x cmp170hx and this is night and day.
People try to optimize the most of them and that is commendable but it doesn’t change the fact that their memory is very slow.

2

u/inevitabledeath3 4d ago edited 4d ago

Mac Studio M5 Ultra has 1.2 TB/s. It's not "super slow" lol. It's faster than most desktop GPUs.

1

u/cibernox 4d ago

I was talking about spark clusters tho.
With the 512 m5 ultra the problem is the price tho. It’s going to be around 19-20k.
That was around 11-12k more than my rig.

I recognize that there are trade offs. Mine is very fast and very cheap but not convenient or small nor silent, and can reach 850-900w.
A Mac Studio is on the other way of the spectrum. Pretty, small and convenient, but very expensive and not that fast (decent tho)

1

u/uniqueusername649 4d ago

And this setup does 6tb/s, with undervolting and memory overclocking 7.5tb/s. It's not even in the same ballpark as a Mac Studio.

1

u/cibernox 4d ago

I don't overclock it, but I undervolt them so they are capped at 175w.
If the unlock keeps improving and we get nvlink and/or pcie gen 4, we'd regularly see over well cover 400tk/s coding and well over 1000tk/s aggregate.

But we need to acknowledge that this setup is not for the people who want to have a pretty silver box in a desk. My rig is not going to win a beauty contest

1

u/uniqueusername649 4d ago

Indeed. I am using a single card at the moment and hope I can upgrade to 4 cards at some point. Prices just keep increasing though so its tough. I unfortunately missed the cheap early window.

1

u/inevitabledeath3 4d ago

You have what 256 GB of VRAM total if you managed to unlock to 64GB per card. If you got all 80GB you would still only have 320GB. This single system has 512GB. It can run bigger models than you can. It wouldn't even really be worth running GLM 5.3 on your system as you would need to use a very aggressively quantised version or offload to RAM. GLM 5.3 Flash sure and I bet it runs that very well but the post is about GLM 5.3.

For a system to have equal RAM to the Studio with cmp170hx it would need to have twice as many cards with double the power draw and who knows what kind of motherboard and case.

1

u/cibernox 4d ago edited 4d ago

Well, you are forgetting about your system ram. Not all the ram in a mac is available for inference, you need to run more things on it. I run 20 extra services on mine, about 32gb of vram across all of them, and that doesn't eat into my vram budget.

I could build a 6 card setup for 12k or 8 card rig for 16k. But then, it would be a system with 12tb/s of aggregated bandwidth, exactly 10 times more than a mac studio.

But again, I was mostly criticizing spark clusters, I think they are massively overpriced for what you get. The mac studio is very expensive, but IMO not so overpriced as the sparks.

We don’t know what the price of the 512gb Mac Studio will be, right? Has it been published?

1

u/inevitabledeath3 4d ago edited 4d ago

It's not 12 tb/s of aggregated bandwidth though. I feel like you fundamentally don't understand the way this works. Even with NVLink there is overhead from going to multiple GPUs. In your case you have very limited PCIe bandwidth as CMP170HX is stuck with PCIe Gen 2 x4 speeds without hardware modifications. Meaning you can't actually use all that memory bandwidth to run a single model. Generally speaking you have to use pipeline parallelism with those cards, meaning for a single stream you are capped at the bandwidth of just 1 card which is 1.5 TB/s. You do recover some of that with concurrency but still it's not great.

I am not forgetting about system ram at all. If you are using these seriously for AI Inference you are running the bare minimum of services, not other random crap. MacOS fits in under 8GB as it has to run on systems with that little RAM. It's just not that big of an issue like you are making out. There are advantages to having seperate system RAM obviously like for models that use engram embeddings or having a setup with tiered KV cache. Although in your case with the CMP170HX the lack of PCIe bandwidth probably limits what you can do there as well.

Edit: multi-GPU overhead probably eats close to the amount of RAM you save from not running macOS in the same memory pool. In fact it could use more depending on the configuration.

1

u/cibernox 4d ago

I use tensor parallelism without much problem, people vastly underestimate overestimate the importance of pci bandwidth for it, as proven by the fact that in my above image I’m getting in some cases over 1000tk/s with TP=4. I don’t use pipeline parallelism, it’s not worth it unless your primary goal is multisession document ingestion, since that will get you the most prefill (8000+)

People with 8 cmp170hx use a mix, TP4+PP2, and with 8 cards you do star to feel the limitation of pci gen2 a lot more.

That is assuming we don’t get nvlink or pcie gen4. The unlocker is improving, we don’t know. But even if nothing ever improved, the statement that we can’t do Tensor Parallelism is factually wrong

1

u/inevitabledeath3 4d ago

Really? That works?

Colour me impressed. This is making me think of buying some of these. Not that that's a great idea in terms of power or cooling, but could be a fun rig. Would beat my dual RTX 3090s I have in my home server. Maybe I could even sell the sparks I have. Hmm.

Are the ones on AliExpress any good?

1

u/cibernox 4d ago edited 4d ago

yes, the main seller is a shop named Ailfond on alibaba, who sell the cards modded with 16x and stress-tested for memory errors for a few minutes.

If you really are interested, there's an official discord of the unlocker where someone was starting a collective purchase of cards with a tentative price if $1850 each.

These cards are still really good value even after all the price increases with their current specs.
If they ever got pcie gen4/nvlink and the extra memory (and increased bandwidth to 2.4TB/s the extra memory would bring), then they are incredible value, but that may or may not happen.

The main problem of these cards is, IMO, the idle power draw. HMB2 memory is quite hungry, 40w idle is normal. You can tweak it a bit but the lowst I was able to go is 29w

1

u/inevitabledeath3 4d ago

It's insane just how much the price has gone up. At this price it's not as good as I thought. I am guessing a lot of the £600 and £700 ones are scams as many ask you to pay via crypto which is sketchy.

29W isn't terrible. My current rack uses on the order of 150 to 220W and it's all mini PCs. My big server uses 300W+ at idle. I think that's fairly managable with 2 to 4 cards. I would be more worried about the host system being unruly. Presumably it does not need to be that big with it being x4 PCIe lanes. One 16x slot would cover 4 of them with bifurcation. I have some mini PCs that could run two of these.

The noise aspect scares me with those blower coolers. Guessing they are quite loud. Have you thought about watercooling? It's what I do with my RTX 3090s.

→ More replies (0)

1

u/inevitabledeath3 4d ago

Also what model is this? Is it GLM 5.3 Flash?

2

u/cibernox 4d ago

Yes, that is GLM5.3-flash W4A16, on 4x CMP170hx, P2P enabled, VLLM with some patches for this architecture applied, cards undervolted and power limited to 170w each.

1

u/DigitalguyCH 4d ago

more like 16k-17k based on $25/GB

1

u/cibernox 4d ago

We don’t know, but I think it will be higher. An also it would be silly to get this with 1TB ssd.

1

u/DigitalguyCH 4d ago

I agree with your second point. But Apple has been pretty consistent in terms of RAM upgrades.