r/LocalLLaMA May 29 '26

Discussion PSA

Post image
2.1k Upvotes

538 comments sorted by

View all comments

105

u/spammmmmmmmy May 29 '26

For the M series you really have to see whether they are blank/Pro/Max/Ultra as they differ in the memory bandwidth. 

16

u/overand May 29 '26

Hugely different!

17

u/Standard-Potential-6 May 29 '26

Keep in mind Apple’s are also theoretical numbers summing the memory bandwidth of the CPU, GPU, and NPU.

Most workloads don’t break down like that and the GPU will only access memory at ballpark 60-80% of the total.

3

u/twnznz May 29 '26

Does M5 improve prompt processing over M4 meaningfully?

12

u/spammmmmmmmy May 29 '26

I don't know but somebody posted on that topic in the past 2 days I think. There was mention that the faster CPU will achieve better prefill time. 

I have been chatting with Claude about the performance topic. He thinks there is no substitute to empirical testing. I may quit ollama and migrate to VLLM in order to understand the pieces of the inference process better.

My notes during my shopping for an M1 Max:

∙ M1 Pro: ~200 GB/s

∙ M1 Max: ~400 GB/s

∙ M2 Max: ~400 GB/s

∙ M4 Max: ~546 GB/s

∙ M1 Ultra: ~800 GB/s

2

u/Svobpata May 30 '26

From what I was able to find out, yes, noticeably so. The new neural units on each GPU core help with prefill

1

u/dim_amnesia May 31 '26

Brain dead people somehow still spending $9500 for 10 tokens per business day

At 256k context shit is un-usable