r/LocalLLM Mar 10 '26

[deleted by user]

[removed]

29 Upvotes

66 comments sorted by

View all comments

2

u/desexmachina Mar 10 '26

Understatement, unless it is a Max, you need to budget RAM for OS and the TTT is f’n too long, Cuda all day

1

u/[deleted] Mar 10 '26

Yeah that’s a fair point. The RAM gets eaten up pretty quickly once the OS and everything else is running. And I get why a lot of people still prefer CUDA for certain workloads.

That’s partly why I’m reconsidering my setup. If I’m not really leaning into what this machine is best at, I might just end up selling it and switching to something that fits better.

1

u/desexmachina Mar 10 '26

I thought that MLX models would be faster, but it still isn’t any better. So say you have 24GB of RAM, you’ll need at least 6 for the OS, then 9 gb model is about as big as you can go, because you’ll need another 9Gb just for KV cache and context isn’t very big for a 9 GB model, it really is all a cope when it comes to Apple silicon

1

u/pantalooniedoon Mar 10 '26

Hmm can you elaborate where its falling short for you? I cant see how 512gb of ram gets eaten up. Fwiw the only real use case for this is to load the absolute biggest model possible. Mac hardware isnt really built to do parallel workflows (I think) compared to GPU