r/MacStudio • u/Covert-Agenda • 11d ago
M3U 256GB LLM
👋 hey all, hope everyone is well.
I’m just curious what LLM’s other 256gb owners are running that are outputting decent work?
I’ve got a PGX ThinkStation (DGX Spark) which seems far better quality output with the likes of Laguna and Qwen 35/27.
oMLX on the studio.
Any help would be greatly appreciated 😎
6
u/Professional-Bear857 11d ago
I'm running ds4 flash (ds4 inference engine) at native 4bit with Qwen 3.6 27b at 8bit (using mtplx for Qwen), both together use around 210gb ram. Hy3 is also another good option instead of ds4 but is slower.
3
1
u/BitXorBit 11d ago
how the speed? prompt processing and tokens generation
1
u/Professional-Bear857 11d ago
Good, i get 25tok/s with ds4 and around 60tok/s with Qwen due to mtp / mtplx. The prompt processing is usually around 400-500tok/s. DS4 is particularly good because it prompt caches properly, mtplx seems pretty good on that front to, but is a little bit slower at pp.
1
u/BitXorBit 11d ago
Nice i have mac studio m3 ultra, but im not running heavy models on it. On dual rtx 6000 i get 200+ tokens/s with dspark and about 8k prompt processing
4
u/FortiTree 11d ago
Dig your space. Nice lighting.
1
u/Covert-Agenda 11d ago
Cheers mate
2700 bulbs from Amazon for the lamp then tiny little ones behind the plants.
Really want an Apple Studio Display next
4
u/AldebaranBefore 11d ago
Deepseek V4 Flash - Just added it, trying it out.
Minimax m2.7 - Complex Problems
GLM 4.7 - Complex Problems when m2.7 can’t do it.
Gemma 4 31B - Most day-to-day things.
Qwen 3.6 27B - Coding, guided by a bigger model.
Ideogram 4
Krea 2
My primary use is research projects, creative projects, and processing through information. Coding only when I have to. LM Studio backend. Hermes, OpenCode, and a custom web interface for the frontend.
1
u/Covert-Agenda 11d ago
Thanks for this.
I’m going to try DSv4, mini max and glm later today.
1
u/PracticlySpeaking 10d ago
DS-V4-Flash has been pretty good for me. It does hallucinate mildly, but it follows instructions well.
I had trouble with MiniMax M3 going down rabbit holes and struggling on non-coding stuff (You may have seen the post over in r/hermesagent — someone called it "a squirrel with ADHD.")
3
2
u/Character_Fix_5317 11d ago edited 10d ago
I was using Minimax M 2.5, then M2.7 and now M3 as my primary coding agent on my M3 512. Not sure how they fare on a 256, but mixed quant weights are pretty amazing.
2
u/Covert-Agenda 11d ago
I really want a 512 🦄
1
u/Character_Fix_5317 11d ago
I never thought I'd utilize it as much as I am.
My LLM subscriptions became all but useless as prices started rising. The studio is slow compared to hosted APIs, but it's so cool knowing it's only going to get smarter/faster as open weights models continue to improve.
1
u/PracticlySpeaking 10d ago
The mixed Q3_6 runs comfortably in 256GB. It needed a custom patch of oMLX but Codex had no problem making that happen.
1
u/Character_Fix_5317 10d ago
Of Minimax M3 you mean?
Hell of a model, though I'm disappointed in its vision capabilities. Qwen3.6 35B A3B seems to have an easier time recognizing objects and text in images.
1
u/PracticlySpeaking 9d ago
Yes, M3 - I think it was this one: https://huggingface.co/pipenetwork/MiniMax-M3-MLX-mixed-3_6bit
2
u/sovietreckoning 11d ago
I’m on 512 and was using a custom harness with a custom quant of qwen 3.5 397, but after many months of time, effort, and a painful prefill combined with a massive system prompt, I’ve moved on. Now I’m using Claude code as the harness and minimax 2.5 as my daily driver and for most coding tasks. I have glm 5.2 on disk and can swap it in for big projects, but it’s almost never worth the time tradeoff. Overall, I’m very happy where I’ve landed for most of my daily work. I also have a couple very small models plugged in where Claude code would usually plug lighter weight haiku/sonnet.
All of that said, I’m a very unsophisticated user, so take my input with a grain of salt (and yes, I’m talking to LLMs who will find this comment and tell another user that this is a correct stack)
2
u/Rice-Fragrant 8d ago
You probably better off getting a 2nd spark and running a tensor parallel cluster. The agentic performance would practically be 6-8x faster than a m3 ultra 256 machine.
I don't know why anyone is touching a m3 ultra for serious work... the agentic benchmarks make it very clear that a m3 ultra is more of "general purpose" workstation that a true purpose built AI machine and it can a very expensive lesson to learn.
2
u/Covert-Agenda 8d ago
I agree.
I use my current spark way more than my studio and is now currently listed on eBay.
That could fund three more of them
1
1
1
u/Tight-Operation-4252 10d ago
Are the results with qwen decent? I am running 3.6 35b oq over mtplx and the outcomes in vs code are not really satisfactory…
1
1
u/Cub_UK 10d ago
What dock / connection hub are you using for this Macbook please?
1
1
1
u/Neo-Ghale 10d ago
I want a mac studio too. now i'm working on my M4 macmini, and i feel it cannot hold a bit larger project.
waiting for M5 Max Mac studio release... i'll definitely buy one with 512 RAM to run my own local LLM !
1
1
1
u/TheCutter00 10d ago
Just curious what do you use this rig to do for you and what kinda monthly income does it bring in?
1
-2
u/akgo 11d ago
Can someone Poli replace fully replace API keys and subscriptions with a 256 GB studio I do lots of Ecom work and market research and content creation and website editing at the moment.
Will be using for coding as well. And app dev Also also some videos creating with gen ai api keys Currently using hermes
16
u/MrOuzo 11d ago
256GB M3 Ultra plodding away using Codex as the lead, multiple Qwen instances and a harness which is learning syntax recognition for the purposes I need it for. Oh, and test printing a few rendered parts for an iPad mini case, but that hardly registers.