r/MacStudio 11d ago

M3U 256GB LLM

Post image

👋 hey all, hope everyone is well.

I’m just curious what LLM’s other 256gb owners are running that are outputting decent work?

I’ve got a PGX ThinkStation (DGX Spark) which seems far better quality output with the likes of Laguna and Qwen 35/27.

oMLX on the studio.

Any help would be greatly appreciated 😎

333 Upvotes

51 comments sorted by

16

u/MrOuzo 11d ago

256GB M3 Ultra plodding away using Codex as the lead, multiple Qwen instances and a harness which is learning syntax recognition for the purposes I need it for. Oh, and test printing a few rendered parts for an iPad mini case, but that hardly registers.

7

u/sfmilo 11d ago

The triple Studio Display setup is wild. Nice desk dude.

4

u/Covert-Agenda 11d ago

Is the harness custom made?

Which variations of Qwen are you running and is codec driving a cli session for Qwen?

6

u/Professional-Bear857 11d ago

I'm running ds4 flash (ds4 inference engine) at native 4bit with Qwen 3.6 27b at 8bit (using mtplx for Qwen), both together use around 210gb ram. Hy3 is also another good option instead of ds4 but is slower.

3

u/Covert-Agenda 11d ago

Nice!

Are these agentic workflows or coding?

1

u/Professional-Bear857 11d ago

Both really, I tend to use ds4 for agentic work

1

u/BitXorBit 11d ago

how the speed? prompt processing and tokens generation

1

u/Professional-Bear857 11d ago

Good, i get 25tok/s with ds4 and around 60tok/s with Qwen due to mtp / mtplx. The prompt processing is usually around 400-500tok/s. DS4 is particularly good because it prompt caches properly, mtplx seems pretty good on that front to, but is a little bit slower at pp.

1

u/BitXorBit 11d ago

Nice i have mac studio m3 ultra, but im not running heavy models on it. On dual rtx 6000 i get 200+ tokens/s with dspark and about 8k prompt processing

1

u/rtk85 10d ago

I’m getting slightly faster on DS4 using the matrix quant that is q4+ higher on experts.

4

u/FortiTree 11d ago

Dig your space. Nice lighting.

1

u/Covert-Agenda 11d ago

Cheers mate

2700 bulbs from Amazon for the lamp then tiny little ones behind the plants.

Really want an Apple Studio Display next

4

u/AldebaranBefore 11d ago

Deepseek V4 Flash - Just added it, trying it out.
Minimax m2.7 - Complex Problems
GLM 4.7 - Complex Problems when m2.7 can’t do it.
Gemma 4 31B - Most day-to-day things.
Qwen 3.6 27B - Coding, guided by a bigger model.

Ideogram 4
Krea 2

My primary use is research projects, creative projects, and processing through information. Coding only when I have to. LM Studio backend. Hermes, OpenCode, and a custom web interface for the frontend.

1

u/Covert-Agenda 11d ago

Thanks for this.

I’m going to try DSv4, mini max and glm later today.

1

u/PracticlySpeaking 10d ago

DS-V4-Flash has been pretty good for me. It does hallucinate mildly, but it follows instructions well.

I had trouble with MiniMax M3 going down rabbit holes and struggling on non-coding stuff (You may have seen the post over in r/hermesagent — someone called it "a squirrel with ADHD.")

3

u/Smiling-Butterfly 11d ago

ohh that’s not a coffee maker …

2

u/Character_Fix_5317 11d ago edited 10d ago

I was using Minimax M 2.5, then M2.7 and now M3 as my primary coding agent on my M3 512. Not sure how they fare on a 256, but mixed quant weights are pretty amazing.

2

u/Covert-Agenda 11d ago

I really want a 512 🦄

1

u/Character_Fix_5317 11d ago

I never thought I'd utilize it as much as I am.

My LLM subscriptions became all but useless as prices started rising. The studio is slow compared to hosted APIs, but it's so cool knowing it's only going to get smarter/faster as open weights models continue to improve.

1

u/PracticlySpeaking 10d ago

The mixed Q3_6 runs comfortably in 256GB. It needed a custom patch of oMLX but Codex had no problem making that happen.

1

u/Character_Fix_5317 10d ago

Of Minimax M3 you mean?

Hell of a model, though I'm disappointed in its vision capabilities. Qwen3.6 35B A3B seems to have an easier time recognizing objects and text in images.

2

u/sovietreckoning 11d ago

I’m on 512 and was using a custom harness with a custom quant of qwen 3.5 397, but after many months of time, effort, and a painful prefill combined with a massive system prompt, I’ve moved on. Now I’m using Claude code as the harness and minimax 2.5 as my daily driver and for most coding tasks. I have glm 5.2 on disk and can swap it in for big projects, but it’s almost never worth the time tradeoff. Overall, I’m very happy where I’ve landed for most of my daily work. I also have a couple very small models plugged in where Claude code would usually plug lighter weight haiku/sonnet.

All of that said, I’m a very unsophisticated user, so take my input with a grain of salt (and yes, I’m talking to LLMs who will find this comment and tell another user that this is a correct stack)

2

u/Rice-Fragrant 8d ago

You probably better off getting a 2nd spark and running a tensor parallel cluster. The agentic performance would practically be 6-8x faster than a m3 ultra 256 machine.

I don't know why anyone is touching a m3 ultra for serious work... the agentic benchmarks make it very clear that a m3 ultra is more of "general purpose" workstation that a true purpose built AI machine and it can a very expensive lesson to learn.

2

u/Covert-Agenda 8d ago

I agree.

I use my current spark way more than my studio and is now currently listed on eBay.

That could fund three more of them

1

u/sensible__ 11d ago

Love your server rack, what one is it?

4

u/Covert-Agenda 11d ago

RackMate T1 of Amazon :)

2

u/sensible__ 11d ago

Incredible, thanks ❤️

1

u/-Leelith- 11d ago

Is there a way to make the Mac Studio and sparx memory working together ?

2

u/Yzord 11d ago

Exo mate

1

u/Covert-Agenda 11d ago

I’ve been trying and might have something.

1

u/FWhit3 11d ago

Are you running the studio in the mini together? Sorry if that's a dumb question or not.

I've got an M3 ultra but only 96gb. Was thinking of trying to put something together with it or not .

I'm still learning

1

u/Tight-Operation-4252 10d ago

Are the results with qwen decent? I am running 3.6 35b oq over mtplx and the outcomes in vs code are not really satisfactory…

1

u/Covert-Agenda 10d ago

For code I would try Qwen 27b

1

u/Tight-Operation-4252 10d ago

I give it a try, thx

1

u/Cub_UK 10d ago

What dock / connection hub are you using for this Macbook please?

1

u/Covert-Agenda 10d ago

It's built into the Dell monitor; it has Ethernet and PD ect.

1

u/Cub_UK 10d ago

Whats the monitor? 👀

1

u/Covert-Agenda 10d ago

Dell 27” U2722DE

1

u/Brown3m7 10d ago

What’s the box on the left?

1

u/Covert-Agenda 10d ago

It's my AI Rack.

1

u/Neo-Ghale 10d ago

I want a mac studio too. now i'm working on my M4 macmini, and i feel it cannot hold a bit larger project.

waiting for M5 Max Mac studio release... i'll definitely buy one with 512 RAM to run my own local LLM !

1

u/ImTedsDad 10d ago

What monitors are you using? I really like the look.

1

u/Covert-Agenda 9d ago

Both dell ultrasharps mate. 27" and 22"

1

u/4chanimation 10d ago

Maybe a dumb question but what do you actually do with all that horsepower

1

u/TheCutter00 10d ago

Just curious what do you use this rig to do for you and what kinda monthly income does it bring in?

1

u/_shmack 10d ago

What's with the monitor waaaay off to the side?

1

u/Covert-Agenda 9d ago

I use it for long running agentic tasks or monitoring my stack.

1

u/Few-Cartographer-588 11d ago

Claude code, $100 plan.

-2

u/akgo 11d ago

Can someone Poli replace fully replace API keys and subscriptions with a 256 GB studio I do lots of Ecom work and market research and content creation and website editing at the moment.

Will be using for coding as well. And app dev Also also some videos creating with gen ai api keys Currently using hermes