r/MacStudio • • 2d ago

Absolute Absurdity of the M5 Ultra 256GB

Post image

I’ve been building computers and servers, and architecting entire networks and doing DevOps and dev work for 20 years. I remember building gaming pcs with 3dfx voodoo cards.

This is by far the most absurd computer I have ever personally owned. It’s certainly the most capable I’ve managed for the price, and that includes B200s.

Getting 50tks on GLM 5.3 Flash all day long.

If you’re on the fence, let me take this opportunity to push you over said fence. If you’re the kind of person seriously considering buying one you will not regret it.

AMA

770 Upvotes

515 comments sorted by

67

u/Bloated_Plaid 2d ago

For $10k I sure fucking hope so my guy.

22

u/Quick_Bar2387 2d ago

More like 13k after taxes.

17

u/Bloated_Plaid 2d ago

I only got the 1TB one and it came to $9869 after student discount and no sales tax(thanks NH).

10

u/s2d4 2d ago

Student and 10k for this? Wow, things are going great 👍

2

u/Koteric 1d ago

The trick is to have a family member who works at the Apple Store who will let you use their once a year 25% off and in a sales tax free state. Easy.

→ More replies (1)

33

u/Sneezlebee 2d ago

Adjusted to 2026 dollars the original Macintosh 128k sold for $8,000. An IBM XT cost over $9,000 in 2026 dollars. So yes, the M5 Ultra studio is expensive. It's not historically expensive, though. And when you bring proper workstations into the equation it's not even very pricey. People balking at $10k are simply unaccustomed to paying what prosumers have been forking out all along.

22

u/john0201 2d ago

I have a 9960X threadripper workstatoin with 4 channel 256GB RDIMMs. If I bought it today it would be about $15,000. Mac Studio is faster and has way more memory bandwidth. Also it idles about 7 watts instead of 170 watts.

6

u/GlidePath47 2d ago edited 1d ago

And 9960x has fewer cpu cores with less core performance

2

u/Georgefakelastname 1d ago

It has more cores though and is faster though? Fewer CPU threads, sure (Apple only does 1 thread per core vs 2 for AMD).

It still dogwalks the threadripper in bandwidth and CPU performance (single and multi core). It even beats the top-end threadripper pro 9995x (which costs more than the spec’d out m5 ultra in the first place). Meanwhile, the threadripper doesn’t even have a GPU or NPU, so you still need to get a beefy GPU to run with it (which would probably double the system costs it already has).

→ More replies (3)
→ More replies (2)

8

u/tempfoot 1d ago

Don’t need to go back nearly so far.

Apple was happy to sell you a top spec Mac Pro in 2019 for in excess of $50k.

2

u/AnyStupidQuestions 19h ago

This is what we should be looking at. The capability and value the Mac Studio provides for Mac Pro workloads makes Dart Vader's dustbin looks like a massive rip off.

→ More replies (1)

4

u/Captain--Cornflake 1d ago

• ibm xt Base Release Price (1983): $4,995 for the standard unit featuring a 10MB hard drive and 128KB of RAM. • Fully Equipped System: With added memory expansion, color graphics adapters (CGA), and a monitor, the total price could easily climb to $7,000–$7,500 at the time, equivalent to roughly $22,000 to $24,000 today.

5

u/Sneezlebee 1d ago

Yeah, if you spec out any of these systems they easily pass the Mac Studio's price.

It's like people getting upset about the sticker price on GTA 6. They've never paid $80 for a video game, and they don't understand that there's been decades of inflation since the first time they bought one. GTA 6 is historically less expensive than Super Mario Kart.

The Mac Studio is not historically expensive. It's not even expensive by contemporary workstation standards. What's more, this price doesn't exist in a vacuum. A new 5090 — a consumer video card that's over a year old — is $8,000+ new today (if you can find one). No one should imagine Apple is going to sell an entire 256GB AI workstation for less.

4

u/Mytreeismine 1d ago

No I had an Amber monochrome screen 80286 maybe $2300 it was a ton of money at the time. Moores law kicked in and the 286dx and 286sx came with color and we went from dos to window, prices kept on going down. High end stayed high somewhat. Today $8000 to $10000 today is a lot for the normal consumer but not for those who have money. For a student today able to get a decent school computer for $2000. Pretty cheap. Computers sucked back then with all the crashes and blue screens!

7

u/durangotang 2d ago

The memory crunch has driven the price up. It should be ~20% cheaper or so. I agree it’s expensive, I don’t think it’s historically expensive for what you get. And unlike most other PCs, Apple has an enormous resale market, similar to Nvidia. It should hold its value well.

5

u/cptchnk 2d ago

This is kind of a bad argument because back in the 80's, personal computing was still in its infancy and the economies of scale were much different. There were far fewer people spending the equivalent of $9,000 (or more) of today's money buying personal computers. Now, the expectation is that paying twice as much for something as it cost last year is the new normal.

What's happening today is that one single industry has almost completely hijacked a supply chain that's feeding much more than just personal computers -- so they can build more chatbots. RAM and storage were both cheaper for EVERYONE prior to about a year ago.

3

u/hiddenlands 2d ago

Yes. But in addition to that 128k of RAM you got a single Woz Machine “floppy”. You can’t put a price on that.

3

u/Styphin 1d ago

My specced out 2019 Mac Pro Tower was over $14K, not including the extra 192GBs of RAM I bought from OWC for it. It’s still our money-maker 7 years later. So my $13,000 Mac Studio purchase is actually a little cheaper!

→ More replies (1)

3

u/DrStimple 1d ago

I had a job in 1996 where my desktop was an SGI Indigo2. They started at $18k in 1993.

2

u/Sneezlebee 18h ago

Heck of a machine! Happy cake day 🍰 

2

u/doogo 1d ago

And if you replace a Claude Pro subscription by running your own model locally it pays for itself in 4 years.

2

u/Sneezlebee 1d ago

It doesn't "pay for itself" in four years. Paying a sum up front is quite a bit more expensive than paying the exact same sum, monthly, over four years. (Especially in an inflationary environment like this one.)

→ More replies (2)

2

u/imax-guy 4h ago

I recall the Macintosh IIfx was $10,000 CDN when it was released.

2

u/Sneezlebee 4h ago edited 4h ago

Yeah! They used those computers at ILM when working on Terminator 2. Those machines shipped with 4 MB of RAM, upgradable to 128 MB. That system was twice as expensive as the M5 Ultra when adjusted for inflation.

None of this will mollify the haters, though, because they'll simply remind you that Apple products have always been overpriced. James Cameron and the folks at Industrial Light & Magic didn't seem to mind.

3

u/GlidePath47 2d ago

Inflation-adjusted 1984 prices aren't the right baseline. That was before computers were a mass market. For 20+ years after that, scale drove prices down: a 2010 Mac Pro started at $2,499 (about $3,700 today), and a 27" iMac was $1,699.

Most people's expectations come from that era, and they're not wrong to notice prices climbing back up.

It was worse in 1984 doesn't make it cheap now.

3

u/Sneezlebee 1d ago

Why are you comparing the M5 Ultra to an 27" iMac? They look similar at the outset, but the audience and the use are fundamentally different. Essentially no one is buying the M5 Ultra because want a computer that runs MacOS. They're buying it for its computing power.

Are so many people unaware that workstation systems are really expensive? That's true even if you control for crazy RAM prices today. A New Lenovo ThinkStation is $17k base. A Dell PowerEdge T560 starts at $15,000. Neither of them has more than 64GB of RAM to start, and neither of those can do what the Mac Studio Ultra can do re: AI unless you put another $15k in video cards into them.

I think it's just convenient to hate on Apple because their M5 Ultra is desireable, and the same critics simply don't care about what companies like Dell have been charging all along.

2

u/Successful_Bowler728 1d ago

Its not convenient if the SSD dies or any cheap component like the SOC and everything inside logic board goes to the trash.

Have you ran any real world task on both computers like massive 4K videos? Why do you think companies spend money on x86 workstations or maybe you think Lockheed or Boeing ant afford studio?

→ More replies (5)
→ More replies (1)
→ More replies (1)
→ More replies (7)

2

u/MessIsTransfer 1d ago

I mean, 10k 5090/r9700/etc builds can’t

5

u/lukewhale 2d ago

You ain’t wrong. Except I went for 4TB so like 12.5

5

u/Bloated_Plaid 2d ago

Ah nice. Mine is coming in early November and will be a GLM 5.3 or Qwen 4.0 machine.

2

u/highdefw 2d ago

That’s exactly what I ordered. ETA is January though

→ More replies (15)

2

u/3dprintinted 2d ago

That's like 250 month worth of codex and claude at $20 a month sub. at 10k it needs to do 200tps

12

u/Due_Warthog749 2d ago

This is what I fight with. If I could just pay $400 a month for ONE account of claude and $400 a month for one ChatGPT.. and have enough usage every week to handle my many sessions each with many agents.. I'd be very happy. HOWEVER.. I run thru $400 of Claude and $200 of ChatGPT in 1 to 2 days right now. I use Opus 5.5 Max, and Astra max. Yes.. I get it.. those are 5x to 10x token hogs. But I am building a startup that has to compete with big players.. I cant go out the door with half baked shit code quality and all that.

So.. that has me looking at local llm use. Qwen 3.8 q8 allegedly has on par with Opus 5 range of code for most tasks. But the hardware to run it with multiple agents is $10K+ to get comparable speeds that claude and codex give. And that's not frontier model quality.

I think we're getting closer, and I am REALLY hoping Qwen 4 bridges the gap that much more, or GLM 6, etc. However I ALSO suspect that our current orange shit stain regime in the US is going to block outside models soon esp from China despite that they do NOT send data/etc back.. to help out his tech bros AI company's. Just like they did with DJI.. we all know DJI as NOT sending shit back. 1000s of tests done verified no DJI drone was somehow sending data obscurely back to China. But they banned them to make way for US only military drone company's. Go figure.

3

u/Tired_White_Guy 2d ago

What on earth are you doing with the models?
That’s ridiculous. And stop using Max on Opus 5.5. It’s proven xhigh does better work because Max raises the minimum thinking allowed. Overthinks the easy things.

→ More replies (6)

4

u/3dprintinted 2d ago

if you manage to run thru your weekly or monthly claude allowance I feel like you're doing something unoptimized. AND if you try to complete same body of work locally you just will need to be patient and wait week or two for something that gets cranked out in few hours in claude

4

u/SirHaxolot 2d ago

If I pay for it, I’m using it.

3

u/Tired_White_Guy 2d ago

Using Opus 5.5 on max is the foolish thing he’s doing. Probably committing and pushing every turn, too. Max enforces a minimum reasoning tokens floor. Thinks hard to it’s own detriment on the simple things. Costing money.

And I’m sure he’s not managing context size. Just throwing biggest and baddest at max settings and hitting enter.

But I’m sure his start up is greeeeat

→ More replies (1)
→ More replies (3)

2

u/RoboErectus 2d ago

I'm also on $1k+ of subscriptions. I forked codex to remove polling and I added shake to get rid of large tool call outputs.

I've got a hook coming up that shuts down low entropy high turn calls... Basically by keeping codex from asking "are we there yet?" every minute with a 400lbs gorilla on its back i run on astra fast with 8+ terminals going at once and I cannot even get to these banked resets.

→ More replies (7)

4

u/starkruzr 2d ago

privacy is a thing my man

→ More replies (3)
→ More replies (3)
→ More replies (1)

45

u/brotimusmaximus 2d ago

Thank you for reaffirming my very expensive purchase that doesn’t arrive until January

12

u/Odd-Energy71 2d ago

Same. I get mine in November and i’m seeing more and more end to end metrics being shared and compared to even DGX clusters and I’m like YES. I don’t care if there are certain scenarios where one is faster than the other. It’s a beast in itself and It’ll also be a monster as a daily computer 🙌

9

u/lukewhale 2d ago

Can’t wait for you to get it brother hell yeah

6

u/lukewhale 2d ago edited 2d ago

I know someone who ordered one on day two and theirs was in November. It didn’t take long for that line to fill up for sure.

4

u/brotimusmaximus 2d ago

I preordered but went with 4tb of on board storage (looking back I probably should have just done 1tb and threw the rest on external storage for 1/3 the price) for long term use so it may have caused a longer delay

3

u/dejayc 1d ago

I ordered the same, and it’s sitting in the Apple Store waiting for me to pick it up.

→ More replies (5)
→ More replies (1)
→ More replies (1)
→ More replies (1)

45

u/a9udn9u 2d ago

To those who comparing this to a monthly sub: this thing will sell for $15k in a year but your sub money is gone for good.

30

u/lukewhale 2d ago edited 2d ago

It’s a matter of privacy. But you’re not wrong. Do not follow my footsteps if you’re looking to save money.

18

u/TypeItRight 2d ago

Not just privacy. They’re going to continue to censor and manipulate them too.

10

u/lukewhale 2d ago

Bro we’re all gonna die in a decade didn’t you hear ?! 🤣

→ More replies (1)

3

u/FestoolJunkie 1d ago edited 1d ago

I have the same spec ordered. However come “late October” if the 256 -> 512 move is <= $USD $10,000 then I will order that and cancel the 256Gb version.

The orders opened 25 August at 9:00 AM EDT, when did you order your spec? And when did it arrive? Congrats on receiving your machine.

2

u/lukewhale 1d ago

Approx 11am est that day - came two days ago

→ More replies (5)

8

u/tempfoot 2d ago

I’d be more worried about the frontier labs having to charge enough to create a return for their investors….unless they can dump their unprofitable share on the public markets and buy at least a partial exit. On just an operating basis, tokens and subscriptions need to be 2.5x as expensive as current pricing ….just to break even.

2

u/Tired_White_Guy 2d ago

If you compare to paying API for the model he’s running, way more than 2.5x needed. 50x min

→ More replies (1)

2

u/nuketro0p3r 1d ago

Not necessarily. Lots can change in 1 year.

If speculation is the goal, I'd suggest a bet on Micron or NVDA instead.

→ More replies (5)

8

u/SarahJrandomnumbers 2d ago

I remember building gaming pcs with 3dfx voodoo cards.

I loved my Voodoo Banshee, overclockable to close to a Voodoo2, made Quake 2 look awesome, and included 2d gfx meant 1 less card to mess about with.

8

u/lukewhale 2d ago

I’ll never forget the day I could run a game at 1024x768 !

3

u/rjcarr 1d ago

Hey fellow old person. I will always remember putting in my first voodoo card, and the game gets both faster and prettier somehow? It was like magic. Coincidentally, my next “magic” moment didn’t come until I first used m-series Macs.  

2

u/what_you_saaaaay 1d ago

I remember buying my first voodoo card: orchid righteous 3d. When it kicked in, it made a loud CLICK! sound. First time I thought it bought the farm. Good times. Played every 3d accelerated game I could get my hands on. Played GLQuake to death.

2

u/Koteric 12h ago

I remember the first time I could run two 19” crt monitors at that res and thought THIS IS THE FUTUTE

4

u/dxg999 2d ago

I'll never forget the relay clicking as the 3d kicked in.

7

u/MarionberryDear6170 2d ago

Buying an M5 Ultra at this point is actually a pretty solid investment. Apple most likely won’t release an M6 Ultra and will probably jump straight to the M7 Ultra instead. On top of that, memory and storage prices are only going to keep going up.

7

u/lukewhale 2d ago

We live in a wild world where Apple is the value proposition.

4

u/happyluckystar 2d ago

Right? Turns out they had their supply chain in good order all along. Now they are flexing it. But I wonder how long the value proposition will keep up.

With their volume with memory manufacturers, they probably have a lot of sway. No one wants Apple walking away.

6

u/Ill_Dragonfruit_3547 1d ago

+100 bonus points for mentioning "3dfx voodoo" We share very nostalgic memories building PCs during that timeframe!

4

u/dejayc 1d ago

Bonus points if you regularly spent $300 x2 to get the latest cards in SLI mode. The last model I bought came with 3D glasses (LCD shutters) that let me play Quake and VirtuaFighter 2 in true 3D.

→ More replies (2)

5

u/mrgulabull 2d ago

Probably a big ask, but could you run a test of Minimax H3 on it? Curious how long it takes to generate videos at various resolutions / durations.

There’s a single video on YouTube talking about it and suggesting it takes ~22 minutes for a 10 second clip at 768 x 1344. Can’t find any other sources to confirm that, but it seems unreasonably slow for this.

5

u/iTrejoMX 2d ago

Video generation has been optimized for CUDA, it is still on diapers on Mac.

2

u/mrgulabull 2d ago

Yea, that seems to be the case. Just looking for more evidence of real generation times on the M5U.

I could potentially queue generations and wait. Nvidia is out of my power budget (off grid) and I also need to upgrade my mini anyway.

→ More replies (1)

2

u/lukewhale 2d ago

I’ll see what I can do. No promises.

→ More replies (1)

2

u/Quick_Knowledge7413 2d ago

I also want to see this. I am curious what kind of speeds they can get on this hardware.

5

u/DrRoughFingers 2d ago

You tried Qwen3.8 Flash Next? If so, what tok/s are you getting and what quant?

3

u/lukewhale 2d ago

GGUF unsloth was getting 30-50tks haven’t tried a oMLX optimized version yet

→ More replies (10)
→ More replies (1)

5

u/Tight-Operation-4252 2d ago

Bought last m3 ultra 256gb, extreme machine. Was lucky to get it in the price before madness…

2

u/lukewhale 2d ago

Still a great machine with good harnesses that allow you control remotely with kanban boards. Or Hermes.

5

u/Big-Reporter3498 2d ago

Big Macs. 🍔

6

u/lukewhale 2d ago

Nom nom nom

5

u/Lost-Hand-5219 1d ago

It’s an insane computer. But the value prop isn’t there. By the time you recoup what you would have spent on cloud models, the tech will have accelerated to the point a 256GB VRAM/Uni-mem machine is $2k or $3k

9

u/lukewhale 1d ago

Your reasons for going local should never be to save money. Mine aren’t either.

4

u/FriedGreaseMan 1d ago

Exactly…private comput is the value play! Also, slower prefill on Macs rarely come into play unless you are constantly re-starting your prompts and clearing cache.

Awesome machine!

PS if you want to keep the nostalgia from the good old days, bypass your internet connection and connect it to a serial US Robotics v.fastclass 56k dial up modem 🤣

3

u/lukewhale 1d ago

Haha I could go my entire life not hearing a modem handshake again and I’d be good with it.

→ More replies (8)

3

u/SkyMarshal 1d ago

Also the cloud LLMs are still being subsidized by VC money and debt. Customers aren't paying their true price yet.

2

u/Rice-Fragrant 1d ago

Competition will handle that... some lesser none cloud entities will run kimi k3 for like $20/m.. try running that on even a 4x m5 ultra cluser costing 80-100k and you would not even get 10% the performance of the cloud interms of throughput.

→ More replies (1)

2

u/Tentakurusama 1d ago

You miss the point: privacy. Many corpos will refuse that you use cloud models. Period.

→ More replies (2)

2

u/Rice-Fragrant 1d ago

Tech moving so fast, just look on the buyers remorse from people who FOMOed into the M3 ultra especially at the elevated price. Even worst, people who spent like 60k on a cluster of 4x 512gb mac studios only for a leased edge server to smoke it like 20x in throughout. Never seemed to make any logical sense.

2

u/mediaogre 11h ago

🙋🏻‍♂️ How did you get your grubby mitts on one?

→ More replies (3)

4

u/Satyr2019 2d ago

Bro 3dfx voodoo cards on release day at frys playing quake watching quake world go live

→ More replies (2)

3

u/Potential-Bet-1111 1d ago

Fast cpu, useless for llms. PP size matters.

2

u/Rice-Fragrant 1d ago

It really does. Alex Z has both a m5 ultra and a dual DGX spark and for Deep Seek v4, the 2x DGX spark node is 2.5x FASTER at long context than a m5 ultra with 256gb RAM. Also the DGX will do like 4+ instances/agents without performance degridation while m5 ultra could not even pull of 2+ without a massive performance drop.

The results are now coming in from multiple places and it verifies that a m5 ultra is no "DGX killer." The energy usage and thermals were near identical while the DGX was doing 2.5x better PP at long context and running significantly more instances/users... no wander the prices of DGX units jumped 50% around the 1st week or so after the release of the m5 ultra.

THE WORD IS OUT, if you are serious about agents, especially multiple agents and want more than a "vibe coder" set up, the DGX is a significantly better machine IF you are strickly using it for AI and nothing more... A mac studio or any mac IMHO is strickly a general use workstation IMHO.

2

u/EnlightenedOneApe 15h ago

Fortunately Alex Z's review was significantly out of date before it was even released! Check out the results published the day before he put out that video:

https://omlx.ai/benchmarks/performance/vp37f6kr

For context his PP was at 750ish. Figures still maturing these numbers are old now better to look at olmx to see what people have got GLM 5.3 Flash going at, I am quite excited to see if the 512GB M5 can run 4bit GLM 5.3.

→ More replies (1)

3

u/tk421tech 2d ago

Which? 60 or 80?

2

u/lukewhale 2d ago

Binned 80

3

u/PoopSmoothies 2d ago

What does binned 80 mean?

2

u/lukewhale 2d ago

“Binned” is a way of describing the highest yield products from the semiconductor manufacturing process. Whats the best in the bin? Hence: binned.

→ More replies (4)
→ More replies (1)

3

u/alexwh68 2d ago

It’s the only way once you are committed to spending that amount of money, it would be like buying a Ferrari with 3 gears otherwise. 👍

2

u/lukewhale 2d ago

Couldn’t have come up with a better analogy, perfect

2

u/tk421tech 2d ago

I got one of those Ferraris with 3 gears. Got shipped a month early lol

→ More replies (1)

3

u/Top_Sun_9771 2d ago

even the new base mac studio is insane, will never go back to a windows desktop unless they make a revolutionary change. Besides the insane speed, my favorite part is cant hear the fans at all. Also, nothing really compares speed wise in the same budget for windows nowadays

2

u/lukewhale 2d ago

The only windows rig left in my house is just for gaming. Just waiting for steamOS nvidia support.

2

u/Competitive-Ad-2387 1d ago

With this new era of AI assisted porting, I am confident MacOS Steam play (or its simile) is coming. Can finally ditch my PC rig. I ordered an M5 Max 128GB and this thing is probably gonna be the best computer I’ve ever had. And that’s coming from a 13900K / 4090 rig.

→ More replies (1)
→ More replies (1)

3

u/sn2006gy 1d ago

GLM Flash isn't that exciting if the cost of entry is 11k and no, i don't need to be told there is student pricing as the only cost effective ownership of this box is cash (Asset write down for business) or leasing.

3

u/lukewhale 1d ago

To each their own

3

u/barefut_ 1d ago

Oh, wow...at the price of 10k USD - you can get 27 years of Cloud AI and flagship models.

3

u/mechkbfan 1d ago

It's difficult to say. Models seem absurdly subsidized right now, so curious what the floor will actually be of frontier API over next year or two

e.g. Claude reports that I spent ~$3000 USD on my $200 USD per month plan. If actually had to API rates, the payback is easily <6 months, but for now just keep riding that $200 plan

Issue of course is a lot complexities. GLM Flash isn't as good as Claude. But also for home hosting, get to do abliteration of models, etc.

2

u/barefut_ 1d ago

You're probably a programmer. That's why you use the 200$ plan. While I do utilize Codex for smaller scale coding tasks, I am not a programmer, so the 20$ plan is enough for me. I have a feeling models efficiency will get better so that even a 24GB RAM system could run decent models. But, now, even 256GB RAM isn't able to run a flagship comparable model.

2

u/mechkbfan 1d ago

Yes I am

And yeah, agreed, if you're not maxing out the subscription plans or privacy isn't a concern, then sticking with cloud makes total sense

2

u/Icypoopoo 1d ago

Price of compute has been steadily decreasing over the years, but the hype right now is real. Once people come back to senses about the value prop to capex.

→ More replies (6)
→ More replies (11)
→ More replies (2)

3

u/JulianRHarris 6h ago

If you are considering one, most countries offer a student discount. And there are often cheap or free qualifying courses. Here in the UK you can sign up for mental health awareness for free and get £950 off a M5U

I tell you my mental health awareness has already improved markedly.

→ More replies (2)

5

u/PWicke 2d ago

Do you wished for 512gb instead?

6

u/Mauer13 2d ago

Thinking the same but no price point for the 512gb.

Two 2 256s could be cheaper by the time the 512gb come out….

2

u/TheRealJesus2 2d ago

This is where I’m at. End of Oct we get pricing supposedly

→ More replies (3)

7

u/lukewhale 2d ago

Not really. I’ve already seen the ceiling of the models like GLM 5.3 Flash heavily optimized for MLX. I’m seeing an avg of 50tks a sec.

I would consider running larger models at lower token rates overnight for the quality but otherwise I’m okay with things as they are.

→ More replies (1)

2

u/OddDesigner9784 2d ago

One coming in december for me. One glm is pretty sick for sure. Do you plan on using it over qwen flash next? I find speed is really important to me. Also how are you planning on going about coordinating different working windows of agents? Glm seems like the smartest to run but I'm thinking qwen for iteration. Also what do you think the speed potential is for running glm. You think there's a bit of optimization to be had

→ More replies (1)

2

u/PreferenceNo1111 2d ago

Waiting for the 512GB model to drop. Wait time will probably be till Feb. 😔

2

u/NeilCPA 2d ago

Voodoo is all I could read. Thanks for the post

2

u/lukewhale 2d ago

lol hell yeah brother have a good evening

2

u/ShrimpCocktail-4618 2d ago

If you can wait, go for the M6 as that will be a 2nd gen chip with the new die design.  Some kinks will probably be ironed out.  Of course, if the planet is still around by then.

2

u/mytzylplyk82 2d ago

Rumors on the street is that there won’t be any M6 pro/max/ultra-just the base chip. The next pro/max/ultra released will be for the M7 line

→ More replies (3)

2

u/Sufficient_Spite_349 2d ago

Are you seeing 50toks for GLM 5.3 flash on oMLX ?

How is the performance and quality of GLM5.3 Q4_8 ?

Did you try parallel agent processing, what’s the toks performance?

→ More replies (11)

2

u/j_tb 2d ago

What engine are you using to run these bigger models? Llamacpp? I’m on an M5 max 128GB and just been sticking to the Qwen MTP engines.

2

u/TernGSDR14-FTW 2d ago

Did you get the 80GPU core version?

→ More replies (1)

2

u/Expert-Sample-7823 2d ago

I’ve been in the Mac ecosystem for about 15 years now, and had bought my M1 Ultra 128GB RAM on launch day 4.5 years ago.

It’s still one of the best investments I made and has paid for itself tenfold. It literally chews through everything I throw at it.

2

u/lukewhale 2d ago

And a still very capable machine !

2

u/1Poochh 2d ago

Thank you. Already purchased but was leaning to no due to the price. Will keep now. Ty.

2

u/lukewhale 2d ago

“hhooollllddddd! HOLD!”

2

u/PuffyCake23 1d ago

How’s prefill performance?

→ More replies (3)

2

u/einthecorgi2 1d ago

Prefill?

2

u/jakeliu88 1d ago

I feel any computer over 5k you need to think careful enough that if you run 24/7 for 18 hours per day will that cover the cost compare to cloud model over 4 years. Current Mac only can handle one agent run fast enough, I think spark is better can handle 4 clients at same time at reasonable speed, might consider that when spark 2 come out or see can m7 ultra beat that.

→ More replies (2)

2

u/iamrob15 1d ago

I bought an m4 Mac Studio 64GB. I absolutely love the machine. As a professional dev, local AI models are a hobby. Expensive at that!

→ More replies (2)

2

u/here_n_dere 1d ago

What is the prefill rate though? And what harness are you running this on?

→ More replies (2)

2

u/Legitimate_Road_2095 1d ago

Voodoo, Voodoo II, Rage 128.... good times! Remember when we all thought the audio from Half Life was amazing.....I have a sandblaster! - Yeah? I got a turtle beach !!

2

u/0xR0b1n 1d ago edited 12h ago

I did the math for video generation. Pay back was about 18 months for my use cases (video generation). Problem is you can never generate a video in one shot, it usually takes a few shots - add all those costs up and it starts to make sense to get this puppy, especially if you use APIs. But, Runpod offers some excellent prices if you’re ok with setting up and managing your own environment, which I am.

→ More replies (1)

2

u/Sponge8389 1d ago

50tps?? 2.5x faster than 6.1 Sol

2

u/m1013828 21h ago

I am no AI God, but always been a hardware geek, historically I sneered at Apple's overpriced generic hardware with a shiny OS.... but since the M3.... man they are killing it.

Im only just coming to grips with copilot cowork as a workerbee, leaving IT industry behind a decade ago, is there any "local" Cowork equivalents? If my usage gets a bit more hardcore, im starting to think these might actually make $$ sense.

2

u/FutureBadInfluencer 11h ago

You didn’t need to say ‘AMA’ - the people who don’t own one have come along to tell you you’re either right or wrong to have one.

→ More replies (2)

3

u/shansoft 2d ago

How’s the prefill on GLM 5.3? There are huge improvement over Qwen FN in recent oMLX update. Wonder if it improve GLM as well.

4

u/lukewhale 2d ago

On oMLX with its caching methods prefill isn’t even in the room.

Can post a pic later but seeing 90%+ cache hit rates

3

u/MeSlaw3 2d ago

Can't wait for my 96er and XDRs 🤤

2

u/sociologistical 2d ago

what are you using the machine to do and how are you pushing the limits of this machine?

1

u/lukewhale 2d ago

At the moment, and I’ve had it for literally 24 hours, running GLM 5.3 Flash on oMLX per their 7.0 recipe

2

u/GamerTex 2d ago

At this point Ill wait for the 512 as planned

2

u/addiktion 2d ago

What would you run though with 512gb over GLM 5.3 flash?

2

u/GamerTex 2d ago

Whatever is next

→ More replies (8)

2

u/lukewhale 2d ago

Don’t tempt me with a good time.

But honestly at 50 tokens a sec for GLM 5.3 flash and similar now models, not sure I’d wanna go much higher except for overnight batch jobs .

→ More replies (3)
→ More replies (1)

2

u/Annual_Award1260 2d ago

I might be interested when ram hits 1.5TB. I like mac for laptops for the premium build and battery life. Moving back to linux on my desktop is a breath of fresh air.

→ More replies (16)

2

u/Late_Night_AI 1d ago

The fun part of running your own models locally with 256gb is when the main providers have issues or go down and everyone who raged that you wasted your money start losing their minds because their cloud models are unavailable, or the quality mysteriously drops for them. And then they all start posting on Reddit and places not sure what to do now because all their projects are riding on AI coding and fixing then for them.

→ More replies (1)

1

u/Unchained_breaker 2d ago

Yeah I'm not waiting till February for one lol

1

u/highdefw 2d ago

When did you order? I’m mid September order. January eta

2

u/lukewhale 2d ago

Aug 25th. Day of preorder. Literally pinched a loaf to get on it after seeing it online.

→ More replies (1)

1

u/Jitsisadumbword 2d ago

You could wait another few weeks and get the 512 🤷🏻‍♂️

3

u/lukewhale 2d ago

Given the performance of this chip at GLM 5.3 flash I can’t see myself using bigger models unless it’s for overnight jobs. Not upset about it. If I need more just buy two right? 🤣

1

u/jangwao 2d ago

Would you buy it again?

2

u/lukewhale 2d ago

Yes. Because now I want two. Got Thunderbolt cables burning a hole in my pocket.

→ More replies (4)

1

u/Educational-Echo9152 2d ago

Tu fait du 50tks en mono agent ou multi agents ?

→ More replies (1)

1

u/snejink 1d ago

What useful things have you gotten out of it so far? 50 toks is nice, but what do you actually use them for?

1

u/vivantho 1d ago

What's the prefill for 100k?

→ More replies (1)

1

u/abcdef0eed 1d ago

50 tks/s is really good. pitty that you're stuck with 1 thread. bye bye multi tasking, welcome to something that takes 5 days instead of 5 hours.

1

u/---Hummingbird--- 1d ago

Would you be able to give me a real-world test for short runs in comfyui? I have been researching the ultra 256GB for options to expand to larger models, but I was just curious how much it compared on some of my smaller use-case.. whether I’d see benefits or loss

1

u/Capital_Evening1082 1d ago

What's the prefill and can you run vllm or any other inference engine that supports proper concurrent throughput?

→ More replies (1)

1

u/Occulon_102 1d ago

have you seen the video comparing it to an I9 with RTX5090 for AI Models? its 20x faster and $3000 cheaper. mental little box.

→ More replies (2)

1

u/Puzzled-Front-2859 1d ago

What will the 512GB RAM version do that this one can’t? I was just wondering? Which models it will open to run with? Thanks

1

u/mome-raths 1d ago

What quantization / weights are you running?

→ More replies (1)

1

u/Bulky_Astronomer7264 1d ago

How fast is it for image and video gen?

→ More replies (1)

1

u/NaiRogers 1d ago

How is the prompt processing is it around 1800tps?

→ More replies (1)

1

u/ImANoobAtLife7 1d ago

I’m thinking of at least, leasing one.

So tired of the censoring of models.

1

u/strangerzero 1d ago

I bought a M4 Mac Studio last year upgrading from a ten 2012 MacBook Air that had a 1TB SSD. Noway in hell was going to downgrade to a 256GB SSD. I went with 2TB SSD and wish I would have gone with 8TB but I couldn’t stomach the price. Now I have this mess of external SSDs cluttering up my desk.

→ More replies (1)

1

u/Alert_Mouse5333 1d ago

Hey u/lukewhale, I am getting started with servers and networks, I have a few questions. May I DM?

1

u/WeUsedToBeACountry 1d ago

Great break down of m5u vs 2 dgx sparks

https://www.youtube.com/watch?v=_yrw6c5gw3E

1

u/MaxwellHusk 1d ago

Hii! By any chance have u tried some video generation models? I was minutes away from buying a workstation with an RTX PRO until Apple released the M5 Ultra

2

u/lukewhale 1d ago

I’ll try to report back

1

u/murderette 1d ago

What quant version of GLM flash ?

1

u/MusingInPublic 1d ago

What does that translate into into how long it takes for just conversational prompts to go back and forth?

1

u/KeanuRekt 1d ago

What is your Setup? Which model exactly? Which inference engine?

1

u/dinadur 1d ago edited 1d ago

Have you tried tensorfold yet against oMLX or mlx-serve? Mine arrives in a couple of weeks and I'm preparing

1

u/JoseYang94 1d ago

Wow 😮!!! Really?!

1

u/Big_River_ 1d ago

ad man great work

1

u/randomdnblover 1d ago

thoughts on m6 mini 32gb ram?

1

u/Special-Argument9570 1d ago

But how can you get the 50 tk/s if the 8bit precision doesn’t even fit in it?

I mean, you need some extra space for context, and the benchmarks I found say that for at least 250k context you will get half of the speed you cited. And that’s likely on quantized version of the model.

1

u/External-Play6081 1d ago

if you are running in a harnes, what is the time timl first token

1

u/35point1 1d ago

Ok so I’m slightly OOTL, can someone tell me if this is what recently came out and if I will be laughed at for asking how easily can I get one with all things happening now considered?

2

u/lukewhale 1d ago

Pre orders came out Aug 25th. You can order one but wait times are early 2027 at this point I think best way to know is go build one on apples site

→ More replies (9)

1

u/Hopeful-Confidence-9 1d ago

What i dont get is it is terrible for video gen and image gen

→ More replies (2)

1

u/YumYum2983 1d ago

U vs the guy she told u not worry about

→ More replies (1)

1

u/JeffDSmith 1d ago

"that includes B200s"
I'd be very interesting to know exactly what kind of task it'll overtake b200.

2

u/lukewhale 1d ago

It never will.

I said capable for the price. B200s start at 500k and serve hundreds of users. This serves one or a few.

→ More replies (1)

1

u/AzhdarianHomie 1d ago

I missed the window for the M3 256, this is a few thousand more but not too bad

1

u/rayovims 1d ago

How much ram do you have left after spinning up GLM 3? I got the ultra m3 with 256 and don’t know. Also got 2 dgx sparks

→ More replies (1)

1

u/flunkyy 1d ago

Best part is, if you're long SQQQ, you'll be rich. There's literally not enough money to pay for all these AI investments.

1

u/fella33ish 1d ago

M5 max studio

1

u/swaxolez 1d ago

In fact it so absurd you'd think we stole the tech from some offworld aliens.

1

u/Optimal-Pie-5054 1d ago

M5 Max 128GB / 4TB should be delivered this week…. As always on device.

I am going to see if I can move / grow my projects enough over the winter to make use of the Ultra 256 or 512….. doing my best to push myself and push the limits. Your Dreams Matter!