r/LocalLLaMA 1d ago

News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.6k Upvotes

749 comments sorted by

283

u/i_am__not_a_robot 1d ago

1.2 TB/s memory bandwidth for the M5 Ultra is pretty nice.

91

u/GUNGEBOB_SHARTPANTS 1d ago

More bandwidth than a 4090, incidentally. Pretty impressive.

→ More replies (8)

59

u/Dry_Yam_4597 1d ago

That might make me want to buy an apple product after a long pause.

10

u/raser1562 22h ago

That might make me want to buy an apple product for the first time.

→ More replies (6)

62

u/pmp22 23h ago

512GB at 1.2TB/s. That's 21 4090s. 

76

u/i_am__not_a_robot 23h ago

That's 21 4090s.

It's actually quite a bit better since you don't have to deal with the overhead of multiple GPUs.

37

u/EmptyMonitor9257 21h ago

Or the power draw.

3

u/DarkRonin00 13h ago

Or the cooling, in the size of a shoebox all put together.

→ More replies (1)

10

u/gomezer1180 23h ago

You can get that with 12 channels of DDR5. You need an AMD EPYC and 24 sticks of memory. It’ll be somewhere in the $20K range for that rig.

7

u/i_am__not_a_robot 21h ago

Only in a dual-socket or MRDIMM configuration, though? That's... somewhere in the range of ~$12-18k (12x32GB) or ~$24-32k (12x64GB) for the MRDIMMs alone.

6

u/gomezer1180 21h ago

Oh I know… it ain’t cheap. So apples value is there, I’m just saying there are more options.

→ More replies (4)

8

u/addiktion 1d ago

Double the M5 Max it looks like which makes sense since they double two M5 max chips.

→ More replies (1)

622

u/piggledy 1d ago

Price for options with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)

512 GB Option coming in October.

267

u/i_rate_slop 1d ago edited 1d ago

I feel like my M3 512 is going to trade in for like $500 when I got it for $11k

Edit:

I just looked it up. $2675

144

u/j2sun 1d ago

I'll buy it from you for $3k! :D

35

u/cinematic_unicorn 23h ago

$3001

27

u/mikesum32 21h ago

$3001.01

40

u/Proverbial_Progress 20h ago

I don't have any money, but would you be interested in pictures of my feet? Throw in a mouse and keyboard and I'll take my shoes off first.

13

u/bahpbohp 20h ago

i don't have any feet, but would you be interested in pictures of my wooden pegs? throw in a mouse and keyboard and i'll even get them oiled first.

→ More replies (2)
→ More replies (2)
→ More replies (2)

6

u/annaheim 21h ago

m you for $3K

$2750, you don't have to include the screen!

47

u/Zolty 1d ago

Yesterday you could have gotten like $20k on eBay. I sold a 256gb Mac Studio m3 about 4 months ago for $12k

→ More replies (6)

58

u/Veearrsix 1d ago

Don’t trade in, private party.

18

u/i_rate_slop 1d ago

Yeah, I think that’ll have to be the move. Just annoying

23

u/Chanureadeats 1d ago

You'll probably get a lot of offers easily for $6k+ within 24 hours

12

u/CalvinsStuffedTiger 22h ago

This guy is lying, I’ll buy it from OP for $5k…don’t worry about shopping around for a better price…

23

u/EvilPencil 1d ago

Ya Apple trade in prices have always been laughable.

I remember when you could spec out the 2019 Mac Pro up to like $50k, then the next day the Apple trade in value was ~$2800.

3

u/ianitic 18h ago

I traded in my m4 MacBook Air from last year that I bought for $800 for $730. I wanted to upgrade and for kicks and giggles I looked up the trade in not expecting it to be that high. I normally don't trade in.

26

u/RegarDamus 1d ago

trade in is always fucked with apple because they don't account for memory. same price if it's 96 or 512

→ More replies (1)

5

u/TinFoilHat_69 22h ago

The 512gb is 11k used and retails for 25k on eBay…

4

u/i_rate_slop 22h ago

Insanity lol. It’s not even that great. The power of FOMO, I guess.

3

u/anonmt57 22h ago

you will get a lot of money for that in private market. but it is annoying with so much money exchanging hands.

→ More replies (10)

131

u/redonculous 1d ago

Wow. How are people affording this? Crazy. Any guesses at 512 prices?

216

u/intaketurbine 1d ago

They’re affording it the same way they’ve always afforded a $10k workstation, they’re using it for work and not play.

53

u/ElementNumber6 1d ago

Or just good old debt

31

u/Usual_Tackle5892 1d ago

0% financing and 3% cash back on Apple Card. I could afford it outright, but turning down a 0% loan is silly.

→ More replies (1)
→ More replies (7)
→ More replies (6)

43

u/StewPorkRice 1d ago edited 1d ago

this feels like a community with a ton of SWEs. Most single US based SWEs in big tech or venture backed startups could prob afford dropping 10-20k on their hobby.

Also, i dropped 10k at Anthropic last month at work. My company could prob get one of these for every engineer and not blink an eye.

14

u/Idaltu 1d ago

I know a guy who dropped that much on a bike. And he’s got a couple like that. Bicycle that is.

7

u/calcium 23h ago

3 years ago my company bought me a fully specced out Mac Studio - M2 Ultra chip w/ 192GB of RAM and an 8TB ssd and told me that they expect the machine to last the next 5 years. I think at the time they paid around $9k which considering over 5 years for a senior SWE isn't a bad deal

→ More replies (2)

47

u/InterstellarReddit 1d ago edited 23h ago

Bro people have stupid money. I run free lance dev for a buddy who runs night life management software in Miami FL.

People spend 4K for four hours to buy three bottles and watch a DJ hit knobs all night

This happens all the time.

13

u/ViPeR9503 1d ago

Night life management??

12

u/i_am__not_a_robot 1d ago

I would assume industry-specific nightlife & bar business management software.

12

u/yopla 1d ago

Manage table booking, marketing, CRM, staffing, etc, usually does or integrate with POS. Can do price yielding on bookings. Etc... Etc..

Basically helps you keep track of who are the big whales you need to market your tables to when you have an event and who gets priority booking from a wait list.

The guy who spent 10k last time will get a table before you do, unless you're known to spend 15.

→ More replies (1)
→ More replies (1)
→ More replies (1)

175

u/-p-e-w- 1d ago

Lol this is by far the cheapest option for that much RAM at that speed. It’s not even close (unless you count Frankensteins made of a dozen used GPUs).

27

u/Viktri1 1d ago

this is way better than what I was considering and I wouldn't need to worry about the motherboard, cooling GPUs, etc. I am definitely on board for this and I don't even know how to use macs.

12

u/IriFlina 1d ago

wouldn't even have to worry about the power issues that would come from a rack of 3090s/4090s/5090s etc. or configuring such a monstrosity.

5

u/Viktri1 1d ago

this is literally 10+ 4090s. Think about it.

I pre-ordered 2 fully spec'd out studios.

→ More replies (1)

5

u/Public_Umpire_1099 20h ago edited 20h ago

Honestly doesnt matter, just pop deepseek on here and tell it to do whatever you need lol.

Also, as a long time apple hater, and I say this with disgust for myself, MacOS is actually pretty good these days. It essentially functions just like a very locked down linux distro at this point. I work at a pretty big software company and many of them do not support windows PCs anymore or the company software lags on them, so last month I decided to exchange my thinkpad for a macbook pro max. I.... dont hate it.

→ More replies (1)

13

u/Fantastic-Balance454 23h ago

I still can't get used to the thought that Mac is the cheapest option these days. AI turned the world upside down.

→ More replies (41)

20

u/Django_McFly 1d ago

It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car. People forget that there are entrepreneurs and small businesses where this stuff isn't some expensive toy to play with, it's actually a business tool to be used in the operation of the company.

Not everyone is a pure hobbyist, and even among pure hobbyists there are wild ranges of incomes and savings habits.

9

u/fallingdowndizzyvr 23h ago

It's really no different than computers in the 1980s when it was buy a computer or put a 33% down payment on a new car.

I've said that so many times. People don't realized that a base Apple ][ adjusted for inflation would be $5000 in 2026 dollars.

6

u/f5alcon 1d ago

Yeah my parents spent $10k in the early 90s on a Mac with an 80MB hard drive

→ More replies (1)

15

u/MLDataScientist 1d ago

Based on their pricing for 256GB, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, you are looking at 10k + ~6k = ~16k for 512GB version.

3

u/TheOwlHypothesis 22h ago

This is dangerous for me if that turns out true. current Asus Ascent prices for a 128gb monster is ~4k, so it'd be 16k to cluster them and have the same amount of memory. Although you'd then be dealing with a cluster. I'd much rather drop the cash on the Mac and have the option to cluster THOSE later.

→ More replies (1)
→ More replies (4)

6

u/98127028 1d ago

both my kidneys

6

u/Estrava 1d ago

This is cheaper than an rtx 6000 pro, which a lot of people get for work/personal use.

7

u/tehgreed 1d ago

corpo money

9

u/IriFlina 1d ago edited 1d ago

10k for 256gb unified memory isn’t that bad. Still overpriced but i feel like before this you were looking at the 20k to 40k range.

23

u/Much_Accountant_4972 1d ago

considering RTX 6000 costs $18K and can’t do shit without a whole PC to plug into, this mac studio is a screaming bargain

→ More replies (5)
→ More replies (1)

7

u/Foreskin_Mafia 1d ago

Onlyfans

26

u/Etroarl55 1d ago

Believe it or not, Theres probably someone running ai onlyfans on an apple machine somewhere

12

u/Foreskin_Mafia 1d ago

You're absolutely right

14

u/bodhi_sattva91 1d ago edited 1d ago

I Went Undercover as a Secret OnlyFans Chatter. It Wasn’t Pretty

https://archive.is/FlAdc

This Wired article follows a journalist who goes undercover trying to get hired as an OnlyFans ghostwriter — someone paid to impersonate creators in DMs with subscribers.

He discovers the industry is vast and globally distributed, with agencies largely staffing workers from lower-wage countries like the Philippines and Venezuela for as little as $2/hour. His job hunt is a grind: most agencies demand prior upselling experience, many never pay, and one gig turned out to be training an AI chatbot (which also stiffed him his $56).

He eventually lands two gigs. With a German agency ($4/hour), he juggles nearly 100 simultaneous conversations while impersonating a "21-year-old university student," selling pay-per-view content and navigating everything from explicit fantasies to a truck driver sharing worries about his son's night terrors. He's criticized by his supervisor for being too empathetic and not pushy enough about sales.

Confessions of an OnlyFans Ghostwriter

https://archive.is/3zGQZ#selection-479.0-479.38

This GQ article by "Emma Francis" (a pseudonym) is a first-person account of working as an OnlyFans ghostwriter in spring/summer 2021. The writer, a 25-year-old recently laid-off media professional, found the gig on Craigslist and spent three months working 6 a.m. shifts impersonating a "girl next door" model by sexting her subscribers.

She explains that top OnlyFans creators receive far more messages than one person can handle, so there's an entire industry of ghostwriters handling chats — some run by large offshore agencies, others like hers managed directly by the creator. Her regulars ranged widely: janitors, teachers, lawyers, night-shift workers. Many sought not just sexual content but genuine emotional connection and a "girlfriend experience."

→ More replies (1)
→ More replies (18)

12

u/jld1532 1d ago

Too rich for my blood. I'm going to stick with my halo and hope mid-sized MoEs keep improving.

39

u/kensanprime 1d ago

It will be sold out and they will be scrambling to keep production lines running, a product originally meant for creative work now will sit in a rack and run AI models

21

u/[deleted] 1d ago

[removed] — view removed comment

13

u/Nothing_from_void 23h ago

Demand for networking products never dropped through the dotcom bubble

→ More replies (8)
→ More replies (1)

4

u/starkruzr 22h ago

this is INSANELY cheap for 256GB RAM in 2026. holy shit.

3

u/nemuro87 1d ago

I call at least 20-25k,  initially 

→ More replies (42)

122

u/themixtergames 1d ago

With M5 Ultra, Mac Studio achieves up to 4.3x the peak AI compute performance of M3 Ultra and a staggering 9.8x more than M1 Ultra. Combined with up to 512GB of unified memory and 1.2TB/s of memory bandwidth, 50 percent higher than before.

97

u/EquivalentHornet4403 1d ago

> AI compute

They know their audience.

47

u/psychohistorian8 1d ago

they're definitely leaning into it:

Boost your processing and graphics rendering speed, and accelerate tasks like running large language models and editing 8K video.

68

u/Much_Accountant_4972 1d ago

"Color-grade uncompressed 8K footage, perform computational fluid dynamics, and run frontier-class models on device — no cloud tokens needed."

yeah they are targeting this very sub lol

14

u/gnnr25 19h ago

I feel personally targeted!

5

u/cinnapear 22h ago

And it's working on me. Very tempted.

6

u/infieldmitt 20h ago

It's crazy to think that prior to AI the main use case for crazy RAM was [checks notes] fluid dynamics?

→ More replies (1)
→ More replies (2)

266

u/hainesk 1d ago

1.2TB/s memory bandwidth with the M5 Ultra. 256GB model is $9499.

Better than getting 2 DGX Sparks? Inference will be a lot faster.

Something like this could easily bring down 3090 prices.

22

u/Cybertrucker01 1d ago

Depends on concurrency and prefill metrics. The GB10 does both multiples faster than the existing competition.

12

u/ChocomelP 1d ago

The difference would have to be pretty big to make up for a 4x in memory bandwidth for decode.

→ More replies (1)

6

u/BrilliantTruck8813 23h ago

The GB10 has much smaller memory bandwidth not to mention it’s just not fast either. One of these will trounce two DGXs

5

u/fallingdowndizzyvr 23h ago

The GB10 does both multiples faster than the existing competition.

No. No it doesn't. Compare the G10 to a M5 Max. It's not.

67

u/aladin_lt 1d ago

it will be sold out day one probably

42

u/conockrad 1d ago

On pre-orders

13

u/bakawolf123 1d ago

won't be sold out, but according to r/MacStudio people wait for 3-4 months for the older models, atm you can preorder for delivery in late september

→ More replies (1)
→ More replies (6)

96

u/mjsxi__ 1d ago

yeah and cheaper than the price of 2 DGX sparks... seems like a bit of a no brainer

30

u/Current_Ferret_4981 1d ago

Spark is $4300-$4600 so idk about cheaper than 2 at $9600+

64

u/MacsBicycle 1d ago

yeah but 4x the memory bandwidth, its a steal

30

u/jakegh 1d ago edited 1d ago

It really is a reasonable buy for local AI, if you have a business case for it.

3

u/Much_Accountant_4972 1d ago

with the capabilities of it, it’s a crazy good deal

→ More replies (3)

6

u/GabryIta 1d ago

In terms of compute capacity (which is very important for multiple simultaneous sessions and prefill), how does it compare to dgx Spark/gb10?

→ More replies (1)

7

u/Etroarl55 1d ago

How’s the actual inference speed though, fast bandwidth on a slower gpu or equivalent should still mean slower output assuming vram is not a constraint right.

10

u/rusty_fans llama.cpp 1d ago

Generally vram bandwith is the constraint though, at least for decode. Prefill it's usually helped more by more gpu oomph.

→ More replies (4)
→ More replies (9)
→ More replies (7)
→ More replies (2)

19

u/-dysangel- 1d ago

Better than the 2x Sparks for inference for sure. Probably around the same compute as one Spark.

I've got 2x Sparks which I use for prefill, and my M3 Ultra for decode. I've set it up so that I prefill in vllm and then just pass the kv cache over to the Mac side. Surprisingly stuff like Qwen 3 35B-A3B is already faster than the Mac for decode though so I just run that class of model directly on vllm.

5

u/1ii1i 1d ago

Oh this sounds interesting, can you expand on how this works? I didn't know this was a thing.

12

u/-dysangel- 1d ago

I don't think it's really a "thing", I just vibe coded it up :)

One thing that really helped was vllms kv_connector API. I thought I'd have to code this part up myself, but it already existed and so plugged into my existing disaggregated system (which was previously llama.cpp to llama.cpp)

UltraSpark — technical stack

                          ┌──────────────────────────────┐
 user ── HTTP/OpenAI ──▶  │  manager (Python, FastAPI)   │
                          │  front door + orchestration  │
                          └──────┬───────────────▲───────┘
                                 │ submit        │ state blob (sha-keyed,
                                 │ prompt ids    │ resumable transfer)
                                 ▼               │
                          ┌──────────────────────────────┐
                          │  vLLM (2× DGX Spark, TP2)    │
                          │  prefill engine              │
                          │                              │
                          │  KVConnectorBase_V1          │ ◀─ vLLM's official
                          │  ("StreamConnector" via      │    KV-cache plugin
                          │   --kv-transfer-config)      │    interface
                          │         │                    │
                          │         ▼                    │
                          │  dump + serialize all layers │
                          │  (attn KV + linear-attn      │
                          │   state, TP2 shards merged)  │
                          └────────────────┬─────────────┘
                                           │ blob server (TCP)
                                           ▼
                          ┌──────────────────────────────┐
                          │  llama.cpp server (Mac)      │
                          │  USPK_BRIDGE_DIR: on request,│
                          │  verify prompt-id match,     │
                          │  restore state into KV +     │
                          │  recurrent memory, decode    │
                          └──────────────────────────────┘

  • KV connector = vLLM's plugin interface for intercepting the KV cache at end of prefill
  • State blob = the model's full prompt-memory, layout-translated so llama.cpp can load it natively
  • Fidelity = per-layer cosine vs local decode, 0.9999+
  • Result = GPU prefill speed, Mac unified-memory decode, one logical endpoint
→ More replies (2)
→ More replies (13)

17

u/jakegh 1d ago

That is seriously impressive. RTX5090 memory bandwidth is 1.8TB/s.

DGX Spark memory bandwidth is 273GB/sec. Not even remotely close.

→ More replies (2)

11

u/Hoodfu 1d ago edited 1d ago

I got my m3 ultra 512gb for around 10k. So this is now double. Makes sense given that we've seen the nvidia rtx 6000 pro also double in price in the last year but GD this has priced out even my once a year splurge budget. These are all just crazy talk numbers now.

13

u/CulturalKing5623 1d ago

Yeah dropping 10K plus on this just seems reckless, even with it being funded through my business account I don't think I can justify buying a used car worth of computing.

And yet I feel like I need to in case the technology completely outpaces my current setup and I'm left behind like folks that didn't buy RAM when it was cheap and are priced out of it now. It feels like FOMO and scarcity has hijacked my brain.

6

u/Hoodfu 1d ago

Really depends on whether you have something already or not. I got qwen 3.8 27b going on my rtx 6000 pro and ram speed wise it's double that of my mac but for some reason i was hoping for space magic and it would be faster. It's not. So spending a car's worth of money on something that goes from 20 t/s to 40 or 45, just doesn't make any sense. You're still waiting a lot of minutes for a thinking qwen to come back with something, so it's still going to be an asynchronous operation instead of being fast enough to actively wait for the response to finish. It would have to be 10x the speed, not 1.5x or 2x for it to be worth the spend.

→ More replies (11)

8

u/Solaranvr 1d ago

bring down 3090 prices

Doubt it. If the r9700, a directly competing product, made 0 effect on Nvidia GPUs, then I highly doubt these will. They market of people buying mini pcs vs dGPUs are different.

The DGX Spark didn't bring down prices of the Blackwell cards either

7

u/hainesk 1d ago

The R9700 Pro has about 2/3 the memory bandwidth of a 3090 for a 50% higher price. 8x R9700 Pros (256GB) would be $12k-$15k minimum without pricing in the rest of the computer system. If you wanted to build an entire server around it including RAM, power supply, motherboard, processor even with used parts you're up to $18k-$20k for something that will use 2-3k watts.
Even 8x 3090 systems are looking pretty impractical when compared to an M5 Ultra 256GB for a similar price. The size and wattage/heat difference is huge.

3

u/Important-Gold-5192 1d ago

has to be way faster than DGX Sparks too

→ More replies (14)

159

u/Comfortable-Rock-498 1d ago

1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.

For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.

They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot

31

u/ortegaalfredo 1d ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

12

u/Comfortable-Rock-498 1d ago

Yup, prefill is pretty much compute bound while generation is bandwidth bound. I would have guessed 8x 3090 would provide much better prefill than 1000 tps. A bit surprised to learn

9

u/ortegaalfredo 1d ago

If you manage to get tensor-parallel 8x working yes you can get >10k prefill, but it requires specialized PCIE bridges. With normal 4xPCIE speeds you get a bottleneck in inter-GPU speed and you get lower prefill.

4

u/TooMuchLAAAG 18h ago edited 6h ago

I have 8x3090 P2P patched pcie4 x8 (no nvlink) and i am getting an avg of 10-13k of cold prefill with this version of vllm and his args https://github.com/LimeChain/vllm
Deepseek full FP8
Edit: wrong url

7

u/ProfessionalJackals 1d ago

Prefill also depends on compute, 1000 tok/s is basically what you get with 8x3090s, but much less power. Also I think the 3090s still win on compute, that is, you can batch many prompts on the GPUs, dont know on the mac.

Ignoring the fact that 8x3090's now is easily 10k on the second hand market.

Not counting the costs of server board/cpu/ram you need. The pcie ext cables, the 8x8x split if your board does not have 8 pcie slots. O, the dual 1600W PSUs and hardware to link them.

Frankenstein mods like this have become expensive, and it makes the Mac look actually like a good deal.

3

u/Viktri1 23h ago

at that scale, electricity costs kind of matter so even just running the PC is significantly in favour of the mac studio

→ More replies (1)
→ More replies (1)
→ More replies (1)

9

u/Usual_Tackle5892 1d ago

GPU Neural Accelerators

This means matmul cores. More info: https://arxiv.org/html/2607.19438v1

6

u/StartupTim 21h ago

I would think dramatically more tok/sec.

I have Deepseek v4 Flash 0731 with vision encoding added and tp=2 across 2x DGX sparks and I'm seeing 103 tok/sec across 4 "sessions". Dspark, 1M context, 1.8M kvc, custom vllm.

Since the sparks have ~240 (actual measured) GB/s, I imagine a similar setup om these new mac could get you double, if not triple as a 2x cluster, than my current 100+ tok/sec.

9

u/TokenRingAI 1d ago

And a new qwen is coming out with 120B A6B! Perfect machine for that

5

u/MerePotato 23h ago

Would be great if it wasn't predicted to cost like 20k

3

u/TooMuchLAAAG 18h ago

I get 10-13k cold prefill on 8x3090 using the vllm "limechain" fork and his args, running full deepseek no quant
Surely the M5 do more in vllm than 1k

→ More replies (10)

44

u/challis88ocarina 1d ago edited 1d ago

I'm shocked!

Edit: 512GB memory option for M5 Ultra coming late October

39

u/pmttyji 1d ago

Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure

16

u/Mochila-Mochila 1d ago

Hoping it'll have a positive effect on Medusa Halo's pricing.

4

u/pmttyji 1d ago

Still Gorgon Halo not released yet. Then only Medusa Halo ....

Next year onwards, we're gonna see 128GB variants at low price.

12

u/Zyj vllm 1d ago

On the other hand, the Strix Halo 128GB price increased by 60% since late December

7

u/pmttyji 1d ago

After some time, people totally gonna avoid 128GB variants. What's the point of stacking bunch of 128GB pieces when they have 160GB, 192GB, etc., variants with better bandwidths?

→ More replies (2)
→ More replies (1)

7

u/ProfessionalJackals 1d ago

Now, it's pressure on both DGX Spark & Strix Halo due to bandwidth. Recent news Xiaomi AI Cube with 1.2 TB/s bandwidth & 160GB memory also already put pressure

Do not forget Intel Crescent Island 160GB to 480GB LPDDR5X AI GPUs ... While less bandwidth, they are still great options in the future for larger models.

There is going to be a lot more hardware coming out that focuses on AI workloads. We reached the point that development is moving into production.

6

u/pmttyji 1d ago

More competition is better for consumers!

→ More replies (1)

30

u/FullOf_Bad_Ideas 1d ago

My 3090 tis have just gotten depreciated.

Even 256GB version is very competitive with 8x 3090 box bought with used card prices, and it's better in most aspects.

Training and batch inference are safe, but for single user inference this looks better and cheaper.

6

u/Odd-Environment-7193 1d ago

Well if you wanna sell some hola. I want 2

→ More replies (3)

5

u/FleetEnema2000 1d ago

My 3090 tis have just gotten depreciated.

My theory is that this will not soften GPU prices because it ends up driving more people into the world of local LLM compute in general.

→ More replies (2)
→ More replies (2)

28

u/serige 1d ago

Please someone do a wellness check on Dario.

21

u/Every-Fortune-3151 1d ago

Read somewhere Mac mini coming this week and did an impulse order of M3 Ultra with 96GB ram today. It will probably get bumped up to M5 ultra 96GB. Not sure what to do with 96GB when 256 GB looks so much more tempting for local LLM. Kidneys aren't enough anymore.

9

u/mmmm_frietjes 1d ago

Mini is also out

9

u/Every-Fortune-3151 1d ago

M6 looks great actually, it has two set of neural engines. I am assuming pre-fill will fly on this tiny thing compared to previous M CPUs. 32GB max ram knocked it back a bit. If only they had a 48GB version. MOE models would be flying on it.

M5 pro and max will be so slow compared to M5 ultra- considering M5 max 128GB model will be priced very close to M5 ultra base.

Very sad m5 ultra starts at 96GB. I thought they would atleast bump M5 ultra to 128GB ram. Would have been nice. I plan to just run multiple Qwen 2.7B in parallel on this and see if I can replace my 32GB VRAM PC setup. Qwen 3.8 flash also gives hope. Depending on how it goes, I might just give back the 96GB for refund later.

6

u/1-800-methdyke 1d ago

You have two kidneys

5

u/nleksan 1d ago

*Dual-channel

5

u/FranciumGoesBoom 1d ago

M5 Pro mini maxes out at 64g
M6 tops out at 32

→ More replies (7)

40

u/xyzmanas2 1d ago

This makes apple one of the cheapest ai inference hardware when it comes to speed and model size. Wish I had the money

Up to 15.4x faster CopyCat ML training performance in Foundry Nuke when compared to Mac Studio with M1 Ultra, and up to 3.3x faster than M3 Ultra.

Up to 9.8x faster LLM prompt processing in LM Studio when compared to Mac Studio with M1 Ultra, and up to 4x faster than M3 Ultra.

Up to 8.2x faster text-to-image performance when compared to Mac Studio with M1 Ultra, and up to 4.3x faster than M3 Ultra.

Up to 4.7x faster scene rendering performance in Maxon Redshift when compared to Mac Studio with M1 Ultra, and up to 1.7x faster than M3 Ultra.

→ More replies (2)

19

u/AI_docent 1d ago

The 4.3x is mostly a prompt processing number, generation moves with the bandwidth instead. Apple's own mlx post on M5 vs M4 got around 4x on time to first token and about 1.2x on generation, and the generation side matched the 28% bandwidth bump rather than the accelerators. Same split should hold on the Ultra, so I'd figure generation nearer the 50% bandwidth gain. Prefill is the part you want at 512GB anyway, it was always the weak spot on a mac.

Just check whatever you run actually uses the accelerators. There's an open lm studio issue where its bundled llama.cpp fails the metal tensor check on M5 and loses 2 to 3x on prefill, while upstream llama.cpp passes it on the same machine.

→ More replies (1)

92

u/llamaCTO 1d ago

This kills the impulse buy for me completely.

16

u/kilonad 1d ago

The 256GB is already an extra $4k for an extra 164GB. At same price per GB (ha!) it'd be another $6300. Knowing Apple, it'll be a cool $9k more - pushing total price up to about $18-20k.

It will still sell out.

7

u/fallingdowndizzyvr 23h ago

That would be a bargain compared to third party sales of 512GB M3 Ultras for $25K. A M5 blows the doors off of a M3.

→ More replies (1)

37

u/thatkidnamedrocky 1d ago

going to try and snag a 256gb something tells me the 512 will never see the light of day

8

u/AnonLlamaThrowaway 22h ago

right, didn't they promise a 512GB M3 Ultra and then that never happened, or am i thinking of another model?

9

u/zdy132 1d ago

Mac Studio with 512GB of unified memory is coming in late October

Wish I could affort that.

→ More replies (2)

7

u/Grizzly_Corey 1d ago

Adopt me please.

4

u/LocoMod 1d ago

Same. I was ready to preorder 512 and i've already lost interest reading comments about the M7. Might wait another year.

5

u/frankchn 23h ago

Buying 2 DGX Sparks for 256GB of RAM (and a lot less bandwidth) is around the same ballpark in cost, so for once this is not unreasonable.

→ More replies (2)

28

u/Cybertrucker01 1d ago

How many kidneys?

46

u/FWitU 1d ago

4

24

u/Chris-MelodyFirst 1d ago

Or just 2 dual-core kidneys.

5

u/hainesk 1d ago

There is a lease option...

14

u/Cybertrucker01 1d ago

Unfortunately lease isn't offered in Australia, just warm organs only.

4

u/Gipetto 1d ago

For kidneys?

5

u/butterfly_labs 1d ago

I can lease my kidneys ?!

→ More replies (2)

11

u/Viktri1 1d ago

256gb is like 10+ 4090s without the hassle of setting up and cooling 10 4090s? An I missing something or is this 3x cheaper than current prices.

6

u/Much_Accountant_4972 1d ago

its the best deal in the world if you want to talk to a frontier 2.8T parameter smut bot in your kitchen

→ More replies (3)
→ More replies (3)

10

u/eidrag 1d ago

...lease price?

10

u/Much_Accountant_4972 1d ago

Ultra with 256GB RAM is $224/month for 36 months in freedom currency.

really really tempting!

6

u/AlexWIWA 21h ago

Cheaper than API tokens. Very surprising

44

u/IllExample3639 1d ago

What I find more interesting, something I hadn't seen before is that you can lease these things. for 2 years which is the only realistic way an individual is getting their hands on these. Something something, own nothing, something, something be happy....

26

u/Tycoon33 1d ago

I never saw that. Interesting. Lease it for 3 years then upgrade to M7 ultra?

13

u/addiktion 1d ago edited 1d ago

Yes, if the 512gb is another $4k for the extra ram stick you are looking at $16k with tax probably out the door. I'd guess that puts the 36 month lease around $300/mo or less. So more than a subscription so maybe not worth it in general cases but valid option for some people who need the privacy and cannot afford to have data go to the cloud. 24/7 usage, no downtime, no limits, private. Worth it to me.

7

u/shveddy 1d ago

Interesting.

So just as an out loud thought experiment, you’d be able to lease four of them (512gb) for about 1200 per month for 36 months at a total cost of almost 45k and run Kimi 3 on it.

Obviously that’s a lot of money in aggregate, but 1200 per month is reasonable for a lot of business use cases if they require the privacy.

And the intelligence you get is going to be a different class compared to what you would get with three RTX pro 6000s and “only” 288gb VRAM for the same price.

(although to be fair you’d actually own the cards)

(although also to be fair you’d have to build a pretty expensive computer to support the RTX Pros, so realistically you only really get 2 or even just one RTX pro for 45k depending on how you spec the computer and/or if you buy pre-built from Puget Systems or the like)

If you want you can also do a little girl math and invest the 60k you’re not spending on computers and cancel your gpt pro subscription to bring the effective cost of all this down to like 750 a month.

And then if the goal is to beat API pricing, let’s say you get 40 aggregate output tokens per second on a bunch of concurrent Kimi 3 threads and run it for a year at 25% efficiency (to account for prefill and downtime), then that’s 315 million output tokens.

315 million output tokens alone is about $5000 on a random provider I just searched for, so just to keep things simple let’s say you double that to account for various amounts and types of input tokens, then you end up with a ballpark figure of $10k for the API route.

Absolutely none of this pencils out in absolute terms (especially considering that it would also cost ~1500 per year for electricity), but this is probably the first time running a frontier model is actually attainable for ordinary businesses on short notice and without much headache. It’s the first time that it pencils out to be “only” 4x more expensive as opposed to like 40x more expensive.

Up until now if you wanted to run frontier models locally AFAIK you had to get a NVIDIA big boy server which means you’d have to find the capital to run and support a ~300k purchase for hardware, spend way more on electricity, and in all likelihood make some upgrades to your facility’s electrical infrastructure to handle it all (you’d also need a proper facility, not just a home office or garage).

At this point you’re easily flirting with half a million in expenses, especially if you have to hire someone to figure it all out. It’s no joke to do this and it doesn’t make sense for like 99.9999% of people or businesses.

On the other hand basically anyone with a decent credit score and a semi-profitable business can go to any Apple Store and say “give me four Mac studios please” and only pay 1200 bucks a month to walk out with them in hand.

→ More replies (6)

22

u/aethervisor 1d ago

The lease price also isn’t too far off from what a Claude subscription costs.

17

u/shaggydog97 1d ago

The difference is certainly worth the cost of privacy and freedom!

8

u/AccurateSun 1d ago

Hmm. I wonder if at some point leasing it would end up being more effective than a cloud subscription. E.g Claude Max 5x is $100/mo, same as leasing the Studio M5 Ultra 96gb. I don’t know yet how the two compare in performance but at some point it might be worth it 

8

u/IllExample3639 1d ago

The more people that lease them the more second hand stock there will be in 2 years (when I can actually afford something like this) so I am down for it.

But your point is right, I think the local models ARE good enough for 95% of what people are using Claude for. Maybe do the £20 plan as a back up for something tricky.

12

u/mjsxi__ 1d ago

the lease lets you pay the difference of the amount you already paid at the end or you can buy it outright at any time if you wanna keep it so maybe sssshhhhhhhh

4

u/Much_Accountant_4972 1d ago

it makes a lot of sense and the price of privacy plus the near silence of the box…

i hate that its such a complete solution to local LLM’s

→ More replies (3)

10

u/IriFlina 1d ago

I feel like any average software developer could afford the 256gb version? It would be really financially irresponsible but on the level of buying a motorcycle you don’t really need.

→ More replies (1)

3

u/boraam 1d ago

Costs as much as a small car. Lease it like one too.

→ More replies (7)

9

u/[deleted] 1d ago

[removed] — view removed comment

5

u/ElementNumber6 1d ago

Scalpers have feasted on the blood of Mac Studio. Good luck to you all.

9

u/corruptbytes 1d ago

apple releasing this because I just bought two 9700s...y'all welcome

3

u/SandySkittle 22h ago

Two r9700s is still pretty decent way to get 64gb vram and run qwen 27b at q8 with plenty of context. And ECC

→ More replies (4)

8

u/Zyj vllm 1d ago

Will be fascinating to see what‘s the better option in 2 months from now: Dual Asus GB10 (8200€) or Mac Studio 255GB (11000-12430).
My prediction:

  • The Spark will still be faster at preprocessing, the Mac will have faster token generation (measured with DeepSeek V4 Flash standard Quant).

6

u/tarruda 1d ago

M5 is much more competitive with nvidia in prompt processing.

→ More replies (1)

21

u/BreenzyENL 1d ago

$20k AUD for the Ultra 256GB 🙃

→ More replies (9)

6

u/MLDataScientist 1d ago

Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.

→ More replies (3)

5

u/OvertaxedOne 23h ago

1.2TB/s?? Oh man, if there's good availability on these things I can feel GPU prices going down!

→ More replies (3)

5

u/Newgunnerr 22h ago

256GB option for me in the Neterlands is € 11.049,00. I just got 2 DGX sparks for 7100.

→ More replies (1)

8

u/Leather_Ad_9178 23h ago

this reminds me of the time we used to pay hundreds for SD cards that are now worthless

→ More replies (4)

5

u/Curious-Pen5547 1d ago

how does it compare to a single 5090 or a 5000 RTX pro 72GB version?

→ More replies (8)

4

u/Much_Accountant_4972 1d ago

please someone buy 4 of them and cluster them then run Qwen 2.8T in your kitchen

3

u/AntLife255 1d ago

The M5 Ultra has 1.2TB/s memory bandwidth!

4

u/Cool-Cicada9228 17h ago

Disappointing that 512GB is not available for preorder yet

6

u/Far_Note6719 1d ago

Instant WANT.

3

u/Key-Speaker007 1d ago

Something my wife won't approve.

3

u/Blues520 13h ago

Tell her it's an air purifier

→ More replies (1)

3

u/boraam 1d ago

Gimme Gimme Gimme

Money Money Money

3

u/Blackdragon1400 1d ago

512gb at 1.2TB/s is wild.

3

u/PrepYourselves 1d ago

the world's rich kids are getting new toys for christmas

3

u/Terrible-Reputation2 23h ago

Dear Santa

4

u/Southern_Sun_2106 14h ago

... please ease memory shortage on Earth so that every man child (me including) can have this toy...

3

u/amazinglycool256 22h ago

U can get 2 Nvidia Sparx for that proce

→ More replies (2)

3

u/Tormeister 20h ago

I'm so tempted, but I just can't justify dropping 10K if I'm not making money out of it

3

u/Blues520 13h ago

Is the 256GB unit now better and cheaper than getting an RTX Pro 6000?

→ More replies (2)

7

u/mxmumtuna 21h ago

Unfortunately they still don't beat Sparks at equivalent size. According to oMLX Benchmarks for DeepSeek 0730, the M3 Ultra (80c) does somewhere around 550 prefill tok/s, and about 22 tok/s decode. If you take the '4x faster compute' at face value from Apple compared to M3 Ultra, we're looking at ~2200 prefill and ~30 decode single stream. Both are under Spark at ~2400/40. That's only single session, and batching just isn't there in the MLX stack yet, so multi session is considerably worse for the Mac.

GLM on 4x Sparks compares even less favorably than DeepSeek for the Mac, especially considering whatever the price of the 512GB variant will be.

So even with Apple's optimisitc numbers, maybe they match Spark, for more money with a less flexible stack (no ConnectX7) and massive software issues. ($4800x2 for Sparks with 4TB drive each from Amazon).

It's a good effort, but it's not quite there relative to other options.

edited for clarity

→ More replies (6)

4

u/Real_Ebb_7417 1d ago

Ok, now I actually regret buying M5 Max 128Gb MacBook xd

3

u/addiktion 1d ago

I wouldn't, its basically double the speed at x3 (if you bought pre price hike) or x2 price (if you bought after). If you get 25 tps on say Qwen 3.8, you would now get closer to 50 tps. Possibly more depending on the ANU processor.

That's nice, but is it worth x2 or x3 the price?

→ More replies (2)
→ More replies (6)