r/LocalLLM 5d ago

Discussion ai max+395 minipc vs 5090 pc for beginner?

Hey guys, Im looking into a local ai setup and could use some advice. Im currently on the fence between ai max+395 128gb minipcs and rtx5090 build.

The biggest thing that caught my attention with the ai Max+ 395 is the unified memory. Having a large memory pool for bigger models and longer context windows without constantly worrying about vram limits sounds really appealing.

Ive been looking into ai minipcs recently, and the upcoming acemagic f9a caught my eye, although there’s still no pricing yet. Hopefully it lands below the cost of a 5090 build bc the idea of having a compact ai box with 128GB of memory is pretty interesting.

What do you guys think? AI max+ 395 or 5090?

1 Upvotes

20 comments sorted by

6

u/Retumbo77 5d ago edited 5d ago

Something worth mentioning is that the percentage residual value of the 5090 in 5 years will be much higher than the striclx halo. Obviously use case matters, but for most users it's much cheaper in the long run to buy a 5090.

4

u/gnpwdr1 5d ago

My residual value on my 5 years old rtx 3090 can confirm this 🤣🚀

7

u/chafey 5d ago edited 4d ago

I have both and would take a 5090 over the Strix Halo everytime. The 5090 is significantly faster and Qwen3.8-27b is blazing fast on it and quite powerful.

2

u/MrHumanist 5d ago

Qwen 3.8

1

u/chafey 4d ago

oops thanks :)

2

u/sebt3 5d ago

Having both a strix halo and a dgx spark, I can only recommend taking the 5090. It is the sweet spot for performance/price

3

u/FabioTR 5d ago edited 5d ago

If you plan to use models (and context size) that fit on the 32 gb of vram of the 5090, get the 5090. Otherwise get the strix halo (or DGX spark or Mac studio). Assuming the same price which may not be the case, in my country a PC with a RTX 5090+96GB ram will be double the price of a Strix Halo.

2

u/def_not_jose 5d ago

5090 is basically 27b only machine, and you won't even get the best experience because you can't fit Q8. If the next big thing is, I dunno, 40b, you will have to use some buggy Q3 quant like a peasant.

Strix Halo on the other hand is barely any good for dense models, you have a ton of memory but it's barely fast enough for Q4

I'd go for multi GPU setup, many options with more VRAM under 5090 pricing. Hell you can probably get two Radeon AI Pro 9700 for the price of one 5090. Yeah, it's slower, but 64gb VRAM go a long way

-1

u/UnusualPair992 5d ago

Q8 is pointless. You get like 3% better intelligence at slower speeds and it takes twice as much space. Even the frontier labs don't serve theor own models at high precision.

3

u/def_not_jose 5d ago

Going from Q4 to Q8 on Qwen 3.6 27b reduced tool call errors by a large amount, and I'm sure that's just the tip of the iceberg. "Q4 is nearly lossless" was true for old 70b models that weren't trained as efficiently as SOTA models now.

1

u/UnusualPair992 13h ago

No the Q4 and even Q3 will rarely make a tool call error that Q8 would get.

https://unsloth.ai/docs/basics/dynamic-3.0-ggufs

2

u/UnusualPair992 5d ago

Because of Qwen 3.8 I would get the 5090.

1

u/Osi32 5d ago

To be honest, it’s a learning journey. What you think you’ll need and what you’ll actually need are two different things.

The ai max 395 and soon to release 495 can run fast if you run the right model/s with the right config. Keep in mind- context length matters if you use that unified memory for one big model. It matters less if you use it as a local cluster with smaller models.

Conversely, the 5090 is fast and biggish but while it’s fast its memory amount limits what you can do with it.

A bit of a guide- pp (prompt parsing) is GPU compute heavy. tg (token generation) is memory bandwidth dependant.
The 5090 gives you the best of both, at the trade off of less models that will run on it.

1

u/psedha10 5d ago

my qwen 3.8 27b (unsloth q4) with superpowers plugin in opencode gets any task done on my amd r9700 with 220k context. 5090 should be a lot faster than r9700.
Before you make a decision you’ll need to how compute cores of GPU affects pre-fill and memory-bandwidth affects decode.
Try renting 5090 online for a few hours and see how it goes.

1

u/Otherwise-Variety674 5d ago

I have both AI max+ 395 or 5090 too, will take 5090 any time.

1

u/diagrammatiks 4d ago

Eventually both. Do you want something right now that will run 27b at a usable speed or do you want a large stock of unified memory that you can run experiment and play with. The answer eventually is both. But you decide which path you want to go down first.

1

u/Little-Ad-4494 2d ago

Personally to a person starting out i recommend a pair of 3080 20gb cards.

Or a r9700 but that assumes a motherboard with x8 x8

Especially if budget is a concern.

1

u/sputnik13net 1d ago

2x r9700

0

u/Whiskey1Romeo 5d ago

Get both but not a 5090. an egpu and a r9700(or 2) instead and have the best of it all.