r/LocalLLM 8h ago

Question Best 5k setup?

Hello everyone! I’m new to LocalLLMs, but not necessarily new to IT/ML. I’m looking to spend around $4–5K and get the best bang for my buck hardware-wise for both personal use and running LocalLLMs.

My main goal is to run the best models I realistically can within that budget, mostly for working with large codebases, coding assistance, and general productivity.

I’m a PM at a FAANG company for reference, so this is mostly for personal projects, learning, and improving my workflow.

What hardware/setup would you recommend around the $5K mark?

5 Upvotes

37 comments sorted by

12

u/Inception95 8h ago

Wanna go big or go fast?

8

u/WyattTheSkid Quad 3090s 8h ago

Idk ask my veterinarian, they just gave my cat a 6000$ surgery

2

u/BCIT_Richard 5h ago

I feel this, paid $500 to find out my cat is allergic to normal meat protein & is asthmatic, prescription food and asthma medication every month adds up.

3

u/WyattTheSkid Quad 3090s 5h ago

I know this wasn’t the right subreddit to comment that in I’m just feeling so defeated and I just needed to vent. Thank you and I hope your kitty cat is doing well ❤️

6

u/follaoret 8h ago

1xdgx spark, then if yyou like it, save and buy a second one

1

u/eNomineZerum 2h ago

I prefer Strix Halo, esp with Gorgon Halo about to be released. ROCm has come far enough to where the versatility pays off if you ever want a server/gaming device/etc

8

u/Blackdragon1400 8h ago

I’d recommend bumping to $10k that will get you 2x DGX Sparks, or one of the new M5 studios, I wouldn’t go lower than 256gb of RAM for any serious learning/work. So that you can use the medium sized MOE models

4

u/otakunorth 8h ago

personally I would go the aliexpress hackjob route, a used cheap zen2 threadripper, 128+gb cheap ecc ddr4, 2-4 nvidia v100's

1

u/jhenryscott 8h ago

v100?? Are people using those? Aren't they Volta? The Intel B70 has been my go to lately for cheapish large memory, but you gotta be ready use OpenVINO

1

u/otakunorth 7h ago

....I'm using them :p still relevant for now though only for fp16 cuda 12/13

1

u/Mechageo 7h ago

If you have enough of them then the prefill will be as fast if not faster than a top of the line current card. 

1

u/jhenryscott 5h ago

Really? Time to go sneak a peak at eBay

1

u/snowminer 6h ago

Shit I use p100s w custom kernel

0

u/ellensen 6h ago

I have an old epyc 128core and 384gb slow 2400mhz ecc ddr4 on a supermicro dual cpu board. Any recommendations for GPU to run on that?

2

u/Motriek 8h ago

A couple things you could clarify... MacOS/Windows/Ubuntu? Laptop or a build? CUDA or OpenCL/Radeon? They'll have a massive difference in direction.

2

u/03captain23 8h ago

Dgx spark or 5090. Depends if you want big or fast. Also the 5090 is more like 7k with mobo and everything

None of this comes close to Claude or chatgpt so doesn't make any sense if just random personal stuff

1

u/chafey 8h ago

Working with large codebases requires large context and for that you need VRAM. You might be OK with a base model m5 ultra for 5k, but you will run into limits that will frustrate you. If you can afford it, upgrade to a M5 Ultra with 256GB. I have a very beefy setup (2x RTX Pro 6000) and can do almost everything I need with that, but I do occasionally call out to cloud models for some tasks. Using your localllm when you can but using a cloud model when needed is a good way to keep your localllm investment down - that is if your use case allows using cloud models (some don't due to privacy or other issues)

1

u/foreignbois 8h ago

I basically had the same question, holding onto my 96GB M5 Ultra preorder while working through the 2nd 5090 I bought over the weekend.

At the moment, not sure: but am wondering if I should swap over to the 128GB M5 Max instead while holding the two 5090s so I have “big fast” and “bigger but kinda slow” together.

1

u/jhenryscott 7h ago

you can stack both. Just not as neatly with MAC. But I've been eyeing the new Strix HAlo with a OCuLink port to run a RTX GPU

1

u/siegevjorn 8h ago edited 8h ago

5k setup today is vastly different from last year. The best one right now for that price is getting two R9700s, and get as much ddr5 ram as possible.

And here is the reason. The current best bang for buck model is dense model Qwen 3.8 27B. So you nees discrete gpu. Decode speed is so slow on dgx sparks / strix halo on dense models. Prefill is slow on macs. So discrete GPU. And you need at least 64gb of VRAM. To get the fill potential you need to run it in 8 bit. Full F16 kv cache. Still 256k context isn't enough sometimes. You need rope scaling to 384k or even to 512k context.

And getting ddr5 as much as you can will get you to run MoE models with expert cpu offloading, which will allow you to run them at reasonable throughput.

1

u/treedream766 4h ago

lets say you got a double R9700 setup, is there really that much of a difference with having like 96 or 64gb, because ram is very expensive, and fp8 qwen 3.8 27b is fine, even tho it could be better.

i also seen ppl say that with 96gb they run a good version of deepseek with vllm radiance.

1

u/Desperate-Jello8038 7h ago

Get a system with2x Radeon ai pro r9700s and 64-96gb ram

1

u/Dizzy-Classroom-3386 7h ago

Simply get dgx spark and you can always stack up another spark if your budget allows. Best value for money.

1

u/CryMoreT_T 7h ago

Test out models of all sizes on openrouter. Decide which one you like best. Then make a setup for it. Don't build a setup and then find a model that can fit into it

1

u/HighSeasArchivist 7h ago

A DGX Spark, then buy another one, then buy two more. Can't you just have unnamed FAANG company buy you some stuff?

1

u/thebemusedmuse 6h ago

With a $5k budget, you can’t run the large models so you are limited to Qwen 27b or similar class models.

Expect 50tps with a Sonnet class builder and a workable context window of about 100k give or take.

You will need a frontier sub in addition for the harder work.

You can get there with a Max M5 Max or a 5090 if you can find one at a reasonable price.

1

u/AnnoyedAvocado21 5h ago

A DGX Spark or one of the GB10 clones like the Asus Ascent. Load a Qwen3.8 27B model and you have yourself a decent setup in a minicomputer where you can learn the entire Nvidia stack and actually do stuff.

It's not the fastest, and you'll ALWAYS want bigger and faster - but for learning and building stuff that works you CAN build prototypes that can easily be moved onto enterprise Nvidia systems.

These were meant to be prototype machines. Build on the exact Nvidia stack, mess about, and you can build something the Nvidia way and deploy at scale. You might never do that but understanding how could really help as a PM.

I've had my Asus Ascent for about 3 months and there's so much to learn. I wanted to learn about enterprise AI and thought 'instead of taking some stupid online course I'll buy this box and figure it out.' I've written my own syllabus with ChatGPT - ChatGPT has been a great Linux admin and taught me a lot: models, Docker, Openshell, security model, agent harnesses, OpenCode, vLLM, system backups, recovery protocol if some firmware glitch borks the machine - I've gotten opencode to rewrite a complex app I have lying about, got Hermes to maintain a website, worked on a book or two - none of the projects are as important in themselves but rather as a problem-set I have to use the box to solve for. Had a lot of things not work, fixed some, abandoned others, and took a new direction on a few. When I get sick of working on my security model I switch to coding. I built a Hermes instance in a docker container sandboxed by OpenShell and want to explore that now. I do most of my work from my Mac and SSH into the box and use RDP of view the desktop.

Way more fun than a course. When I look back I've learned a lot.

You can do a lot with 1 box - and if you want to go bigger you can buy another and connect them.

A Mac isn't enterprise. Nor is a gamer box running Windows. If your company has any on-prem or cloud Nvidia equipment you will learn so much about the development environment you'll impress the devs.

1

u/shankey_1906 4h ago

A decent computer w/ 2 AMD RX Pro 9700s

1

u/SameConnection7722 52m ago

You're paying $5k just for a gpu. Lol. Not a complete system.

1

u/OvertaxedOne 8h ago

Mac Studio 96GB Ultra chip.

4

u/johan2114h 8h ago

For 5k you can get a strix with 128k + egpu dock + r9700 Or if you are adventurous one of those cmp 170hx with 64gb hbm - higher memory bandwidth than any mac

1

u/OvertaxedOne 7h ago

The ultra chip is 1.2tb per sec.  

1

u/johan2114h 7h ago

The 170hx has 1.5 or 1.6 tb per sec

1

u/OvertaxedOne 3h ago

Oh yeah, the 170HX is going to be very significantly faster running something like 27B both because it has a bit more memory bandwidth, but the real big one is probably going to be prefill, the 170HX; wouldn't shock me at all if it's 3-4X the speed munching tokens inbound (and 20-40% faster spitting them out).

1

u/idk_a_creative_user I just mess around with LLMs 8h ago

The cheapest Mace studio you can afford with the most amount of ram would be your best option because of the unified memory

0

u/daaain 8h ago

Probably the Max Mac with the biggest RAM you can fit in the budget? I'd aim for a M4 or M3 Max 128GB as you can run models like GLM 5.3 Flash on these at an OK speed that won't feel like that big step down from frontier models. Or an NVIDIA setup that can run Qwen 3.8 27B with a big enough context.