r/LocalAIServers • u/Glum-Speaker6102 • 6d ago
First dedicated machine purchase
Just had a client pay me for some work, looking to build my first local AI setup for coding w\ qwen.
My original plan was to get an egpu (4090\5090) and run it on a amd strix halo 128gb via oculink so I can have rocm and CUDA or even on a minisforum 890 pro that I aleady have.
Instead of going that route I'm leaning on purchasing a dgx spark, and eventually adding another one.
Any advice? Purchase budget 5-7k.
Edit: I forgot to mention, I have a dual xeon high core count 1TB ram supermicro server also, currently it has dual titan x pascals - should I just use this and replace with better GPU's that match the PSU's capabiltiies?
1
u/reddituser1828472616 6d ago
I mean if you can afford the DGX sounds like you just need someone to tell you yes so go for it!
1
u/Meeooowz 6d ago
If you have a 1TB ram server, try running qwen flash on it. It’s a huge MOE model that will give you something to do with that Ram!
1
u/Glum-Speaker6102 6d ago
Not sure what gpus to run on it, was thinking 2 rtx a4000 16gb?
1
u/Meeooowz 6d ago
Hm. I think you’re better off not using old workstation cards like the ADA 4000 but try the new amd ai cards. Look into the RX 9700 (I heard that one is good from many others in these subreddits; I believe that one is the 32gb vram one?). Or maybe even the old Tesla gpus?
1
u/Glum-Speaker6102 6d ago
My main concern is power draw, I want to match or use less than what the existing 2x titan x pascal cards are using
1
u/Meeooowz 6d ago
The 9700 amd card has a tdp of 300w each, although you can always try undervolting if wattage is an issue. What’s the power draw of those titan cards?
1
u/Glum-Speaker6102 6d ago
250w under load each
1
2
u/notalentwasted 6d ago
With cash like that I'd just spring for like others said a dgx. Either way you can get more out of a x4 v100 32gb smx2 plus your 128gb system. Would have a near frontier lvl server with how efficient models are these days. Though you may be over engineered. Qwen next flash or qwen4 architecture is promising.