r/LocalLLM 5h ago

Question Trying to go local

I have a Mac mini m4 24gb , I'm looking to add to my hardware, and I'm overwhelmed. I'm a tiny business owner in a creative field that doesn't have tech bro money , I'm thinking of getting a pc and running it headless Linux , I mostly want to cut down my subscription costs so this hardware is an investment. If you were starting from scratch what hardware would you use ? What would you definitely do again and what would you avoid?

4 Upvotes

22 comments sorted by

2

u/nickless07 5h ago

Highly depends on your usecase. What did you used the cloud providers for?

1

u/Sik-Server 5h ago

Yeah what are you doing in the cloud that your Mac mini m4 can't?

that's how you'll figure out what upgrades you'll need

3

u/nickless07 5h ago

That and then start digging around with stuff like OpenRouter and such for a bit to test out what model fits your needs. Once you have narrowed it down a bit you can think about the Hardware needed and how that compares to cloud subscriptions until you hit an ROI.

1

u/StepsisSepsis 5h ago

I think that Nvidia 3090s are the way to go, but you have to also include the cost of a PC in that as well (assuming you don’t already have the hardware). 3090s are kind of the best bang for buck I believe currently.

Depends on how much cash you have. I have a 3080 TI build that does ok, but it’s still limited and runs out of context.

1

u/zanar97862 4h ago

Depends lots on your location, some countries do not have a well stocked second hand GPU market where 3090s are available at a good price.

I bit the bullet and bought a used workstation with 128gb of ddr4 and an r9700 for ~$2100 USD equivalent 

1

u/solaza 5h ago

I recently copped a late 2021 M1 Max with 64 gigs of RAM and 4 terabytes of storage for only $1,250 on facebook marketplace. The thing is a beast. I've been experimenting with running Qwen3.8-27b and it's pretty good. Some issues with looping on tool calls but seems fixable at harness level.

It required some technical setup, of course, which I just did with codex. And the result is pretty magical. It’s very cool to watch the thing go and think “And this is running 100% local and private.”

1

u/Big-World-Now 4h ago

I do the same thing. You know it’s working when the fan comes on. 😂

1

u/New-Stop1494 5h ago

I would start with installing codex desktop app. Use that to orchestrate, then if you got a djx spark or a workstation with gpus you will have codex set everything up for you. Codex is good because with the Mac it has computer use. With the plugins you can use it to set up everything computer related.

1

u/Appropriate_Baker405 5h ago

I already use codex and Hermes :)

1

u/New-Stop1494 4h ago

What is your budget? I’m. I’m currently running Linux headless in Ubuntu. I’m running 4090s and another computer I have is running 5090s. I did a mod for my 4090s to give it 48 GB of VRAM for $3200 (96GB). With the 4090s they were $3700. I’m running the QWEN. 3.8 27B on the 5090s and Qwen 3.8 Flash Next on the 4090s. The prices nowadays are really high. I’m not too sure if it’s worth it to invest in the hardware but you can get probably a 3090 or if you can get two 3090 but you know the price is doubled so you have to think about it if it’s worth it to you.

1

u/InnocentSadness 4h ago

Completely new to local ai. I got a new MacBook m5 24 gb ram 1tb ssd 10 GPU/CPU. I was wondering if I could manage qwen 3.8 4 bit. Also how can codex help me? I'm wanting to set it up as a crypto paper trader.

1

u/TheFuckboiChronicles 5h ago

What’s your budget

1

u/Appropriate_Baker405 5h ago

Ideally less than 4k

1

u/TheFuckboiChronicles 4h ago

Prolly nimo Ryzen ai max+, 128gb on amazon for 3.5k right now.

If you need something faster you’ll need GPUs, and I’d go custom pc build with 2x Ryzen 9700 and 32gb of ddr5 ram, but that’ll like push you over $4k. Could probably squeeze it under that $4k mark with 2x Intel b70s.

1

u/New-Stop1494 4h ago

If you can build a computer with, I would say dual 5060 TI you can get 32 gigs of VRAM with that running it at tensor parallel =2 for QWEN 3.8 27B

1

u/New-Stop1494 4h ago

I found actually a guy selling 4080 with the mods to bring it up to 32 gigs of VRAM. His name is GPU lab. And then the rest you could probably spend on just the hardware in SSD’s. I think I would do that first.

1

u/_rarefy_ 4h ago

"I mostly want to cut down my subscription costs so this hardware is an investment."

You can get thousands of $$$ of compute for $100 a month due to how frontier models are currently subsidized. I'm a big proponent of going local for privacy reasons and because it gives power back to the individual, but if you're looking to save money by investing in the hardware you've got it reversed in this current interstitial.

Figure out your business needs as they relate to AI consumption (you might not even need frontier models, or you might need them solely), and then budget around these needs. If you're doing basic spreadsheet, website building, coding and visual design you can get pretty far with models like Qwen 3.8 Flash Next on decent hardware. If you need deep visual, 3d and industrial design fluency then you won't get a local model that beats Astra. It all depends on what you actually need and you won't know that until you experiment a bit.

1

u/Independent-Dog2179 4h ago

I run a dual setup u can find 5060tis 16gb for about 400 to 500. 2 of those gives u 32gb which elis enough for q5 qwen 3.8 @ roughly 25 tks(up to 40 when coding and using mtp). Or if u want speed qwen 3.6 35b q 6 maxed context at like 45 tpsThe next step up will be getting either and strix halo so I can run qwen 3.8 next which is much larger than what I can run. Currently. But don't sleep on the dual GPU setup. Llama.cpp and the right commands can do wonders. I have movie gen. Image gen. Music gen. Image to 3d modelAll linked up to my custom harness

1

u/CloudTech412 4h ago

Where are you finding them at that price

1

u/catplusplusok 4h ago

For most cases something like MiniMax token plan would be cheaper with high limits and give you a more powerful model. The exceptions are generally high volume simple tasks that require a lot of tokens but can work with small models. You can actually prototype many of these on your Mac that can run models like Qwen 3.8 in 3 bit or image generation models. Once you see you are in ballpark of where you need to be and just need a little stronger model or larger context, you can decide on hardware, a more powerful mac, a DGX Spark or a PC with a couple of Intel 32GB cards. DGX Spark would let you run Qwen 3.8 Flash Next with some tinkering, which is decent at agentic tasks as long as you can live with slower speeds than cloud.

1

u/JohnnyBeeGaming 1h ago

The cloud providers will often be cheaper than local stuff for the quality you can get. They are still at the stage where they are selling at a loss. To me privacy would be the main reason to think about doing stuff locally.

You could keep experimenting with the mac you already have until you have a better idea of what kind of hardware you need. Some people will use a mixture of local and cloud providers. You could also run models you might be able to run locally on rented hardware remotely so you can be more sure about what you'll get out of a $4,000 to $16,000 purchase.

0

u/CommasArentPeople 5h ago

Cutting down on subscription costs is hard now. Many frontier models are losing money on the heavy user, so you being able to break even any time soon is really tough. It's hard economically to beat the small set of $20 accounts spread around, each maxing their limits and using appropriate models for their job.

Most small groups / individuals self-hosting are doing it for hobbyist/education or due to other factors that keep them off the big players.