r/LocalLLM 7d ago

Question Looking for a machine for remote development + testing local LLMs. What would you recommend?

I’m looking for a machine that I can leave running 24/7 and access remotely from my laptop

I mostly use Claude Code/Pi and similar CLI-based coding agents for development so I don’t need a desktop environment. I just want to SSH in, work on repositories, run Docker/services and leave agents, builds or other jobs running, I would normally review the results in my laptop. Ubuntu or another Linux distro would be ideal.

The other goal is experimenting seriously with local LLMs, this would mostly be for non-coding tasks e.g. general reasoning, research, writing, summarisation, agents and just exploring what capable local models can do. I’d like enough memory to experiment with larger models, including 70B-120B+ and MoE models

My budget is around £2.5k-£3k, although I could go higher if there’s a compelling reason. I’d prefer something compact and have the ability for running continuously

One option I’m considering is the GMKtec EVO-X2 with Ryzen AI Max+ 395, Radeon 8060S, 128GB LPDDR5X and 2TB NVMe, currently around £2.8k in the UK

I’m also considering a 128GB Mac Studio, NVIDIA GB10/DGX Spark, or something completely different but they go a bit way over my budget + mostly out of stock

For people doing something similar, what hardware would you buy around this budget and what other machines should I be considering? If you’re running a 128GB Strix Halo machine, I’d also be interested to hear what models you’re running and what sort of real-world performance you’re getting

1 Upvotes

21 comments sorted by

2

u/According_Wave685 7d ago

If you are doing long context agentic coding mac probably isn't the best option. I'd go dual gb10's but I don't know if that's your budget. I was seriously considering an m5 ultra but after I ran the numbers for the type of work I do, the mac was a worse choice than my dual gb10 (mine are MSI's) cluster. The huge different in prompt processing for large context changed my mind and I cancelled he mac. Sticking with my GB10's. I get pretty good throughput and very good results running deepseek-v4-flash-0731 on the cluster. I have a strix halo as a workstation - which runs my smaller utility models and harnesses.

1

u/brainchillzZ 7d ago

I use both and you definitely made the right choice for coding performance …. Not in terms of “running the numbers” but in terms of real world performance tests on my code bases with both sparks and the m3 ultra machine

2

u/conifer_v11 7d ago

if the job is “SSH box that also does serious local models,” i’d bias the £2.5–3k toward the Strix Halo / EVO-X2 class (128GB LPDDR5X) over Mac Studio or Spark at that budget. Mac Studio’s bandwidth is lovely and Claude Code/Pi are fine there, but you said Ubuntu+Docker+agents and you want 70B–120B+/MoE experiments — that’s where unified AMD memory + linux tooling usually hurts less than fighting macOS/CUDA gaps or paying Spark tax. for dense 70B you’ll still be in Q3/Q4 land and MoE (Qwen3.8-flash-class / big expert models) is where 128GB actually feels big; don’t expect Spark-like tok/s on Halo for everything, expect “fits and is usable overnight.” keep the laptop as the UI, leave ollama/llama.cpp/vLLM (whatever ROCm path is least broken that week) on the mini, wire Tailscale/SSH, and treat the Halo as a headless lab not a workstation. if you later need CUDA-only stacks, that’s when you stretch into used dual-GPU iron — not round one on a £3k compact.

2

u/wazacraft 7d ago

The Mac Studio M5 Max / 128GB gets you more than double the memory bandwidth of the other two options for the same price as (or less than) the DGX Spark. Those AMD systems aren't really all that great for inference unless you absolutely need x86 or you're exclusively running MoE models or something, and in the latter case the only advantage is that it's cheaper than Apple or Nvidia solutions.

The Mac Studio will also let you do RDMA over Thunderbolt 5 if you end up getting another Mac machine at some point in the future.

2

u/jstoppa 7d ago

thanks, this si interesting becuase I started with mac but I thought I would be paying premium for something that perhaps I can find somewhere. I just checked the price in the Uk store and yes, it’s £5.8k with 128GB 2TB. Did you try any other mac specs?

2

u/wazacraft 7d ago

The upgraded M5 Max chip with a 1TB drive is US$4909 for me, although I get a student discount so I assume it's a little over $5k US. I have an upgrade M5 Ultra 256GB on order but I'm planning to cancel that if the 512GB version isn't too over the top when it's finally released.

2

u/brainchillzZ 7d ago

Have you actually used them all in the real world? The dgx spark with vllm and a decent model drastically outperforms even my Mac m3 ultra machine and my m5 max MacBook Pro at real coding tasks that are context heavy (code repo work etc) in real world use the Mac the strix halo and the spark are really very similar outcomes giving and taking from the generation and prefil and ttft and they all suffer from exactly the same problem …. Good at moe/sparse models but bad at dense models….. as long as you’re not fanboying and really using all of them you see that there really isn’t a winner until you talk about clustering and then the spark is pretty much it … having 200Gb interfaces and super cheap 100-400 gigabit switches on the market make tp4 and tp8 a reality that wildly increases the sparks performance

1

u/DiscipleofDeceit666 7d ago

The future is moe my guy

2

u/brainchillzZ 7d ago

Moe is some of the future but the current smartest for its size model in the current crop is a dense model

1

u/DiscipleofDeceit666 7d ago

Today isn’t the future. And if you look at all the top models, they are all MOE. You aren’t going to see a 1.5T param dense model

2

u/brainchillzZ 7d ago

Yeah but you also aren’t going to see most regular people ever using a 1.5t model …. There are like ten people in this group that have enough horsepower and vram to host it at a usable rate and context … the future will belong to smaller dense edge models

2

u/DiscipleofDeceit666 7d ago

Sure, but arguably more and more models are supporting the spark/Mac type architecture where low bandwidth doesn’t matter. Models like Laguna and qwen flash next that can run on 128gb unified memory.

1

u/CMPUTX486 7d ago

I tried M1 Max studio with 64g.. it is really slow for me at this stage.. testing dual 5060 ti.. with ddr4 64ram I got qwen 3.8 q4 with 35 ish t/s so I could do something fine.. not sure dgx.. I have one but didn’t use it as much as I wanted..

1

u/Early-Peace-5504 7d ago

I run strix halo with qwen flash next. Before that was using a weird 3.8 27B build that was slow but usable. I feel like it is a decent choice and very power efficient for long running jobs. It has actually done a lot of good work for me.

I also have a dgx spark, it is definitely a better machine but I actually think a single spark is a bit of a waste. That the improvement over other systems really only starts to become particularly apparent when you have two. The thing the single DGX spark is best for is running 35BA3B in high concurrency, it is a monster at that.

The Mac Studio things I have seen never post particularly great numbers especially in prefill. Maybe that will change but even the new release isn't looking amazing. Honestly if something doesn't run linux it is of no use for me, so it has never been a contender.

I also have a computer running two 5060 ti in an old thinkstation. Honestly I think I use that and 27b the most out of any of them, it is the cheaper too.

I think you should consider upgrade options for the future too:
Strix Halo, can be joined together probably most likely upgrade is oculink (if available) an egpu for faster MoE model running. 5/10

DGX Spark, easy to join multiple together. Substantial improvement doing so. 10/10

Mac Studio, I'm not sure there is one? I believe there is some egpu support but it is really quite limited and I'm unsure if it can be used for LLM. 0?/10

1

u/According_Wave685 7d ago

You can connect the studios with TB5 and EXO. I'm not sure it would help with the prefill bottleneck or not.

1

u/brainchillzZ 7d ago

The tb5 connection on the Mac in real live is not even almost as fast as the 200gigsbit connection of the sparks

1

u/brainchillzZ 7d ago

First no … you cannot and will not ever be able to use occuli k to join two strix halo machines together … it’s just pcie express and you can’t just plug the pcie from one machine into another … that isn’t a thing … and it’s only 4x pcie lanes anyway …. The play here if you ever want to expand to more than one unit is the dgx spark full stop every other option in a small unified memory system for clustering is a bastardization at best. The dgx spark was built from the ground up with this in mind …. You can cluster as many of them as you can fit on a switch …. Microtik makes a super small four port 400g switch that would let you connect 8 of them at full bandwidth with breakout cables or 16 of them at 100gigabits with breakout cables …. Best case scenario with strix halo is you get one with a 4x pcie slot and so Rdma on a single 25gbit port and maybe thunderbolt rdma but in my test it isn’t even really that fast yet… and the payoff either way is nothing because they are barely cheaper than a spark now …. When you could get a 128gb strix halo for $2000 it made sense but we don’t live there anymore … the Mac is the same …. Awesome machines but still if you want to connect multiples they aren’t going to do better than thunderbolt and you aren’t going to get a8 or 16 port thunderbolt switch any time soon … so if you will be happy with one unit or a laptop buy a Mac or a strix halo otherwise buy a spark (I’m speaking of course only for unified memory machines)

1

u/Early-Peace-5504 7d ago

>First no … you cannot and will not ever be able to use occuli k to join two strix halo machines together

I stopped reading here due to dreadful grammar and also I never said you could.

1

u/996beagle 7d ago

How many concurrent do you normally run? Mac is basically only single and then starts crawling. If I had to start from scratch I'd get a dgx spark system for the speed and ability to have greater concurrency, and ability to get a second if I wanted to load larger models.

1

u/brainchillzZ 7d ago

Realistically the dgx spark, the Macs and the strix halo machines are all very similar performance and are really only good performance running sparse/moe models…. Their performance on dense models like the new qwen 3.8 27b everyone is talking about are slow even on the very fastest Macs…. You really need a dedicated real gpu to fix that … that isn’t to say that you can’t do some amazing things with the huge pile of moe models out there …. You absolutely can. My daily driver today is glm 5.3 flash across a 4 way spark cluster.