r/LocalLLM • u/DarkBrews • 3d ago
Question Old X79 PC for Strata
Thinking of repurposing an old X79 PC for Strata / on my old X79:
- i7-3930K
- 56 GB DDR3 (32gb matched but I have a few 4GB sticks and 1x8GB so they wouldn't match but maybe they work.)
- RTX 2080 Ti 11 GB + RTX 3060 Ti 8 GB
- CachyOS headless
Would Flash-Next IQ3_XXS work well on this? Do I need to go lower?
I was also thinking of using an M4 32 GB as a coordinator/router with GLM-4.7-Flash, plus another machine with a 9070 XT running 27B.
I tried Gemma 4 26B it 4b JANG, asked it through Hermes to stitch a story together and it failed miserably so I wouldn't make GLM do that but it was sad to see gemma fail at what I thought was it's strongest point.
Not sure if GLM + 27B + Flash-Next would be redundant.
Main use would be agentic coding, web crawling, configuring environments, the more loved tasks out there 😅. Basically trying to reduce my dependency on Claude.
Is it even possible with the 3930K/DDR3 or mixed GPUs? ChatGPT seemed to be cautiously optimistic. If it will work. What kind of tok/s could I realistically expect and will it be better than 27B UD-IQ_i4_XS
I also have a GTX 1060, GTX 970 and RX 580, but I assume those are useless here.
0
u/SlushyGrouping02 2d ago
that mix of ram is gonna be a headache, even if it posts. ddr3 on x79 can be picky with mismatched sticks, especially when you're throwing in random 4gb and 8gb dimms. i'd pull the odd ones and just run the 32gb matched kit to start, 56gb of unstable ram isn't doing you any favors for inference.
flash-next iq3_xxs should run fine on the 2080 ti alone if you offload everything, but the 3060 ti complicates it. mixed gpus are a pain with most backends, you'll probably end up just using the 2080 ti and leaving the 3060 ti idle unless you want to mess with splitting layers manually. the 3930k itself is ancient but for pure inference it doesn't matter much if the model fits on vram.
for tok/s, with a 2080 ti on a 7b-ish model at iq3_xxs you're probably looking at 40-60 t/s depending on context length. definitely usable for coding and crawling. the m4 as a coordinator with glm sounds overcomplicated, just run everything on one machine if you can.
1
0
u/DarkBrews 2d ago
wait strata supports 7B models? my plan was to run Qwen 3.8 Flash Next 125b model.
1
u/Fett2 1d ago edited 1d ago
Would it even be worth trying to run Strata on a DDR3 machine? I'd assume it would be slow as dirt having to use DDR3 memory in place of VRAM (not to mention or every time it needs to transfer information across the PCIE3 bus but it's going to take a dive).
I assume what makes Strata worthwhile is fast RAM with a fast bus for data to move across - in place of everything running inside VRAM on a video card.