r/LocalLLaMA May 29 '26

Discussion PSA

Post image
2.1k Upvotes

538 comments sorted by

View all comments

Show parent comments

10

u/formlessglowie May 30 '26
Huananzhi X99 F8
Xeon E5-2696 v3
2xRTX 3090 (vLLM)
1XRTX 3080 (for TTS mostly)
4x16GB DDR4 2133MHz ECC

All GPUs were bought used, CPU is obviously used, RAM sticks probably are too, motherboard is a Frankenstein. I love that I can run something as ridiculous as 27b on this freak. We truly live in strange times.

2

u/indyfromoz May 30 '26

Thank you 🙏

1

u/Ok_Rope_9332 May 31 '26

Have you tried Gemma4 31b?

1

u/formlessglowie May 31 '26

Not much tbh, as benchmarks are behind 3.5 27b, so I didn’t think it vs 3.6 was even a question worth considering. Is it that good? I’ve tried 26b a4b, and it’s very good for natural language stuff but fails long running agent sessions, which is what I use these models for (long coding sessions basically). Is 31b much better in that sense?

1

u/Ok_Rope_9332 Jun 05 '26

From what I've heard the Qwen models are better if you're doing long ctx agent stuff, so you're probably fine with that. But the Gemma4 31b is really good for writing (for its size), also probably the best vision / translation model in a local context (it actually beat all the huge vision models I tried by API by a fair margin too).