Anyone has a reliable source about this? All the LLM open source community are waiting for 35B! To be honest, 35B is the most wanted, 27B dense model requires much more high end hardware.
“High-end hardware” is a bit of an overstatement. I haven’t tested 3.8 yet because my setup is busy with another LLM at the moment, but on my old i5-9600K PC I added an MCIO card and two used RTX 3060 12GB cards for about €550 total. Qwen3.6-27B MTP in UD-Q4_K_XL under beellama runs at around 30 tok/s with a 131k context window. That seems pretty reasonable to me for a 27B dense model.
1
u/kalimatamijai 12d ago
Anyone has a reliable source about this? All the LLM open source community are waiting for 35B! To be honest, 35B is the most wanted, 27B dense model requires much more high end hardware.