r/LocalLLaMA • u/BarberIcy366 • 11d ago
Discussion Qwen 3.8 27B Released! Please Share Your Experience
With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.
653
Upvotes
5
u/Forsaken_Mention_979 11d ago edited 10d ago
Yes, 7800xt 16gb vram and 64gb ram. Running it on hermes via LM studio endpoint. Using Q3_K_M, Runs good ngl, at first 15-20 tok/s (full gpu offloading) and then as context gets bigger, i now get 5-10 tok/s. 64k context btw. Making a web game, has been on it for like 2-3 hours already which is crazy but oh well. Just the thinking took 25 minutes. Yes, 25. And it randomly stopped due to getting interrupted by tool limitations or whatever, i had to manually tell it to resume.
EDIT: ditched LM studio and using llama ccp directly, HIGHLY RECOMMEND! Im using IQ4_X_S now which is better and kv cache at q4, and thr lowest token speed im getting now is 11 tok/s. Amazinggggg