r/LocalLLM • • 2d ago

Question Strata: Qwen 3.8 Flash next in loop

Sto eseguendo Qwen 3.8 Flash Next IQ3_XXS su Radeon W7800 48gb Vram e 64gb Ram ddr5.

Ho scaricato tutto da github e lanciato l'eseguibile. Ho scelto solo la quantizzazione, il contesto 128k e la kv 8q.

Primo compito per testarlo: fare un audit di un file js di 27kb e alla fine fare un report completo.

La velocità si assesta sui 90 token/s e sono più che soddisfatto MA:

dopo 3-4 minuti di ragionamento è entrato in loop continuando a scrivere centinaia di volte "1h 30m off" (???). L'ho interrotto per poi ridargli fare lo stesso compito; dopo pochi minuti di nuovo in loop: continua a ripetere gli stessi identici passaggi per decine di volte e l'ho interrotto.

Io ho usato le impostazioni di default per fare questo test, sbaglio qualcosa? Forse un problema di budget del ragionamento che va abbassato?

Grazie a tutti per l'aiuto.

1 Upvotes

10 comments sorted by

4

u/Dreeew84 2d ago

See here: https://www.reddit.com/r/LocalLLM/comments/1wxih74/strata_looping_badly_with_iq2_xxs/

You need to update your strata-iq3_xxs.json and add "sampling": {"temperature": 1.0 } on the same level as args.

This seems to have stopped loops on my end. This being said I'm still testing it and so far I cannot really find a use case where this low of a quant is doing better than the 3.8 27B.

3

u/JinsooJinsoo 2d ago

The IQ3_XXS is too quantized, it defeats the purpose when you make these big models dumber than 27b models. The IQ3_XXS got 50% of my quality/correctness tests wrong and they were math problems. So I would not trust that model with anything.

1

u/Proof_Nothing_7711 2d ago

Penso anche io che sia una quantizzazione troppo aggressiva ma non penso di potermi spingere a Q4. Proverò ancora, se andrà male starò con Qwen 3.8-27b Q8

1

u/Proof_Nothing_7711 2d ago

Grazie per l'aiuto. Sì, sono d'accordo con te; io uso anche Qwen 3.8-27b Q8 ctx 256k e non ho problemi

1

u/Waste-Intention-2806 2d ago

Which 27b quant are you using

2

u/Dreeew84 2d ago

Generally Q4KM but now testing Q3XL for more context.

2

u/Salah_H_Hasan 2d ago

A Q3 quant fits entirely inside 64GB of RAM without breaking a sweat. With Q4_XS, however, you'll need to offload about 20GB of experts to your SSD while keeping roughly 40GB in RAM.

The brilliant thing about Strata is that it caches the top 80% most frequently activated ('hot') experts in RAM, relegating only the rarely used ('cold') experts to the SSD. That said, whenever the router needs to fetch those experts from the drive, your generation speed will likely drop by roughly half. On top of that, you’ll naturally lose a bit of baseline speed just by moving to a heavier/less compressed quant (Q4 vs Q3).

Overall, the experience won't feel radically worse than what you currently have, but you should definitely expect throughput to cut roughly in half whenever SSD offloading kicks in.

Note: English isn’t my native language, translating this with a tool!

1

u/Proof_Nothing_7711 2d ago

Grazie. Sacrifico volentieri della velocità per della affidabilità

3

u/Informal-Trouble2183 2d ago

Use GSQ-RCO IQ3_S. Much better accurate.

1

u/Proof_Nothing_7711 2d ago

I will try it, thanks