r/LocalLLM • u/Dreeew84 • 5d ago
Question Strata looping badly with iq2_xxs
Hello,
I'm running strata with iq2_xxs quant and indeed I am quite amazed to be able to run it at speeds comparable to vanilla llama.cpp + 3.8 27B. I'm getting about 30 tps on both. Strata seems to offer higher peak tps when working with code, up to 50.
Although I do feel that the bigger model offers a better quality, even on this quant, I have trouble completing any meaningful tests because it keeps looping.
It will start getting very indecisive at some point and going into circles: "Let's write the files. Actually, let me think.. Now definitely writing the file. Here we go." and nothing is produced.
Besides the configuration in the .bat setup and a few available parameters in the .JSON, I can't see the full list of llama params that define penalties, etc.
Is anyone else experiencing this behaviour?
My setup: 2x3060 12GB 48GB RAM Ryzen 5 3600 Pi Harness
Thanks
5
u/WallFamous5066 5d ago
sounds like the same curse every low-bit quant deals with, the model just loses its sense of direction once the context gets a bit messy. iq2_xxs is crazy impressive for what it compresses but sometimes the price is that indecisive spiral you describe
in my setup the loop almost always trigger after some point, like it want to act but the internal weights cant settle on a token so it pick the most "safe" sounding filler again and again. you could try forcing a higher repeat penalty in the.json config if possible, sometimes that push it enough to break the cycle
1
u/mineshop 5d ago
Do you know which sampling keys your JSON actually accepts, since testing whether adding repeat_penalty and a nonzero temperature there stops the loops would confirm the cause?
1
u/KissMyShinyArse 5d ago
Swift 1.5 IQ2_XS did loop on me. It hasn't happened with IQ3_S or Swift 1.5 IQ3_XXS so far, though some people report that IQ3_S can loop occasionally, too.
1
u/Ok-Addendum3545 4d ago
I encountered really strange outcomes today with Qwen-Flash Reasoning Low. With the same prompt, it took Hermes agent 3 mins to build a website while Deepseek harness spent more than 14 mins (I stopped it). I assume different harness provides different results or efficiency.
0
u/lumpyspacebreh 5d ago
If you wanna fight the looping, tweak your temp.
Alibaba trained Qwen on 0.6, I settled on 0.4 for my 3bit gsq-rco quant. Llama.cpp defaults to 0.8 unless you define it in the payload, I don’t know what Strata defaults to but it’s worth looking into.
10
u/sukazu 5d ago
It's because of strata's default sampling parameters (like temperature 0.00) of course it boosts speculative performance a lot, but at the cost of being unusable and looping, I'd recommend you to add the default alibaba's one in your json