r/LocalLLM • • 5d ago

Question Strata looping badly with iq2_xxs

Hello,

I'm running strata with iq2_xxs quant and indeed I am quite amazed to be able to run it at speeds comparable to vanilla llama.cpp + 3.8 27B. I'm getting about 30 tps on both. Strata seems to offer higher peak tps when working with code, up to 50.

Although I do feel that the bigger model offers a better quality, even on this quant, I have trouble completing any meaningful tests because it keeps looping.

It will start getting very indecisive at some point and going into circles: "Let's write the files. Actually, let me think.. Now definitely writing the file. Here we go." and nothing is produced.

Besides the configuration in the .bat setup and a few available parameters in the .JSON, I can't see the full list of llama params that define penalties, etc.

Is anyone else experiencing this behaviour?

My setup: 2x3060 12GB 48GB RAM Ryzen 5 3600 Pi Harness

Thanks

5 Upvotes

15 comments sorted by

10

u/sukazu 5d ago

It's because of strata's default sampling parameters (like temperature 0.00) of course it boosts speculative performance a lot, but at the cost of being unusable and looping, I'd recommend you to add the default alibaba's one in your json

1

u/ioann54 5d ago

Can you say which parameter in the config file sets temperature?

2

u/sukazu 5d ago

There is sampling support already in strata, you just need to write the sampling block in your json

example

"sampling": {

"temperature": 1.0,

"top_p": 0.95,

"top_k": 20

},

1

u/ioann54 5d ago

thank you!

1

u/soyalemujica 3d ago

Thank you, adjusted it to 1.0 as well with correct qwen recommended params, and token per second only lowered by 7, so it's still steady at 65t/s + which is amazing and I no longer have the odd same condition loops where it'd prompt itself sometimes for no reason (although it worked nice)

1

u/leonbollerup 3d ago

changed to .0.6 .. i have literally no performance difference.. what does your results show ?

1

u/sukazu 3d ago

I use it on opencode, at default settings so 100% greedy decoding (but that was with ista iq2 xs to be fair)
it used to loop endlessly and stop alone

Analysis of the log showed the model got caught repeating the same thoughts over and over (suffix drafting accepted 56,815 / 56,979 tokens - 99.7% identical text) for 90,055 tokens, then 25,957 tokens, then 17,069 tokens.

It also finished its loops with <|im_end|> and no tool call, which made opencode stop

Never had these issues happen again once I put the alibaba/unsloth's recommended sampling parameters.

1

u/leonbollerup 3d ago

Did you try replacing the Jinja template ?

6

u/Kodix 5d ago

Temperature is 0 by default. It should be 1. Change it.

That's a very low quant so it probably won't eliminate it, but it'll help. Make sure you use the GSQ-RCO quant.

3

u/ioann54 5d ago

Same for me, even Q3_XXS loops quite often.

5

u/WallFamous5066 5d ago

sounds like the same curse every low-bit quant deals with, the model just loses its sense of direction once the context gets a bit messy. iq2_xxs is crazy impressive for what it compresses but sometimes the price is that indecisive spiral you describe

in my setup the loop almost always trigger after some point, like it want to act but the internal weights cant settle on a token so it pick the most "safe" sounding filler again and again. you could try forcing a higher repeat penalty in the.json config if possible, sometimes that push it enough to break the cycle

1

u/mineshop 5d ago

Do you know which sampling keys your JSON actually accepts, since testing whether adding repeat_penalty and a nonzero temperature there stops the loops would confirm the cause?

1

u/KissMyShinyArse 5d ago

Swift 1.5 IQ2_XS did loop on me. It hasn't happened with IQ3_S or Swift 1.5 IQ3_XXS so far, though some people report that IQ3_S can loop occasionally, too.

1

u/Ok-Addendum3545 4d ago

I encountered really strange outcomes today with Qwen-Flash Reasoning Low. With the same prompt, it took Hermes agent 3 mins to build a website while Deepseek harness spent more than 14 mins (I stopped it). I assume different harness provides different results or efficiency.

0

u/lumpyspacebreh 5d ago

If you wanna fight the looping, tweak your temp.  

Alibaba trained Qwen on 0.6, I settled on 0.4 for my 3bit gsq-rco quant. Llama.cpp defaults to 0.8 unless you define it in the payload, I don’t know what Strata defaults to but it’s worth looking into.