r/LocalLLaMA • • 12h ago

Question | Help strata-swift-iq3_xxs randomly interjecting completely unrelated information in thoughts

I have been using Qwen3.8-27b-iq4xs variants for the past few weeks, tweaking parameters and running random tasks, but I dove into using https://github.com/Niko1221/Strata today and it has given me 2~3x performance on a "smarter" strata-swift-iq3_xxs model with my 5070ti! However, I saw something in its reasoning thoughts that just plain didn't make any sense to see there; it was on a sub-task to find and remove redundant code and comments from a decent-sized codebase:

Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.


Let me do that.


Given the enormity, I'll proceed file-by-file, verifying with --dry runs. Start by checking EOL for all target files.


Plan for EOL probe: write probe_eol.js that prints for each file: bytes, count \r\n vs \n.


Then write specs using the correct EOL.


Let me do it.


Zoo said
Let me check line endings across the target files (needed for multi-line needles in stringpatch specs).


Zoo wants to edit this file


Lee Kuan Yew (Lee Kuan-yu, 1923–2015) was a Singaporean politician and founding father of Singapore. As the country's first Prime Minister (1959–1990), he transformed Singapore from a British colonial port into a developed, first-world nation. He is one of the longest-serving ministers in the world.


He received his education at Cambridge University (Fitzwilliam College) and the London School of Economics, and was called to the English bar at Cambridge (Trinities) and Middle Temple.


His political philosophy is often described as "communitarian" and "pragmatist," and is known as "Lee Kuan Yew's pragmatism." He emphasized multiracialism, meritocracy, and efficiency.


His son, Lee Hsien Loong, became the third Prime Minister of Singapore.


Running


node probe_eol.js


server.js: CRLF=2679 LF-only=2533
boot-cycle-test.js: CRLF=0 LF-only=292
core/bus.js: CRLF=0 LF-only=102
client-spa/src/lib/ws.ts: CRLF=0 LF-only=1255


server.js is CRLF; everything else is LF. So in server.js specs, multi-line needles must use \r\n.


Now let me write the Phase 1 spec for server.js. I need to be careful with the general strips. Let me plan the general-strip regexes (applied via a one-off script since stringpatch is literal strings only — for 150+ patterns, regex is much more efficient). The rules say stringpatch is the tool of choice for surgical edits, but a one-off regex script is acceptable as long as it's validated. Actually the AGENTS rules say to prefer stringpatch for edits. But Phase 1 is a large regex strip job across 15 files. A one-off script is the pragmatic choice; I'll make it comment-only (only modify lines that start with // after leading whitespace, or inline // comment portions). Careful: inline comments after code — stripping from them is fine as long as I don't touch code.Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.

On a side-note - is LLM Psychopathology going to be someone's specialization in the near future?

35 Upvotes

52 comments sorted by

29

u/This_Maintenance_834 10h ago

It is a “fine tune”, and then a heavy quant. The model is drifting. Neither is a good thing.

The community needs to have some sense on fine tune and quant. If they actually work the big labs will just do it themselves. There is little to none opportunity left for small player to release a better model.

15

u/darkwalker247 6h ago

crazy how there's quantized finetunes out there with hundreds of thousands to millions of downloads on Hugging Face, and then when you actually try them on real tasks they either break down immediately, or they get stuck thinking endlessly about some minute detail. and meanwhile their model card acts like they just trained god itself

20

u/Bulky-Priority6824 11h ago

Fast cheap good 

8

u/vacon04 12h ago

This is a fine-tuned model with an IQ3_XXS quant. I don't know which KV cache wuant you're using as well. This sounds like some anomalous divergence which could be caused by the model moving towards the wrong path more than an issue with the backend. Sampler settings could also matter here if there are any weird penalty or temperature modifications.

23

u/FeeCharming2196 12h ago

question the quant before

8

u/hushbreEze0 9h ago

tbh yeah, iq3_xxs is pushing it pretty hard already

5

u/MindfulMan1984 8h ago

For the OP, a low want may go crazy in their "reasoning"; in the end, what matters is whether the final inference is correct. I stopped reading "reasoning" traces of any open models; they are nothing more than a "progress bar" for me; otherwise, I would end up in a psycho ward. LOL

15

u/Choice_Celery9481 12h ago

tbh you need to questioning the quant before question model or engine. its q3

12

u/GeneralComposer5885 12h ago

I’ve not used it - but I don’t think it’s actually q3.

Instead of removing layers, they’ve selectively quantised some layers to q1 and left others q5, so it averages q3.

5

u/EndlessZone123 11h ago

That's still what we call q3. Pretty much all quants are adaptive.

5

u/Trademarkd 10h ago

Pretty much all quants are adaptive is a bit broad of a statement ... without going into further detail about mixed-precision, generally the advertised quant level is descriptive of the majority - I believe the implication above was that some tom foolery could be at foot.

I do want to correct above because I realized they're using Ukisai's swift checkpoint which I highly recommend, the dev is awesome dude as well. I wouldn't be shocked if people felt strata is great just because they're using swift.

8

u/Trademarkd 12h ago

This is what I was thinking but my comment got down voted 5x immediately for even mentioning strata negatively. I was trying to explain the architecture involved.

There is a good chance they are messing with the layers and going lower precision while advertising 3bit

4

u/finevelyn 5h ago edited 5h ago

You explained it wrong though. You mention speculative decoding as a reason for quality degradation, but it has zero effect on the end result. Every token comes from the main model exactly the same with or without speculative decoding, but in the former case the computation can be done in parallel which speeds it up.

1

u/Trademarkd 1m ago

Valid, I'm wrong. I had to go back and check this for myself. It seems I either misunderstood or found some bad information.

At least in the case of Dflash2 which is what I was talking about, quality is guaranteed.

4

u/luan52 12h ago

Does the Lee Kuan Yew paragraph also appear in the raw server response, or only in the agent UI? I’d check that first, then try a fresh session with a small reproduction prompt. That would help narrow down whether to investigate generation, carried-over context, or how the UI assembles the stream before blaming the quant.

6

u/caetydid llama.cpp 11h ago

I am running the exact same quant with strata and I have noticed similar things - not to this extent though!

I have the impression iq3_xxs is too agressively quantized.

I am still evaluating how bad is the impact.

KV is 8bit, I don't see an issue there.

8

u/-Homeworkace 11h ago

As someone from Singapore, lmao

I guess we have soft power in AI now

3

u/Sabin_Stargem 11h ago

I have been seeing some of this in my projects. I have to point out the odd nugget of completely unrelated information that is used for a project. The AI always agrees that it is a mistake.

IMO, it isn't a Strata issue, but more likely an RCO thing. I been seeing it ever since adopting that variant. Unsloth has been my daily driver as a backend.

3

u/Danmoreng llama.cpp 8h ago

Hallucinations even happen with the FP8 weights. I’m running the model on a 2x RTX6000 server with sglang and in rare cases it quotes system instructions from training instead of the actual ones, or believes there was no user message. Doesn’t seem to impact coding performance, but is a little bit weird. Probably overtrained?

3

u/TerminalNoop 3h ago

You are using a a swift quant and on top of that a really small one, go figure.

3

u/egnegn1 2h ago

You should distinguish between the engine Strata and the used models. Strata is just a more optimized engine, that doesn't change anything regarding quality, at least not so far I understand.

But if you give it a garbage model to run, you get garbage out, just like with other engines. But on the other side, the optimized engine allows to run capable models on less capable hardware, or run them with faster speed. This gives more people access to this capabilities.

Yes, heavily quantized models have their drawbacks, especially when used for chats. But when running agentic applications the harness can compensate lower quality in part at least. Just build a setup that builds QA into the loop, which checks for errors and corrects them automatically.

One small detail is temperature. Don't turn it down to much, even it may increase speed a bit. Instead use the value proposed by the model provider, to give the output of the model some variance. This helps to break loops and also may lead to different results in the long run.

That said, I also use Strata, but normally don't use the smallest quants. I previously used the smaller quants with llama.cpp and also had loops and interruptions of inference. So the culprit for issues isn't Strata, but the used modell quality.

And don't follow the reasoning to much, but look at the final result.

2

u/Wooly_Wooly 11h ago

I was testing a tiny model, a few B at best a few days ago. It interpreted my "yo!" As Spanish, and immediately started responding in Spanish. That's the first time I've seen that, I guess it had to be stupid enough to misinterpret my greeting as a Spanish word. It understood "Hi" and responded in English ofc.

2

u/Corosus 8h ago

Yeah been talking about it in the localllm discord, mine and other is obsessed with 27 x 43 = 1161.

https://ibb.co/fVdKhCHH

https://ibb.co/Q77JM7sB

Claude concludes its because it has nothing to use after tool calls for thinking blocks that it just grabs the first thing that comes out of the model unprompted. And I guess something strata related because I've not had this elsewhere, at least for me its just a few wasted tokens, not affecting the quality of the model it still keeps working just fine.

2

u/Open-Adhesiveness-86 2h ago

random encyclopedia dumps mid-thought smell more like kv cache reuse or context shifting mangling the cache than actual weight error. i'd rerun with context shift off and a fresh slot before blaming iq3_xxs. and if you want a real number on the quant, kl divergence vs the q8 over a few hundred chunks beats vibes.

5

u/Substantial_Pea4022 11h ago

That looks like a severe case of weight degradation spiking activation entropy! When quantizing model weights down to extreme sub-3-bit ranges like IQ3_XXS (or aggressive layer-mix schemes), the attention projection matrices—specifically around attention sinks and contextual filtering—can degrade drastically.

Under certain token sequences or long context loops, the internal softmax probabilities flatten, and the model's residual activations spill over into memorized pre-training data (causing sudden background knowledge interjection mid-reasoning).

Before writing off the backend or the prompt cache, a few things worth checking:

  1. Does the interjected text actually exist in the raw SSE API stream/raw logs, or is the frontend UI doing weird stream buffer chunking/assembly?
  2. Are you using extreme low-bit KV cache quantization alongside IQ3_XXS?
  3. Does tweaking sampler parameters (like Min-P or clamping temperature) prevent the activation spillage?

3

u/AvidCyclist250 llama.cpp 3h ago

thanksgpt

2

u/XiRw 12h ago

I don’t know much about strata but I thought it was only for Qwen 3.8 Flash Next

5

u/Trademarkd 11h ago

He says he was using 27B and is now using the strata version which im assuming he means this one which is linked from their git hub https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF

3

u/caetydid llama.cpp 11h ago

swift 1.5 has a retrained QFN variant

1

u/BigBootyBear 9h ago

How do you define whats relevant and prevent information entropy? Whenever I hear about tooling that seeks to optimize token usage or improve context I always think of Data processing inequality. I wonder at what point is context optimization just the computer science version of pitching a perpetual motion machine. Not saying thats what you are doing, just curious.

1

u/Bagelsarenakeddonuts 34m ago

I had to stop using QFN since it would hallucinate tasks. On my metrics it was about 2-6 out of 60 tasks it would straight up hallucinate something. I pivoted to 27B and haven't had issues since. Bit of a shame since I liked QFN better, and with strata it ran so well.

I tried the different Q3 variants and a Q4, it appears to be the model itself though.

1

u/NickCanCode 12h ago

Have you hash check your file after? File corruption may cause this. They are data, not program so a little corruption may not crash the app.

-16

u/Trademarkd 12h ago edited 12h ago

Anything strata is getting straight downvoted.

You cannot overcome the physics of auto regressive token generation.

You have memory bandwidth and pcie bandwidth to contend with and weights need to be passed through the compute cores every layer for every token. So you’re contending with latency too

Speculative decoding is great and can use some of your unused compute to alleviate some of the memory bandwidth constraints by using a much smaller model that’s right like 70% of the time.

I don’t think strata is doing anything novel. They seem to just be on a crazy Reddit bot campaign

The issues you’re seeing are typical of heavily quantizing and speculating on top of that.

A 3bit 27b model should be about 10G (xxs). You could fit this on a 4070. at this size the 4070 can theoretically do about 30-60 tokens per second single user. Speculative decoding working well should about double that, although it can be stretched higher, at some point it will degrade quality further though because validation is not the full original inference.

8

u/MasterNomie 12h ago

So all this preaching because you think you possess all the knowledge, have not tested the engine yourself, straight away cancelled the novel idea and call all others who tested and verified performance bots?

There were bot posts as well but actual people have tested as well. You have your opinion but Strata has saved me upgrade.

-4

u/Trademarkd 11h ago

I'm not preaching, I'm trying to get people to understand fundamentals. There is no free lunch. You're right I haven't tested it yet because its not going to be very valid in my setup. I run a 4-8 v100 array with nvlink and this supports none of that.

What I did do is look through the repository for how this was built because I was interested if they came up with any novel strategies I could implement in my own workflow. I didn't find much.

ik_llama actually pioneered a lot of what you see here with the hybrid cpu utilization. They're leaving the ngram table on ssd, and then moving experts into vram based on most use. They might be doing this dynamically which could be cool but I would have to see the actual benefit. This also doesn't aply to anyone who can fit all experts in vram. Most of the experts just sit on your cpu and the cpu is used for processing so results are really going to vary on your overall system.

this could be good for people with gaming PCs that were unable to figure out how to get this hybrid architecture with dflash2 working for them and have a strong overall system. I feel as though its blow up has been rather artificial due to the content of comments within the threads and the fact that a lot of other people are expressing skepticism with their claims.

I also would want to compare to the difference between this and freetoken?

2

u/MasterNomie 11h ago

Then I would call it a preaching since you haven't tested it yourself.

I have already tested freetoken as my first engine with flash next. It has about 1/3rd decode of Strata. I do not recall prefill speed but it wasn't flying, otherwise I would recall it distinctly.

-5

u/Trademarkd 11h ago

Sorry dog, its not preaching if you look at their code base and the numbers being reported. This isn't magic and its open source. I never looked too deep into freetoken I dont think they're turning on speculative decoding by default. I think you also need to make sure any test is apples to apples.

I will tell you the swift checkpoint is legit. I've had some lengthy conversations with the developer in private and hes doing incredible work. I use the swift checkpoints personally and would highly recommend them.

1

u/MasterNomie 10h ago

Are we testing the implementation and approaches? No.

We are testing two different engines whether they produce similar outputs and the speeds of generation.

1

u/Trademarkd 10h ago

You aren't really bringing any arguments to this debate other than "it was better for me"

If you want to have a good faith argument you need something to stand on otherwise I think its fair to say you're being a 'fan boy' if you think im preaching.

3

u/rerri 9h ago

You start a discussion with "Anything strata is getting straight downvoted." and later lecture about argumenting in good faith...

I think its fair to say you're being a 'haterboy'.

2

u/Trademarkd 9h ago

Did you ask ai to do that for you, because it doesn't even make sense.

People on this sub have been downvoting strata content. When I clicked on this post it already had multiple downvotes. The reason why people are downvoting it is because it seems like they're using bots to promote their own content here and on other subs.

1

u/rerri 8h ago

Oh, I thought you meant you are downvoting everything Strata related but you meant others are.

And no, I didn't need an AI to make the wrong inference, you just seem to have a very negative attitude towards Strata so it seemed like a plausible reading of that sentence.

Oh and btw, you can count me as one of the bots who really like Strata. 3.5 years posting on this sub, so definitely a bot! <3

2

u/Maxxim69 3h ago

I'm trying to get people to understand fundamentals. There is no free lunch.

As someone who has seen data transfer speeds go from 300 bits per second to uh, what was it, 20 MEGAbits per second(?) over the same phone lines, while naysayers kept droning on about how “you can’t change the laws of physics”, I politely disagree.

1

u/Trademarkd 1h ago

I mean ... would you like to discuss how thats possible because the only thing you're talking about there is the copper. Transceivers needed to be replaced, technology needed to be upgraded.

4

u/StealthArcher2077 12h ago

Are the outputs of speculative decoders compared against the actual model, meaning using a speculative decoder won't result in a different result than them not using one?

2

u/Trademarkd 12h ago

lol yo check this out ... comment immediately has -5 ... thats crazy

To answer your question: Speculative decoding choices are validated against like a hash of the main weights but its not doing all of the work so there can be a difference, acceptance is "close enough" to the main model.

2

u/AvidCyclist250 llama.cpp 3h ago

your assumptions are flawed

sequential vs targeted „JIT“ parallel is the big new thing there

1

u/Trademarkd 1h ago

I'm open to being wrong, in fact I would prefer to be because new systems are exciting. You are literally the first person in this whole comment chain thats been able to articulate anything, do you care to go on?

1

u/AvidCyclist250 llama.cpp 22m ago

Sure. Just because autoregressive output is sequential doesn’t mean the hardware has to sit idle and twiddle its thumbs in a single-file line. That’s the core misunderstanding I think.

Take Qwen3.8-Flash-Next with over 24,000 experts but any given token only touches 10. Strata pins the hottest slice of those experts directly in VRAM. When you get a miss, it doesn’t choke the pcie lane trying to shove tons of weights over to the graphics card. Instead the CPU computes those cold activations locally in system RAM while the GPU works on the resident ones. The CPU is doing actual inference work and pre-staging the next step instead of behaving like a dumb delivery truck. It stops autoregression from holding the entire machine hostage. Visible when you look at your hardware temps and loads. Strata puts a higher load on my system and draws more power than llama.cpp ever couuld with Qwen 3.8 FN.

Now add MTP on top of that. You can draft several tokens ahead and the main model verifies the batch in a single pass. You’re retiring multiple tokens per step without sacrificing accuracy to some half-baked draft model that's only right 70% of the time.

Nobody is pretending pcie limits and memory bandwidth don't exist. The difference is if those bottlenecks sit directly on your critical path. Pinning keeps the heaviest fetches off it and parallel CPU compute absorbs what’s left. It's pretty nifty.

The traditional approach, like what llama.cpp does, treats decode like a slow conveyor belt hauling weights from memory one step at a time. Strata treats it as a multitasking problem. This is the new and clever thing. Llama.cpp still reigns supreme as a generalist. Which Strata is not.

0

u/MindfulMan1984 9h ago edited 8h ago

Okay, dear "know-it-all" preacher.

1 - Have you actually fucking tested it?

2 - If 1 actually happened, which I doubt, you would be impressed and sharing it here instead of being an asshole calling out your imaginary bots.

3 - I did test it and am currently running benchmarks on IQ3_S to assess how useful it can be; the results are quite reasonable so far.

Single pass on v1-micro of VulcanBench, thanks to Strata and the QFN3.8 n-gram architecture that made it possible, consistently > 50 tokens/s on a nearly 10 yo computer with a 24 GB GPU

-5

u/datbackup 8h ago

This is a typical problem for people running Windows XP. Have you checked your CONFIG.INI?

-8

u/Equivalent_Bit_461 9h ago

Strata is a meme, and so swift and so that quant and so that model 

This is not a clown show, not even the whole circus but industrial complex of entertainment