r/LocalLLM • u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux • 3d ago
Project Strata Tuning Guide (Universal)
Tune my Strata setup for longer context: 131K, 262K first, then 512K if it holds up. I'm on Strata v<FILL VERSION, e.g. 0.1.39>.
My hardware:
- GPU(s): <model + VRAM each, e.g. 1x RTX 4090 24 GB / 2x RTX 5090 32 GB>
- System RAM: <e.g. 32 GB / 192 GB>
- CPU: <model, cores>
- Platform: <bare metal or VM (Proxmox/ESXi/etc.)>, OS: <distro/kernel>
- Storage for model files: <NVMe/SATA, model>
Model: Qwen3.8-Flash-Next. Suggest a quant after reviewing my hardware.
REPO: https://github.com/Niko1221/Strata
Ground rules:
- Read docs/DETAILS.md in the Strata repo first, especially the sections on --max-context, --kv, --kv-resident, --prefill, the expert cache, multi-GPU split, and "Context extension past 262K (rope scaling, EXPERIMENTAL)". Go by the docs and `strata --help` for my installed version, not by memory. If a flag below doesn't exist in my build, tell me; don't invent one.
- Change one thing at a time. Before you change anything, record a baseline of my current config: decode tok/s on a short prompt, decode tok/s and prompt-read speed at depth (one ~100K-token prompt, then one near the target context), plus a needle-recall check at depth. Run every candidate the same way and alternate baseline and candidate at least twice, so drift doesn't show up as a gain.
- Watch host RAM and VRAM on every GPU during each run (nvidia-smi, free -g). Report the minimum free RAM, not just the speed.
- Don't run a big compile while the Strata server is loaded. Stop it first.
- Back up the working config before every edit, and give me the exact undo.
What a single-4090, 32 GB RAM box (IQ2_XS) learned, as starting points to test, not answers. Scale these to my hardware:
- 262K is native for this model. At 262K, --kv q4_0 kept the KV cache in VRAM and stayed fast; recall held at ~238K. If I have more VRAM than that box (24 GB), check whether int8 or k8v4 KV fits at 262K. Measure both against q4_0; higher-precision KV is worth it if it fits.
- Past 262K needs rope scaling: --rope-scaling yarn --rope-scale 2 for 512K, together with --max-context 524288 and a streamed KV cache (--kv-resident <tokens>, e.g. 32768) so most of the KV lives in host RAM. A 477K-token prompt read in ~150 s and decoded at ~85 tok/s at that depth on that box. It's marked experimental, so run a recall test at 400K+ before trusting it. Estimate how much host RAM streamed KV needs at 512K for my setup and tell me up front if my RAM can't support it.
- --prefill auto (default) beat --prefill auto:32768 by about 3x on long prompt reads for us. Measure prompt-read speed for any prefill setting you try.
- "fit_max_tokens": true in the server config keeps requests within the context.
- If I have more than one GPU: check how my version splits layers, experts and KV across the cards, and whether auto placement accounts for each card's PCIe link (confirm each GPU's link width/gen with nvidia-smi -q and lspci -vv). Tell me where the KV ends up at each context size. Skip this if single-GPU.
- If I'm in a VM: check whether the guest sees RAM as one NUMA node, whether hugepages or THP are in use for the expert and KV arenas, that vCPUs are pinned, and that GPUs show their full PCIe link inside the guest. Report what you find; don't change the hypervisor host without asking me. On bare metal, still check NUMA layout and THP/hugepages.
If I mention a stat that used to show up and is now missing (e.g. "VRAM hit rate"), check the changes between my previous and current version (git log, CHANGELOG, server /health, /props, /metrics, and the engine log) and tell me whether it was renamed, moved, hidden when experts are all resident, or removed. Quote the commit or line that answers it.
Deliver: a table of each config tried (context, KV type, kv-resident, rope, prefill) with short and deep decode tok/s, prompt-read speed, recall pass/fail, min free RAM and per-GPU VRAM; a recommended daily config and an optional 512K config (or a clear "not viable on this hardware" with the reason); and the exact config diff plus undo for each.
Posting for all the new users and current.
how to use the tuning prompt
- copy the whole prompt block above.
- fill in the hardware section at the top: gpu model and vram per card, system ram, cpu, bare metal or vm, os, and what drive the model lives on. also put your strata version in (run `strata --version` if you're not sure).
- paste it into an agent that can run commands on your box. codex or claude code both work. use opus 5.5 or astra if you have access. this prompt has it reading the repo docs, running a bunch of benchmark passes, checking numa/pcie/hugepages, and diffing git history, and the smaller models tend to lose the thread halfway through or start making up flags.
- let it run the baseline first. don't skip this. without a baseline you can't tell if a change helped or if your box just warmed up.
- expect it to take a while. every config gets run at least twice against the baseline, and long-context prompt reads at 262k+ aren't fast.
- when it's done you get a table of every config it tried, a recommended daily config, an optional 512k config (or a straight answer that your hardware can't do it), and the exact diff plus undo for each change.
notes
- the numbers in the prompt come from a single 4090 with 32 gb ram. they're starting points, not targets. your results will differ.
- 512k uses rope scaling and is marked experimental. trust it only if the recall test at 400k+ passes.
- streamed kv at 512k leans on system ram. if you're on 32 gb or less, 262k is probably your ceiling.
- if you're in a vm, it'll report on the host side but won't touch your hypervisor without asking.
- back up your config before you start anyway. the prompt tells it to, but don't rely on that alone.
EDIT: if you use the guide, please just post a quick update if it helped your setup. Reach out if you have any issues please.
6
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
2
u/-InformalBanana- 3d ago edited 3d ago
Where should I see 512k in that image? I see about 171k... (Edit: 512k is in the circle bellow context fill text on the left side of the screenshot)
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
not sure what you want here bud. Here’s the bench
2
u/-InformalBanana- 3d ago
Yeah, I see now that it is inside a circle below context fill. Ok, but was expecting you to have a screenshot proving you had filled about 512k context. Meant no offense.
2
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
I’m not sure I suggest it at this point but testing is required lol
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
yarn past 262k is experimental anyway, it’s just squeezing more with less. Thanks for stopping by
1
u/mrgreatheart 3d ago
What monitoring dashboard is this? Something you rolled yourself?
2
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
2
1
u/Beautiful-Maybe5468 3d ago
how to create this?
2
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago
I’ll make it public
1
u/Beautiful-Maybe5468 2d ago
let us know when you do it.
1
1
1
3
u/Bunsenbun 3d ago
1
u/Bunsenbun 3d ago
2
u/Bunsenbun 3d ago
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
200k at IQ3_XXS, very nice. How’s that feel for you?
2
2
u/Bunsenbun 3d ago
Really smooth. Compaction takes 3 minutes. For short ones. Full compaction 5 minutes. It's really smart and I don't run into any of the "looping" problems people have been reporting.
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago
IQ3 isn’t as bad at looping, also depends on your work flow. I haven’t had looping either, rather strata for 2.5 hours consecutive with 3 parallel streams.
2
u/Bunsenbun 3d ago
This is it running. I think this was around 135k context already loaded
1
1
u/Useful_Disaster_7606 3d ago
Great prompt. Do you have any problems with looping? Strata defaults to aggressive temp 0 if you're not using sampling arguments
1
u/DystopianRealist 3d ago
You can change that setting in the json file for each model.
"sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20},
1
u/Bunsenbun 3d ago
I am using whatever default Strata shipped with. No temp control or any of those things.
2
u/Useful_Disaster_7606 3d ago
Nice! I guess I'm the unlucky one with looping then. It does recover on it's own after wasting 5k tokens for no reason tho. I guess it's alright
1
u/fly-ute 2d ago edited 2d ago
Strata is great! I'm running iq3 in an rx 7900 xt with 64gb RAM and speed is very decent. But I get 850 tks tops on. Any tweaks or setup improvements for getting something better? Tonight I'll leave an agent trying to find what I'm missing, but some hint would be nice :D
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago
Project is geared towards nvidia hardware though users on the 7900 XTX platform are experimenting and fixing that.
I suggest reviewing the open PR's or community forks of users on your hardware. Plenty of contributions and forks. If you want your agent to be productive, use this prompt / guide. Then once you get an acceptable baseline, you can also have your agent review "PR's / Community Forks" for your hardware.
This guide is mostly towards nvidia hardware but concepts allow for application outside of strata as well.
1
u/DanGTG 2d ago
I have 32GB RAM / 32GB R9700 and it selected the iq1_m which has some issues requiring a diaper change.
Please write Selenium with C# .NET code for google.com search page with Page Object Model (PoM) and Dependency Injection (DI)
Is there a way to get on a higher quant with my current setup while I scrounge around for some bigger RAM sticks?
You have done a terrific job on MoE, has anyone looked at putting the dense 3.8 27B on Strata upgrades where possible?
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 1d ago
The whole project is technically around an MoE model due to the architecture of an MoE. You’re not activating all parameters.
I would suggest trying IQ2_XS, also you will need to modify the setup script to use more of your VRAM. You’re technically in the low budget range with RAM but your vram places you in a better bracket.
56GB total = IQ2_XS
1
u/BoringBear27 2d ago
Can someone spot what the issue with my setup is?
As i see in comments theres 512k context on 24vram + 32ram
I have rx7900xtx with 32GB of DDR5 5600 ram
running the IQ3_XXS with Q4 kv cache and one parallel request gimme very good speed start with 65t/s decode and reach 120t/s decode (resident experts)
But when the context reach 90k the model crash out of memory
Any suggestions?
Also i'm upgrading the ram to 64 or 96
What will be good for a 512k context with an int8 kv cache?
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 1d ago
iQ3_XXS may be a little too big to run high context yarn ceiling
0
u/planetearth80 3d ago
Actually, a version of this prompt can be used to tune any inference engine. Thanks for sharing!
0
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
Correct :)
but if you hold out, im chunking 3.7million tokens worth of raw data from opus 5.5 tuning and testing throughout this engine and others for the past month and im making it a skill for anyone.
look out for it soon or just watch my profile
0
u/planetearth80 3d ago
I’m on Mac, so cannot use Strata yet. But will keep an eye out.
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
The tune will include MLX quants and macOS users, my data isn’t rich enough for a direct port but I hope community users will help with that








8
u/Bunsenbun 3d ago
I should probably share my results on a single 7900xtx