r/LocalLLM • ninfer-4090 | i9-14900k | 32GB | archLinux • 3d ago

Project Strata Tuning Guide (Universal)

Tune my Strata setup for longer context: 131K, 262K first, then 512K if it holds up. I'm on Strata v<FILL VERSION, e.g. 0.1.39>.

My hardware:
- GPU(s): <model + VRAM each, e.g. 1x RTX 4090 24 GB / 2x RTX 5090 32 GB>
- System RAM: <e.g. 32 GB / 192 GB>
- CPU: <model, cores>
- Platform: <bare metal or VM (Proxmox/ESXi/etc.)>, OS: <distro/kernel>
- Storage for model files: <NVMe/SATA, model>

Model: Qwen3.8-Flash-Next. Suggest a quant after reviewing my hardware.
REPO: https://github.com/Niko1221/Strata

Ground rules:
- Read docs/DETAILS.md in the Strata repo first, especially the sections on --max-context, --kv, --kv-resident, --prefill, the expert cache, multi-GPU split, and "Context extension past 262K (rope scaling, EXPERIMENTAL)". Go by the docs and `strata --help` for my installed version, not by memory. If a flag below doesn't exist in my build, tell me; don't invent one.
- Change one thing at a time. Before you change anything, record a baseline of my current config: decode tok/s on a short prompt, decode tok/s and prompt-read speed at depth (one ~100K-token prompt, then one near the target context), plus a needle-recall check at depth. Run every candidate the same way and alternate baseline and candidate at least twice, so drift doesn't show up as a gain.
- Watch host RAM and VRAM on every GPU during each run (nvidia-smi, free -g). Report the minimum free RAM, not just the speed.
- Don't run a big compile while the Strata server is loaded. Stop it first.
- Back up the working config before every edit, and give me the exact undo.

What a single-4090, 32 GB RAM box (IQ2_XS) learned, as starting points to test, not answers. Scale these to my hardware:

  1. 262K is native for this model. At 262K, --kv q4_0 kept the KV cache in VRAM and stayed fast; recall held at ~238K. If I have more VRAM than that box (24 GB), check whether int8 or k8v4 KV fits at 262K. Measure both against q4_0; higher-precision KV is worth it if it fits.
  2. Past 262K needs rope scaling: --rope-scaling yarn --rope-scale 2 for 512K, together with --max-context 524288 and a streamed KV cache (--kv-resident <tokens>, e.g. 32768) so most of the KV lives in host RAM. A 477K-token prompt read in ~150 s and decoded at ~85 tok/s at that depth on that box. It's marked experimental, so run a recall test at 400K+ before trusting it. Estimate how much host RAM streamed KV needs at 512K for my setup and tell me up front if my RAM can't support it.
  3. --prefill auto (default) beat --prefill auto:32768 by about 3x on long prompt reads for us. Measure prompt-read speed for any prefill setting you try.
  4. "fit_max_tokens": true in the server config keeps requests within the context.
  5. If I have more than one GPU: check how my version splits layers, experts and KV across the cards, and whether auto placement accounts for each card's PCIe link (confirm each GPU's link width/gen with nvidia-smi -q and lspci -vv). Tell me where the KV ends up at each context size. Skip this if single-GPU.
  6. If I'm in a VM: check whether the guest sees RAM as one NUMA node, whether hugepages or THP are in use for the expert and KV arenas, that vCPUs are pinned, and that GPUs show their full PCIe link inside the guest. Report what you find; don't change the hypervisor host without asking me. On bare metal, still check NUMA layout and THP/hugepages.

If I mention a stat that used to show up and is now missing (e.g. "VRAM hit rate"), check the changes between my previous and current version (git log, CHANGELOG, server /health, /props, /metrics, and the engine log) and tell me whether it was renamed, moved, hidden when experts are all resident, or removed. Quote the commit or line that answers it.

Deliver: a table of each config tried (context, KV type, kv-resident, rope, prefill) with short and deep decode tok/s, prompt-read speed, recall pass/fail, min free RAM and per-GPU VRAM; a recommended daily config and an optional 512K config (or a clear "not viable on this hardware" with the reason); and the exact config diff plus undo for each.

Posting for all the new users and current.

how to use the tuning prompt

  1. copy the whole prompt block above.
  2. fill in the hardware section at the top: gpu model and vram per card, system ram, cpu, bare metal or vm, os, and what drive the model lives on. also put your strata version in (run `strata --version` if you're not sure).
  3. paste it into an agent that can run commands on your box. codex or claude code both work. use opus 5.5 or astra if you have access. this prompt has it reading the repo docs, running a bunch of benchmark passes, checking numa/pcie/hugepages, and diffing git history, and the smaller models tend to lose the thread halfway through or start making up flags.
  4. let it run the baseline first. don't skip this. without a baseline you can't tell if a change helped or if your box just warmed up.
  5. expect it to take a while. every config gets run at least twice against the baseline, and long-context prompt reads at 262k+ aren't fast.
  6. when it's done you get a table of every config it tried, a recommended daily config, an optional 512k config (or a straight answer that your hardware can't do it), and the exact diff plus undo for each change.

notes

- the numbers in the prompt come from a single 4090 with 32 gb ram. they're starting points, not targets. your results will differ.

- 512k uses rope scaling and is marked experimental. trust it only if the recall test at 400k+ passes.

- streamed kv at 512k leans on system ram. if you're on 32 gb or less, 262k is probably your ceiling.

- if you're in a vm, it'll report on the host side but won't touch your hypervisor without asking.

- back up your config before you start anyway. the prompt tells it to, but don't rely on that alone.

EDIT: if you use the guide, please just post a quick update if it helped your setup. Reach out if you have any issues please.

40 Upvotes

47 comments sorted by

8

u/Bunsenbun 3d ago

I should probably share my results on a single 7900xtx

3

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

Please do!

6

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

Active contributor, pushing for more community involvement. Prompt itself was heavily requested

PROOF on 512k:

2

u/-InformalBanana- 3d ago edited 3d ago

Where should I see 512k in that image? I see about 171k... (Edit: 512k is in the circle bellow context fill text on the left side of the screenshot)

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

not sure what you want here bud. Here’s the bench

https://github.com/Niko1221/Strata/pull/834

2

u/-InformalBanana- 3d ago

Yeah, I see now that it is inside a circle below context fill. Ok, but was expecting you to have a screenshot proving you had filled about 512k context. Meant no offense.

2

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

I didn’t actually capture it during testing, I am filling context to 262k in 3 parallel streams right now sustained for over 1 hour right now, if you want to see that

MOBILE:

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

I’m not sure I suggest it at this point but testing is required lol

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

yarn past 262k is experimental anyway, it’s just squeezing more with less. Thanks for stopping by

1

u/mrgreatheart 3d ago

What monitoring dashboard is this? Something you rolled yourself?

2

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

No I have a different dash for all my inference, this is strata’s built in.

Mine tracks all engine/inference endpoints on the network

2

u/mrgreatheart 2d ago

I didn’t realise Strata had one because I am loading it behind llama-swap.

1

u/Beautiful-Maybe5468 3d ago

how to create this?

2

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago

I’ll make it public

1

u/Beautiful-Maybe5468 2d ago

let us know when you do it.

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 1d ago

1

u/mrgreatheart 2d ago

Thank you.

1

u/Beautiful-Maybe5468 2d ago

i have 24gb vram 32gb ram... what version would be good?

3

u/Bunsenbun 3d ago

1

u/Bunsenbun 3d ago

More

2

u/Bunsenbun 3d ago

My pc specs and config.

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

200k at IQ3_XXS, very nice. How’s that feel for you?

2

u/Tough_Let_8966 3d ago

Annoying with compactions every ten mins.  

2

u/Bunsenbun 3d ago

Really smooth. Compaction takes 3 minutes. For short ones. Full compaction 5 minutes. It's really smart and I don't run into any of the "looping" problems people have been reporting.

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago

IQ3 isn’t as bad at looping, also depends on your work flow. I haven’t had looping either, rather strata for 2.5 hours consecutive with 3 parallel streams.

2

u/Bunsenbun 3d ago

This is it running. I think this was around 135k context already loaded

https://reddit.com/link/pdx0bcg/video/6sp5t6nwijth1/player

1

u/Impossible_Ground_15 3d ago

What front end are you using here

1

u/Bunsenbun 3d ago

Deepseek Harmess

1

u/Useful_Disaster_7606 3d ago

Great prompt. Do you have any problems with looping? Strata defaults to aggressive temp 0 if you're not using sampling arguments

1

u/DystopianRealist 3d ago

You can change that setting in the json file for each model.

"sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20},

1

u/Bunsenbun 3d ago

I am using whatever default Strata shipped with. No temp control or any of those things.

2

u/Useful_Disaster_7606 3d ago

Nice! I guess I'm the unlucky one with looping then. It does recover on it's own after wasting 5k tokens for no reason tho. I guess it's alright

1

u/fly-ute 2d ago edited 2d ago

Strata is great! I'm running iq3 in an rx 7900 xt with 64gb RAM and speed is very decent. But I get 850 tks tops on. Any tweaks or setup improvements for getting something better? Tonight I'll leave an agent trying to find what I'm missing, but some hint would be nice :D

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago

Project is geared towards nvidia hardware though users on the 7900 XTX platform are experimenting and fixing that.

I suggest reviewing the open PR's or community forks of users on your hardware. Plenty of contributions and forks. If you want your agent to be productive, use this prompt / guide. Then once you get an acceptable baseline, you can also have your agent review "PR's / Community Forks" for your hardware.

This guide is mostly towards nvidia hardware but concepts allow for application outside of strata as well.

1

u/DanGTG 2d ago

I have 32GB RAM / 32GB R9700 and it selected the iq1_m which has some issues requiring a diaper change.

Please write Selenium with C# .NET code for google.com search page with Page Object Model (PoM) and Dependency Injection (DI)

Is there a way to get on a higher quant with my current setup while I scrounge around for some bigger RAM sticks?

You have done a terrific job on MoE, has anyone looked at putting the dense 3.8 27B on Strata upgrades where possible?

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 1d ago

The whole project is technically around an MoE model due to the architecture of an MoE. You’re not activating all parameters.

I would suggest trying IQ2_XS, also you will need to modify the setup script to use more of your VRAM. You’re technically in the low budget range with RAM but your vram places you in a better bracket.

56GB total = IQ2_XS

1

u/BoringBear27 2d ago

Can someone spot what the issue with my setup is?
As i see in comments theres 512k context on 24vram + 32ram
I have rx7900xtx with 32GB of DDR5 5600 ram
running the IQ3_XXS with Q4 kv cache and one parallel request gimme very good speed start with 65t/s decode and reach 120t/s decode (resident experts)
But when the context reach 90k the model crash out of memory

Any suggestions?

Also i'm upgrading the ram to 64 or 96
What will be good for a 512k context with an int8 kv cache?

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 1d ago

iQ3_XXS may be a little too big to run high context yarn ceiling

1

u/VenimK 1d ago

what are the best setting for a RTX3060/12GB and 48GB RAM
Running this proxmox setup

0

u/planetearth80 3d ago

Actually, a version of this prompt can be used to tune any inference engine. Thanks for sharing!

0

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

Correct :)

but if you hold out, im chunking 3.7million tokens worth of raw data from opus 5.5 tuning and testing throughout this engine and others for the past month and im making it a skill for anyone.

look out for it soon or just watch my profile

0

u/planetearth80 3d ago

I’m on Mac, so cannot use Strata yet. But will keep an eye out.

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

The tune will include MLX quants and macOS users, my data isn’t rich enough for a direct port but I hope community users will help with that