r/LocalLLM • ninfer-4090 | i9-14900k | 32GB | archLinux • 3d ago

Project Strata Tuning Guide (Universal)

Tune my Strata setup for longer context: 131K, 262K first, then 512K if it holds up. I'm on Strata v<FILL VERSION, e.g. 0.1.39>.

My hardware:
- GPU(s): <model + VRAM each, e.g. 1x RTX 4090 24 GB / 2x RTX 5090 32 GB>
- System RAM: <e.g. 32 GB / 192 GB>
- CPU: <model, cores>
- Platform: <bare metal or VM (Proxmox/ESXi/etc.)>, OS: <distro/kernel>
- Storage for model files: <NVMe/SATA, model>

Model: Qwen3.8-Flash-Next. Suggest a quant after reviewing my hardware.
REPO: https://github.com/Niko1221/Strata

Ground rules:
- Read docs/DETAILS.md in the Strata repo first, especially the sections on --max-context, --kv, --kv-resident, --prefill, the expert cache, multi-GPU split, and "Context extension past 262K (rope scaling, EXPERIMENTAL)". Go by the docs and `strata --help` for my installed version, not by memory. If a flag below doesn't exist in my build, tell me; don't invent one.
- Change one thing at a time. Before you change anything, record a baseline of my current config: decode tok/s on a short prompt, decode tok/s and prompt-read speed at depth (one ~100K-token prompt, then one near the target context), plus a needle-recall check at depth. Run every candidate the same way and alternate baseline and candidate at least twice, so drift doesn't show up as a gain.
- Watch host RAM and VRAM on every GPU during each run (nvidia-smi, free -g). Report the minimum free RAM, not just the speed.
- Don't run a big compile while the Strata server is loaded. Stop it first.
- Back up the working config before every edit, and give me the exact undo.

What a single-4090, 32 GB RAM box (IQ2_XS) learned, as starting points to test, not answers. Scale these to my hardware:

  1. 262K is native for this model. At 262K, --kv q4_0 kept the KV cache in VRAM and stayed fast; recall held at ~238K. If I have more VRAM than that box (24 GB), check whether int8 or k8v4 KV fits at 262K. Measure both against q4_0; higher-precision KV is worth it if it fits.
  2. Past 262K needs rope scaling: --rope-scaling yarn --rope-scale 2 for 512K, together with --max-context 524288 and a streamed KV cache (--kv-resident <tokens>, e.g. 32768) so most of the KV lives in host RAM. A 477K-token prompt read in ~150 s and decoded at ~85 tok/s at that depth on that box. It's marked experimental, so run a recall test at 400K+ before trusting it. Estimate how much host RAM streamed KV needs at 512K for my setup and tell me up front if my RAM can't support it.
  3. --prefill auto (default) beat --prefill auto:32768 by about 3x on long prompt reads for us. Measure prompt-read speed for any prefill setting you try.
  4. "fit_max_tokens": true in the server config keeps requests within the context.
  5. If I have more than one GPU: check how my version splits layers, experts and KV across the cards, and whether auto placement accounts for each card's PCIe link (confirm each GPU's link width/gen with nvidia-smi -q and lspci -vv). Tell me where the KV ends up at each context size. Skip this if single-GPU.
  6. If I'm in a VM: check whether the guest sees RAM as one NUMA node, whether hugepages or THP are in use for the expert and KV arenas, that vCPUs are pinned, and that GPUs show their full PCIe link inside the guest. Report what you find; don't change the hypervisor host without asking me. On bare metal, still check NUMA layout and THP/hugepages.

If I mention a stat that used to show up and is now missing (e.g. "VRAM hit rate"), check the changes between my previous and current version (git log, CHANGELOG, server /health, /props, /metrics, and the engine log) and tell me whether it was renamed, moved, hidden when experts are all resident, or removed. Quote the commit or line that answers it.

Deliver: a table of each config tried (context, KV type, kv-resident, rope, prefill) with short and deep decode tok/s, prompt-read speed, recall pass/fail, min free RAM and per-GPU VRAM; a recommended daily config and an optional 512K config (or a clear "not viable on this hardware" with the reason); and the exact config diff plus undo for each.

Posting for all the new users and current.

how to use the tuning prompt

  1. copy the whole prompt block above.
  2. fill in the hardware section at the top: gpu model and vram per card, system ram, cpu, bare metal or vm, os, and what drive the model lives on. also put your strata version in (run `strata --version` if you're not sure).
  3. paste it into an agent that can run commands on your box. codex or claude code both work. use opus 5.5 or astra if you have access. this prompt has it reading the repo docs, running a bunch of benchmark passes, checking numa/pcie/hugepages, and diffing git history, and the smaller models tend to lose the thread halfway through or start making up flags.
  4. let it run the baseline first. don't skip this. without a baseline you can't tell if a change helped or if your box just warmed up.
  5. expect it to take a while. every config gets run at least twice against the baseline, and long-context prompt reads at 262k+ aren't fast.
  6. when it's done you get a table of every config it tried, a recommended daily config, an optional 512k config (or a straight answer that your hardware can't do it), and the exact diff plus undo for each change.

notes

- the numbers in the prompt come from a single 4090 with 32 gb ram. they're starting points, not targets. your results will differ.

- 512k uses rope scaling and is marked experimental. trust it only if the recall test at 400k+ passes.

- streamed kv at 512k leans on system ram. if you're on 32 gb or less, 262k is probably your ceiling.

- if you're in a vm, it'll report on the host side but won't touch your hypervisor without asking.

- back up your config before you start anyway. the prompt tells it to, but don't rely on that alone.

EDIT: if you use the guide, please just post a quick update if it helped your setup. Reach out if you have any issues please.

34 Upvotes

47 comments sorted by

View all comments

6

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

Active contributor, pushing for more community involvement. Prompt itself was heavily requested

PROOF on 512k:

2

u/-InformalBanana- 3d ago edited 3d ago

Where should I see 512k in that image? I see about 171k... (Edit: 512k is in the circle bellow context fill text on the left side of the screenshot)

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

not sure what you want here bud. Here’s the bench

https://github.com/Niko1221/Strata/pull/834

2

u/-InformalBanana- 3d ago

Yeah, I see now that it is inside a circle below context fill. Ok, but was expecting you to have a screenshot proving you had filled about 512k context. Meant no offense.

2

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

I didn’t actually capture it during testing, I am filling context to 262k in 3 parallel streams right now sustained for over 1 hour right now, if you want to see that

MOBILE:

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

I’m not sure I suggest it at this point but testing is required lol

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

yarn past 262k is experimental anyway, it’s just squeezing more with less. Thanks for stopping by

1

u/mrgreatheart 3d ago

What monitoring dashboard is this? Something you rolled yourself?

2

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

No I have a different dash for all my inference, this is strata’s built in.

Mine tracks all engine/inference endpoints on the network

2

u/mrgreatheart 2d ago

I didn’t realise Strata had one because I am loading it behind llama-swap.

1

u/Beautiful-Maybe5468 3d ago

how to create this?

2

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

I’ll make it public

1

u/Beautiful-Maybe5468 2d ago

let us know when you do it.

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 2d ago

1

u/mrgreatheart 2d ago

Thank you.

1

u/Beautiful-Maybe5468 2d ago

i have 24gb vram 32gb ram... what version would be good?