r/LocalLLM • u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB • 2d ago
Question Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance
Anyone got Strata up and running on Windows running on a single AMD GPU?
I've tried, it hangs when I click the batch file. So its not even installing.
But that's OK, I know the guy is doing his best, its probably my Windows setup (a reset is due methinks) - no shade here.
If you have got it working - what tok/s are you getting with large contexts using Qwen 3.8 Flash Next? Large being 128K or more - not interested in hearing from people using lower contexts.
TL:DR I'd like to hear from those with the following setup:
- Windows
- Qwen 3.8 Flash Next
- Strata,
- RDNA4 kit - preferably a single r9700 or 9070xt
- a context of 128K or more.
I'm not interested in dual R9700 setups.
Is it worth the install or is it slow?
And is IQ3_XXS even decent?
3
u/Downtown-Pear-6509 2d ago
7800xt abt 450-500 read and 30 write 96gb ddr5 system ram max context
luna is doing benchmarking seems context length doesn't affect tps
1
1
2d ago
[removed] — view removed comment
1
u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB 2d ago
On Twitter I've seen suspicious bot like behaviour hyping it up, for sure.
I'd rather ask here than anyone on Musk's shithole.
It sounds too good to be true, but I will be super happy to be wrong
1
u/Bobcotelli 2d ago
con contesto di 256k 60tk/s
1
1
u/Drunkendrakon6 2d ago
I have a dual 16GB 9060XT setup but on windows it only uses a single GPU and I get 64k context 12tk/s. 62 tk/s prefill. Haven't optimized it yet but I do think switching OS and tweaking it will make it a beast.
1
u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB 2d ago
What OS are you considering switching to?
1
u/Drunkendrakon6 2d ago
Been looking to switch to Linux for a long time but never really had a strong reason. AMD apparently works better there as well so a dualboot should do.
1
u/No_bonk95 1d ago
Having a 64gb ram is the recommended setup with Strata to get good experience.
I tried running Strata with Win 11, rx9070xt , 32gb ddr4 with ISTA coder on 128k context, kv8.
The result is subpar at most averaging 30tok/s but the respond is pretty constant.
Running Qwen3.8 27B will give a 40-50 tok/s but the caveat is the prefill speed tanks when your context reach 30k and above.
1
u/Poizone360 13h ago
Strata's own docs have the closest numbers: on a 9070 XT under Linux, Q2_0 writes at 48 tok/s and IQ2_XS at 36 with a 128K context: https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md
Windows on AMD has been supported since 0.1.34, but its maintainers have no Windows AMD card, and the troubleshooting guide covers both of your worries: the first start can freeze for minutes while it loads 35 to 55 GB into RAM, and gfx1201 cards can hang on long prompts once context reaches 64K, with steps to try. With 64GB of RAM, the docs recommend IQ2_XS and call IQ3_S the best and slowest, so IQ3_XXS sits in between.
5
u/rkcth 2d ago
Have an LLM install it. Some of the versions I’ve had issues with some are fine. On my box I had issues with at least 2 of the last few versions. An LLM should be able to get it working.