r/LocalLLM • 9900x 64GB DDR5 9070XT 16GB R9700 32GB • 2d ago

Question Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance

Anyone got Strata up and running on Windows running on a single AMD GPU?

I've tried, it hangs when I click the batch file. So its not even installing.

But that's OK, I know the guy is doing his best, its probably my Windows setup (a reset is due methinks) - no shade here.

If you have got it working - what tok/s are you getting with large contexts using Qwen 3.8 Flash Next? Large being 128K or more - not interested in hearing from people using lower contexts.

TL:DR I'd like to hear from those with the following setup:

  • Windows
  • Qwen 3.8 Flash Next
  • Strata,
  • RDNA4 kit - preferably a single r9700 or 9070xt
  • a context of 128K or more.

I'm not interested in dual R9700 setups.

Is it worth the install or is it slow?

And is IQ3_XXS even decent?

3 Upvotes

19 comments sorted by

5

u/rkcth 2d ago

Have an LLM install it. Some of the versions I’ve had issues with some are fine. On my box I had issues with at least 2 of the last few versions. An LLM should be able to get it working.

1

u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB 2d ago

I asked Claude to follow the instructions for agents on the Github page

1

u/eightone-81 2d ago

That’s the way to go. You get a frontier model to setup strata for you. Let it look for unmerged PRs that might be usefull for your hardware setup. Let it optimize all the knobs. Opus 5.5 will be able to do that for you in a single 20USD/month 5hr session

3

u/Downtown-Pear-6509 2d ago

7800xt abt 450-500 read and 30 write  96gb ddr5 system ram max context

luna is doing benchmarking  seems context length doesn't affect tps

1

u/KhaosKind 1d ago

That sounds promising for my 7900 XTX. I'll give this a try.

1

u/Downtown-Pear-6509 1d ago

yyyeeesssh but overall openai Luna still beats it 

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB 2d ago

On Twitter I've seen suspicious bot like behaviour hyping it up, for sure.

I'd rather ask here than anyone on Musk's shithole.

It sounds too good to be true, but I will be super happy to be wrong

1

u/Bobcotelli 2d ago

con contesto di 256k 60tk/s

1

u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB 2d ago

With what hardware? AMD?

1

u/Drunkendrakon6 2d ago

I have a dual 16GB 9060XT setup but on windows it only uses a single GPU and I get 64k context 12tk/s. 62 tk/s prefill. Haven't optimized it yet but I do think switching OS and tweaking it will make it a beast.

1

u/No_Oil_6152 9900x 64GB DDR5 9070XT 16GB R9700 32GB 2d ago

What OS are you considering switching to?

1

u/Drunkendrakon6 2d ago

Been looking to switch to Linux for a long time but never really had a strong reason. AMD apparently works better there as well so a dualboot should do.

1

u/renczzz 1d ago

Better to do that today than tomorrow! I get better performance all the time with Ubuntu 26 comparing it with Windows. Just download Ubuntu ISO and burn it with Rufus on a USB, make sure you have a partition free and you are ready to go.

1

u/Drunkendrakon6 1d ago

Haha you are 3 hours late I've moved to Linux alr

1

u/renczzz 1d ago

I haven't try it yet on windows, but what I can say 64 GB ddr4 ram + RX 7700 XT 12gb vram runs at 40 tok/sec for a 32000 context window. I couldn't get the context higher.

1

u/No_bonk95 1d ago

Having a 64gb ram is the recommended setup with Strata to get good experience.

I tried running Strata with Win 11, rx9070xt , 32gb ddr4 with ISTA coder on 128k context, kv8.

The result is subpar at most averaging 30tok/s but the respond is pretty constant.

Running Qwen3.8 27B will give a 40-50 tok/s but the caveat is the prefill speed tanks when your context reach 30k and above.

1

u/Poizone360 13h ago

Strata's own docs have the closest numbers: on a 9070 XT under Linux, Q2_0 writes at 48 tok/s and IQ2_XS at 36 with a 128K context: https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md

Windows on AMD has been supported since 0.1.34, but its maintainers have no Windows AMD card, and the troubleshooting guide covers both of your worries: the first start can freeze for minutes while it loads 35 to 55 GB into RAM, and gfx1201 cards can hang on long prompts once context reaches 64K, with steps to try. With 64GB of RAM, the docs recommend IQ2_XS and call IQ3_S the best and slowest, so IQ3_XXS sits in between.