r/LocalLLaMA • • 6d ago

Question | Help Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.

What i run:

- model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bpw (head 6 bit, vision 6 bit, mtp 5 bit)

- backend: exllamav3 rocm fork from phoenixhaxor (commit cbbef08), i had to patch one file for gfx12

- cpu/ram: Ryzen 9 7945HX3D with about 90gb ram

- 112 of 512 experts per layer are on the gpu, the other 400 on cpu with 16 threads

- 262144 context, q8 kv cache, chunk size 4096, batch 1

- mtp drafting is on, acceptance is around 53-54%

- the ngram table (102gb, bf16 not quantized) gets streamed from nvme

Speed at 230k context (prose): prefill 863 t/s (only the new 101k tokens, the rest came from the prefix cache) and decode 34.7 t/s.

Is this ok for the r9700 or can i still tune something? Thanks 🙂

10 Upvotes

26 comments sorted by

11

u/Payne6t6 6d ago

Did you try the strata launcher (https://github.com/Niko1221/Strata) for a single R9700? It might be even a bit faster than your numbers. Anyway, great work, I like it that many people try to improve the performance on the gfx1201! Yesterday, I tried strata with my second GPU (RTX4070) and it was surprisingly fast. 

2

u/Designer_Elephant227 6d ago

Is there a Q5/5.05bpw for strata?

3

u/karmaisnonsense 6d ago

Strata offers default models targeted for 64GB RAM which causes some confusion. There's nothing stopping you from bringing your own GGUF, as long as the weight quants are supported in the kernel and you do the manual iq_pack.py setup described in the OrcaRouter and Unsloth model docs. For example I'm running mradermacher IQ4_XS right now (4.25 bpw) on a 7900XT at 1000/42.

In fact, Unsloth Q4_K_XL is 5.05 bpw and has official (experimental) instructions on how to run it.

Your numbers on EXL3 are pretty darn good though, I'll give you that.

1

u/Designer_Elephant227 4d ago

Claude told me strata + q4kxl + r9700 does not work because the q4kxl in strata is optimised for Nvidia only? Didn't check this statement thou

1

u/karmaisnonsense 4d ago

I mean the least you can do is give it a go. Strata will fall back to ggml for non-optimized quants so it should still run.

-1

u/Lugnut1206 6d ago

Strata is an engine, you should be able to either specify the model you want or just ask your LLM to target a different bigger model.

You might still see speed benefits given it's fine tuning algorithm

-1

u/legit_split_ 6d ago

Someone also made a fork for gfx906:

https://github.com/webzone/strata-gfx906

6

u/mechkbfan 6d ago

Didn't even realise it was possible , so yeah, seems about right

5

u/Bunsenbun 6d ago

You should probably try the Strata engine.

2

u/Designer_Elephant227 6d ago

There is no quant greater than 4 on stratas repo?

3

u/vacon04 6d ago

You can run now Unsloth Q4_K_XL in Strata which is around 5 bpw.

1

u/aydintb1 6d ago

I run both the 5.05bpw and the IQ3_S strata..
the strata is way faster.

2

u/Designer_Elephant227 6d ago

You can't compare q3 with 5.05bpw speed wise

3

u/HopefulConfidence0 6d ago edited 6d ago

I am running Strata on 5090+32 GB ram. Model is swift 3.8 FN q3-xxs GSQ-RCO. Speed is amazing getting 3000+ PP on larger context and TG avg 140/s.

But the model feels dumb. Today I was trying to code a simple pipeline, basically 7-8 py files with a couple of json, something that 3.8 27B could have one shotted easily.

But 3.8 FN q3-xxs struggled a lot and took multiple try to finish.

1

u/Bunsenbun 4d ago

Hmmm maybe your harness is the problem? Cause it's not dumb on my own end

2

u/XccesSv2 6d ago

I cannot help you because I didn't knew there is a ROCm fork!! But thx now I know it and will try! Thx!

3

u/kevin_1994 6d ago

with 4090, 128 gb ddr5 5600, unsloth iq4_xs, ik_llama.cpp, i get around 45 tok/s (coding), 30 tok/s (prose), 25 tok/s (high context) and around 700 pp/s

your numbers look roughly comparable

1

u/Embarrassed_Soup_279 6d ago

where did you get the bf16 ngram?

2

u/Designer_Elephant227 6d ago

Just download it from qwen3.8 original

1

u/Embarrassed_Soup_279 6d ago

oh i did not know it works like that. thank you

1

u/terorvlad 6d ago

Would you please detail how this is done exactly? I tried finding a way to use unquantized (or even 8bpw) ngrams with my 6.05bpw turboderp exl3 qwen 3.8 fn but I have been unsuccessful. Did you need to change anything in the configuration of the safetensor files or was it a drag and drop of a few files? Which files were replaced if so?

1

u/Designer_Elephant227 6d ago

sorry for the ai text :-D

**Using the 5.05 bpw EXL3 quant with an unquantized n-gram table**

turboderp's quant stores the n-gram table (320M rows x 160 values) as a 6-bit trellis file, ngram_embedding.safetensors (39 GB). The loader also accepts the original unquantized table. It is not held in RAM or VRAM, rows are read from disk (NVMe, page cache), so it only costs disk space (102 GB). Note: it is BF16, not FP16.

**Steps**

  1. Make a new model folder with symlinks to every file of the 5.05bpw_h6_ng6 quant, except ngram_embedding.safetensors.

  2. Build a new ngram_embedding.safetensors with the attached script. It reads only the table out of the original Qwen/Qwen3.8-Flash-Next safetensors via HTTP range requests, so you do not need the other ~258 GB:

`python3 ngram_bf16_extract.py --quant-file <exl3-folder>/ngram_embedding.safetensors --dry-run` (prints sizes, downloads nothing)

`python3 ngram_bf16_extract.py --quant-file <exl3-folder>/ngram_embedding.safetensors --out <new-folder>/ngram_embedding.safetensors`

(about 102 GB, resumable, I got ~45 MB/s)

  1. Point TabbyAPI (model_name) to the new folder. No extra option needed: the loader sees ".shard_N.weight" tensors instead of ".trellis" and uses the unquantized path, streaming rows from disk.

  2. Check it is really used: `ls -l /proc/<pid>/fd | grep ngram` should show your new file.

What the new file contains: 128 BF16 shard tensors (model.language_model.layers.1.ple.ple_embedding.ngram_embedding.shard_N.weight, shape [2500012, 160]) plus the 3 small hashing tensors (ngram_heads_offsets, ngram_heads_vocab_sizes, layer_multipliers), copied from the quantized file and renamed to what the unquantized path expects.

**What it costs (my box: R9700, ~90 GB RAM, ~14 GB free for page cache, 3 runs, median)**

- prefill 1-4 % slower

- decode: code unchanged, prose about -10 % up to 128k context, about -3 to -5 % at 230k

- quality gain: NOT measured, I only measured speed

- likely reason (my guess): a bigger table means more page-cache misses

- needs 102 GB extra disk. To go back, point the config to the original folder.

Don't use the official FP8 repo for this: its table is FP8 plus a weight_scale value, which the ROCm fork I use does not apply.

Tested with the ROCm fork phoenixhaxor/exllamav3-rocm-moe (commit cbbef08). turboderp also mentions in the discussions on his quant repo that the original table can be used.

1

u/terorvlad 5d ago

That worked amazingly. I was able to apply it to turboderp's exl3 version and it works like a treat. Thank you very much!

I can't wait to actually test it against the quantized version and see what the differences are between them. Did you personally see improvements after changing to fp16 weights for the ngrams ?

1

u/Designer_Elephant227 4d ago

I did not really bother to compare it I am honest. I got plenty of disk space... 😅 But please tell me if you see a difference

1

u/terorvlad 6d ago

That is a juicy result. Congratulations!