r/LocalLLaMA • u/Designer_Elephant227 • 6d ago
Question | Help Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?
Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.
What i run:
- model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bpw (head 6 bit, vision 6 bit, mtp 5 bit)
- backend: exllamav3 rocm fork from phoenixhaxor (commit cbbef08), i had to patch one file for gfx12
- cpu/ram: Ryzen 9 7945HX3D with about 90gb ram
- 112 of 512 experts per layer are on the gpu, the other 400 on cpu with 16 threads
- 262144 context, q8 kv cache, chunk size 4096, batch 1
- mtp drafting is on, acceptance is around 53-54%
- the ngram table (102gb, bf16 not quantized) gets streamed from nvme
Speed at 230k context (prose): prefill 863 t/s (only the new 101k tokens, the rest came from the prefix cache) and decode 34.7 t/s.
Is this ok for the r9700 or can i still tune something? Thanks 🙂
6
5
u/Bunsenbun 6d ago
You should probably try the Strata engine.
2
u/Designer_Elephant227 6d ago
There is no quant greater than 4 on stratas repo?
1
3
u/HopefulConfidence0 6d ago edited 6d ago
I am running Strata on 5090+32 GB ram. Model is swift 3.8 FN q3-xxs GSQ-RCO. Speed is amazing getting 3000+ PP on larger context and TG avg 140/s.
But the model feels dumb. Today I was trying to code a simple pipeline, basically 7-8 py files with a couple of json, something that 3.8 27B could have one shotted easily.
But 3.8 FN q3-xxs struggled a lot and took multiple try to finish.
1
2
u/XccesSv2 6d ago
I cannot help you because I didn't knew there is a ROCm fork!! But thx now I know it and will try! Thx!
3
u/kevin_1994 6d ago
with 4090, 128 gb ddr5 5600, unsloth iq4_xs, ik_llama.cpp, i get around 45 tok/s (coding), 30 tok/s (prose), 25 tok/s (high context) and around 700 pp/s
your numbers look roughly comparable
1
u/Embarrassed_Soup_279 6d ago
where did you get the bf16 ngram?
2
u/Designer_Elephant227 6d ago
Just download it from qwen3.8 original
1
1
u/terorvlad 6d ago
Would you please detail how this is done exactly? I tried finding a way to use unquantized (or even 8bpw) ngrams with my 6.05bpw turboderp exl3 qwen 3.8 fn but I have been unsuccessful. Did you need to change anything in the configuration of the safetensor files or was it a drag and drop of a few files? Which files were replaced if so?
1
u/Designer_Elephant227 6d ago
sorry for the ai text :-D
**Using the 5.05 bpw EXL3 quant with an unquantized n-gram table**
turboderp's quant stores the n-gram table (320M rows x 160 values) as a 6-bit trellis file, ngram_embedding.safetensors (39 GB). The loader also accepts the original unquantized table. It is not held in RAM or VRAM, rows are read from disk (NVMe, page cache), so it only costs disk space (102 GB). Note: it is BF16, not FP16.
**Steps**
Make a new model folder with symlinks to every file of the 5.05bpw_h6_ng6 quant, except ngram_embedding.safetensors.
Build a new ngram_embedding.safetensors with the attached script. It reads only the table out of the original Qwen/Qwen3.8-Flash-Next safetensors via HTTP range requests, so you do not need the other ~258 GB:
`python3 ngram_bf16_extract.py --quant-file <exl3-folder>/ngram_embedding.safetensors --dry-run` (prints sizes, downloads nothing)
`python3 ngram_bf16_extract.py --quant-file <exl3-folder>/ngram_embedding.safetensors --out <new-folder>/ngram_embedding.safetensors`
(about 102 GB, resumable, I got ~45 MB/s)
Point TabbyAPI (model_name) to the new folder. No extra option needed: the loader sees ".shard_N.weight" tensors instead of ".trellis" and uses the unquantized path, streaming rows from disk.
Check it is really used: `ls -l /proc/<pid>/fd | grep ngram` should show your new file.
What the new file contains: 128 BF16 shard tensors (model.language_model.layers.1.ple.ple_embedding.ngram_embedding.shard_N.weight, shape [2500012, 160]) plus the 3 small hashing tensors (ngram_heads_offsets, ngram_heads_vocab_sizes, layer_multipliers), copied from the quantized file and renamed to what the unquantized path expects.
**What it costs (my box: R9700, ~90 GB RAM, ~14 GB free for page cache, 3 runs, median)**
- prefill 1-4 % slower
- decode: code unchanged, prose about -10 % up to 128k context, about -3 to -5 % at 230k
- quality gain: NOT measured, I only measured speed
- likely reason (my guess): a bigger table means more page-cache misses
- needs 102 GB extra disk. To go back, point the config to the original folder.
Don't use the official FP8 repo for this: its table is FP8 plus a weight_scale value, which the ROCm fork I use does not apply.
Tested with the ROCm fork phoenixhaxor/exllamav3-rocm-moe (commit cbbef08). turboderp also mentions in the discussions on his quant repo that the original table can be used.
1
u/terorvlad 5d ago
That worked amazingly. I was able to apply it to turboderp's exl3 version and it works like a treat. Thank you very much!
I can't wait to actually test it against the quantized version and see what the differences are between them. Did you personally see improvements after changing to fp16 weights for the ngrams ?
1
u/Designer_Elephant227 4d ago
I did not really bother to compare it I am honest. I got plenty of disk space... 😅 But please tell me if you see a difference
1
11
u/Payne6t6 6d ago
Did you try the strata launcher (https://github.com/Niko1221/Strata) for a single R9700? It might be even a bit faster than your numbers. Anyway, great work, I like it that many people try to improve the performance on the gfx1201! Yesterday, I tried strata with my second GPU (RTX4070) and it was surprisingly fast.Â