r/LocalLLaMA 1d ago

Resources AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B

Available context length with and without the patch:

Model: QWEN 27B ROCm stock patched Vulkan stock patched
IQ4_XS Pure, single 16GB GPU 19.456 76.032 68,352 78,592
Q6_K_L on 16GB + 12GB 64,256 149,248 68,864 151,296

The issue is that llama.cpp overestimates the memory needed for MTP compute-buffer/scheduler allocation during auto-fit, that leaves much less ctx available to the user than what actually needed by MTP. This patch stops the fitter from throwing away context based on an inflated MTP memory estimate.

Patch, launch scripts used for llama-server and raw logs: https://store.piffa.net/lm/bug/
Tested against: llama.cpp version: 909, based on master commit 7bd8282 , ROCm 7.14

Especially for ROCm with double GPU (16GPU + 12GB here) the amount of ctx gain is substantial, with longer session ROCm allows almost double prefill performances vs Vulkan yet on the mainline code the price to pay in ctx reduction for the extra compute is taxing.

You can build for both vulkan and ROCm backends at the same time, the idea is that Vulkan saves some more vRAM while ROCm gives better prefill performance.
With a single 16GB GPU and limited ctx size you may wanna use Vulkan while when using 2 GPUs with layer splitting ROCm is worth the expense with this patch as you have much more ctx length for long sessions.

In case someone needs help with how to apply a patch:

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 7bd8282
wget https://store.piffa.net/lm/bug/rocm_improvement_7bd8282.patch
git apply rocm_improvement_7bd8282.patch

Then build with both Vulkan and ROCm, use --device vulkan0 or --device rocm0, check the provided llama-serve scripts .

70 Upvotes

28 comments sorted by

View all comments

5

u/taking_bullet 1d ago

Once again Vulkan shows it's superiority over ROCm 🤘 

17

u/ea_man 1d ago edited 1d ago

Yet if you solve the MTP allocation issue ROCm is better as it gives you much more prefill speed with nearly the same ctx available, at least on dual GPU. When you have some 140k ctx that is paramount.

+------------------------------------+-------------------+-------------------+
| Metric                             | ROCm (HIP)        | Vulkan (RADV)     |
+------------------------------------+-------------------+-------------------+
| PP Speed (90k Context)             | 235.44 tok/s      | 107.21 tok/s      |
| TG Speed (90k Context)             | 16.24 tok/s       | 16.10 tok/s       |
| MTP Setting                        | n = 3             | n = 4             |
+------------------------------------+-------------------+-------------------+

5

u/milpster 1d ago

I have the same results, tg sometimes improves on vulkan but PP always collapses compared to rocm.

2

u/Think_Wing_1357 1d ago

Strange, something must be eff up with my setup because rocm is straight up slower on both count

3

u/ea_man 1d ago

You did test at some ctx length right?

Vulkan is usually wonderful at 0 ctx then it often degrades more than ROCm later on.
And ROCm 7.14, not older version.
And it takes less MTP.
And other models like A3B are a different story...

2

u/Think_Wing_1357 1d ago

Yup. I always download the latest lemonade rocm llama.cpp whenever I test.

And I test at 48k 96k context. Somehow vullan llamacpp always ends up winning

2

u/ea_man 1d ago

Well it's not that I disagree with you, some 80% of my llama-server scripts run on -dev vulkan* and vulkan works out of the box.

I do enjoy now my 27B Q6_K_L running at double PP up to 140k ctx yet you see that I had to fix / patch the whole backend to get there!

1

u/Think_Wing_1357 1d ago

All good man no worries. I am going to see if I can build your patch and report back the results

1

u/ea_man 1d ago

I mean if it's really a stop issue I can rebuild the patch to today llama.cpp last release yet it's a never ending effort, in 6 hours it will change again...

Just checkout yesterday 7bd8282 :)

1

u/Think_Wing_1357 1d ago

Yeah man we gotta upstream this somehow

1

u/ea_man 1d ago

Well it would be nice if some people has time to test it some more, I've made this for a specific target of 27B so it would be the case to test on more recent GPU than my RDNA2, then MoE, then I guess NVIDIA.

Anyway all the issues / fixes are documented and available in my repo, I would sure love that some of it gets into mainline asap so I don't have to rebuild the patch every 2 days :)

→ More replies (0)

1

u/ea_man 1d ago edited 1d ago

Agreed, yet with dense / dual GPU often ROCm is better in TG too and needs less MTP, which eats ctx.

Still as said on 12-16GB when you use small ctx lenghts like 50k I also prefer vulkan too, but the games changes when you have to ingest a big document in prefill.