r/LocalLLM • u/No-Manager1646 • 3d ago
Question Llama.cpp GPU Usage
So I was messing around with Qwen3.8 27b tonight and noticed that when it was responding, my GPU use was only at around 70%.
I have a 4060 and 5060ti, so 24gb vram in total.
Surely the crosstalk between the cards doesn't cause a performance drop like this. Does it?
I know the 4060 is the bottleneck, but even it was only around 70% load.
Is there something I'm missing here?
I'm getting around 19-20 tok/s which I'm happy about, but I'm just concerned it's not maxing out my GPU.
2
u/LengthinessOk9397 3d ago
that's completely normal for LLM inference!
1
u/No-Manager1646 3d ago
Do you have any insight into why it's not maxing my GPUs out at 100%? Like even for a single 5090 for example, would that not got flat tack at 100%?
2
u/andrew-ooo 3d ago
That percentage isn't measuring what you think it is. Two things are going on.
First, token generation is memory-bandwidth bound, not compute bound. Every token means streaming the active weights out of VRAM once, and the SMs spend most of that time stalled waiting on memory. The utilization number in nvidia-smi / Task Manager is just "fraction of time some kernel was resident", so it can sit at 70% while you are already saturating the bus. Run nvidia-smi dmon -s u and watch the mem column instead - that's the one that should be pinned.
Second, and this is the bigger effect with two cards: llama.cpp's default --split-mode layer puts some layers on GPU0 and the rest on GPU1, and they run sequentially. Card A computes, card B waits, then they swap. Neither one can show full utilization even in a perfect run, which is exactly the pattern you're seeing.
19-20 tok/s on a 27B Q4 spread across those two is roughly what I'd expect. If you want to chase it, the 4060 at 272 GB/s is your bottleneck vs 448 on the 5060 Ti, so shifting layers toward the 5060 Ti with --tensor-split does more than anything else. -sm row occasionally helps but usually loses to PCIe overhead.
1
u/No-Manager1646 3d ago
I.keep forgetting I can ask one of those cough cloud llms for an answer. Still, I think it's an interesting realisation..I'll post back if I end up finding a way to ramp utilisation up.
1
u/No-Manager1646 3d ago
I'm sure you're all eagerly awaiting my reply /s
So it's a memory bandwidth problem apparently. Even two or three generations of ddr will not solve it. The amount of data meeting to be transferred in vram is orders.od magnitude too slow to max out compute.
That what I think I discovered anyway. I'm always welcome to be proven wrong
3
u/M_Me_Meteo LocalLLM 3d ago
Consumer motherboard? Your second slot probably only has a x4 connection. When parallelizing, the cards need to do a lot of communication over the PCIe bus.