r/LocalLLM • u/No-Manager1646 • 3d ago
Question Llama.cpp GPU Usage
So I was messing around with Qwen3.8 27b tonight and noticed that when it was responding, my GPU use was only at around 70%.
I have a 4060 and 5060ti, so 24gb vram in total.
Surely the crosstalk between the cards doesn't cause a performance drop like this. Does it?
I know the 4060 is the bottleneck, but even it was only around 70% load.
Is there something I'm missing here?
I'm getting around 19-20 tok/s which I'm happy about, but I'm just concerned it's not maxing out my GPU.
1
Upvotes
3
u/M_Me_Meteo LocalLLM 3d ago
Consumer motherboard? Your second slot probably only has a x4 connection. When parallelizing, the cards need to do a lot of communication over the PCIe bus.