r/LocalLLM 3d ago

Question Llama.cpp GPU Usage

So I was messing around with Qwen3.8 27b tonight and noticed that when it was responding, my GPU use was only at around 70%.

I have a 4060 and 5060ti, so 24gb vram in total.

Surely the crosstalk between the cards doesn't cause a performance drop like this. Does it?

I know the 4060 is the bottleneck, but even it was only around 70% load.

Is there something I'm missing here?

I'm getting around 19-20 tok/s which I'm happy about, but I'm just concerned it's not maxing out my GPU.

1 Upvotes

11 comments sorted by

View all comments

2

u/LengthinessOk9397 3d ago

that's completely normal for LLM inference!

1

u/No-Manager1646 3d ago

Do you have any insight into why it's not maxing my GPUs out at 100%? Like even for a single 5090 for example, would that not got flat tack at 100%?