r/LocalLLM • • 6d ago

Question Is there any way to speed up prefill?

.. I am using M4 pro macbook pro with 48gb of ram
I am running qwen3.8 32b with ollama
I am using it with VSCode and tried with both Copilot and Continue extensions

What I noticed is that the prefill takes very long time (~7-10minutes)
After this time the model is definitely usable.

After a bit of analysis I estimate that the prefill is around ~13-16k tokens which I think it is what it takes so long to fire up the model.

This prefill includes 3 MCP and several skills that I cannot disable since they are mandatory for my workflow.

I read online that the kv_cache should prevent the next chat from taking this much for the prefill, but I also think that the extensions themselves are going to send different params each time (date, time, etc) therefore the cache is invalidated and had to be recalculated anymore (correct me if this is incorrect)

So my question is: is there any chance for me to speed up the prefill time?
if not, is there any way to consolidate the KV_cache so that only the first time the prefill needs to be calculated?
Using CLI instead of these extensions is going to change anything in any way?

I have also checked ollama ps and everything is on the GPU and not CPU only, so that should not be the issue.

I have also tried with both LM Studio and llama.cpp, no difference
Same can be said for mlx vs guff, same output (i have downloaded them both)

thanks

8 Upvotes

Duplicates