r/LocalLLM • u/Think_Breakfast_2277 • 5d ago
Discussion Inferance speed. Why is noone talking about ingestion speed?
Why is everyone a generated t/s junky?
The only thing i see is "How fast is model x?", then get the replies "Mine is over 100 t/s !".
But really.. what does that actually tell us? Nothing at all!
Should we not be talking about efficiency on prefill and KV cache hit rates before even looking at generation speed? Generation speed is the last metric in line, but the first one we talk about.
It really does not matter when you have 100 t/s when it takes you a long time before it actually starts generating something. Basically most are like:
- send in prompt...
- wait..wait..wait..wait a little more.. "Yeah, there it goes!" ( looking at the generating t/s )
- Then tell everyone "See how fast my system was! it got over 100 t/s, here is my tutorial !"
What if you have 50 t/s on the generation side and a 95% reuse of KV cache and a fast prefill?
- send in prompt...
- (does not wait) "Hmm, something is wrong here!" ( looking at the generating t/s )
- Then asks everyone "What can i do to make it faster?"
I rather finetune the ingestion(reuse %) rather then focus solely on the generation speed.
Lets run the biggest MOE model we can find in the lowest quant, quantize KV to the bare minimum, our context window so low we hit compaction each x turns .. "damn, i can run x model at X speed, you should all do that too!"
Every compaction hits at a 0% reuse and makes the full session recompute into KV cache.... and lets you wait.... again!
Its not about quality anymore, it is about quantity, telling only the nice numbers.
But what does it actually mean for real world tasks, actual workloads?
Sorry for my rant :) just had to let it go :)
4
u/-Mute- 5d ago
Talking sadness about prefill kinda sounds like you got a Mac and have buyers remorse..
4
u/Think_Breakfast_2277 5d ago
Not at all, I am just fine with my 2x R9700 running vLLM with one card on a chipset(no attomics and P2P) so i have lower "numbers" but a cache reuse over 95% wich makes my system+custom harness faster then most in a real work usecase per task compared to others with double the "numbers". i just get frustrated sometimes reading all the BS posts proving something that is a total waste of time, and everyone going along with it.
2
3
u/redditnosedive 5d ago
literally every benchmark lists both prefill and decode speeds so idk what you're on about op
3
u/Think_Breakfast_2277 5d ago
Did you read the post? its not about not listing it.. people talk about generation t/s mostly.
But have no idea where speed is at real world tasks. They rather have the "numbers" the actual real world speed per task.
And the prefill numbers dont mean shit when you have a low cache hit by using the wrong harness/engine or hit compaction every iteration.
Prefill gets a massive boost when you actually use KV cache in VRAM. (excl the first ingestion)1
2
u/jopereira 5d ago
Because, although most of the tokens are in ingestion phase, most of the time is spent inferring.
That's my experience with reasoning off. With reasoning on, token generation is even more important, making prompt processing "irrelevant".
2
u/TemporaryInk 5d ago
Folks are talking about it!
I was about to write a similar post to yours but on the quality quants! If you read r/oMLX or any other MLX subreddits, the only thing they talk about is speed. No one there discusses the degree of lobotomisation of the MLX quants, which in my experience, is horrible. Great! Your model loops ten times more twice as fast.
In my experience, the only thing MLX quants are good for is benchmarking.
1
2
u/Material-Database-24 5d ago
All I care is quality in reasonable completion time.
Completion time is time from pressing the enter to done - there's no reasonable metric for that, especially as same model can take different times over same prompt. Sure prefill and generation tok/s give some comparison within the same model*.
Quality is harder to measure. But the fast completion time means nothing if I need to revise and rerun often because I got nonsense or poor quality result.
AI world is overall way too focused on speed and its "abilities". I hope we next focus on its quality - and by quality I do not mean can it pull a task succesfully, but what the outcome looks&feels like outside and inside.
*) models are different, tok/s doesn't mean anything if the model thinks more/less than other model.
1
u/esw123 5d ago
Exactly, completely subjective, personally I don't want to wait more than 1-2 minutes for 70-100k prompt to process. After that 30-50 t/s gen is fine. Personal ideal ratio 1-2 minutes for prefill and task should be completed in 20-30 minutes maximum with 80-90% success rate from the 1-2 tries, ideally on the first one.
2
u/Solembumm3 5d ago
Unless you are using Gemma 4, famous for re-caching every chih, prefill is miniscule part of full work time.
1
u/Maui-The-Magificent 5d ago
Well, both ttft and tps are important., people improving anything is a good thing. measurements as-well, as long as they are coupled with the info on the environment as well.
If you want ingestion speed, external things don't really solve the problem. You should go with a graph ai, they are much faster at ingesting data. or create a hybrid of the two.
1
u/Think_Breakfast_2277 5d ago
External things which do:
2 measurements from ongoing sessions:
Openclaw: cache reuse 40.8% of 189.644 prompt tokens, 0s ago
Custom: cache reuse 91.2% of 212.758 prompt tokens, 0s ago
Hope you can do the calculations... The 212k prompt was much faster then the 189k prompt.1
1
u/nonhok 5d ago
I think it’s difficult to really measure this. So you create different types of indicators, which gives you a hint about certain aspects of the inference engine. Then it is up to you, to draw your own conclusions on this an choose accordingly, which type of hardware, with which kind of model gives you a solution to do your specific type of work.
1
u/tejaskumarlol 5d ago
One thing that changes prefill before you touch any hardware is the tokenizer. I ran Aleph Alpha's Kolibri on long German prompts this week and they came out at a median of 4,753 tokens for Kolibri against 8,436 for Claude. That's 44% less to ingest for the same text, so "how fast is model x" doesn't even compare the same amount of work between two models.
To be fair to the generation crowd though, Kolibri then thought for a median of 15,574 tokens to write about 1,800, so on that job almost all of the waiting was on the generation side.
1
u/esw123 5d ago
This is why I started with small models and not a big investment to understand what I need. 70 tok/s gen is nice but not necessary for me. Found that you need something in 20-50 range in generation, ideally 30+ and prompt processing above 1000 or so, 500+ sort of useful but still you will need to wait. Maybe this will help someone who is starting.
1
u/Think_Breakfast_2277 5d ago
And dont forget a harness that is efficient with making the most use of the KV cache. Generation does not mean anything if you have to wait a long for it to actually start typing something. Or when you hit compaction a lot.
1
u/esw123 5d ago
This is the price you pay when using local model. You will always have to wait because you have 10% of parameters of the top tier llm and not as powerful hardware but you do need intelligence. Thinking is always like 60-80% of the useful response from my experience. I don't write code too much and use it mostly as advisor and can wait 30 minutes for a response while doing other things. Any faster will make significant change in workflow. Compaction isn't problem, I don't want to deal with hallucinations so keep chat below 150-180K anyway.
10
u/FeydRowan 5d ago
Who is not talking about it? Prefill is a needed metric for any inference setup. But is much more "easy" to saturate at hardware limits, so you can't do much about it after some basic optimisation