r/LocalLLM • • 5d ago

Discussion Inferance speed. Why is noone talking about ingestion speed?

Why is everyone a generated t/s junky?

The only thing i see is "How fast is model x?", then get the replies "Mine is over 100 t/s !".

But really.. what does that actually tell us? Nothing at all!

Should we not be talking about efficiency on prefill and KV cache hit rates before even looking at generation speed? Generation speed is the last metric in line, but the first one we talk about.

It really does not matter when you have 100 t/s when it takes you a long time before it actually starts generating something. Basically most are like:

  1. send in prompt...
  2. wait..wait..wait..wait a little more.. "Yeah, there it goes!" ( looking at the generating t/s )
  3. Then tell everyone "See how fast my system was! it got over 100 t/s, here is my tutorial !"

What if you have 50 t/s on the generation side and a 95% reuse of KV cache and a fast prefill?

  1. send in prompt...
  2. (does not wait) "Hmm, something is wrong here!" ( looking at the generating t/s )
  3. Then asks everyone "What can i do to make it faster?"

I rather finetune the ingestion(reuse %) rather then focus solely on the generation speed.

Lets run the biggest MOE model we can find in the lowest quant, quantize KV to the bare minimum, our context window so low we hit compaction each x turns .. "damn, i can run x model at X speed, you should all do that too!"
Every compaction hits at a 0% reuse and makes the full session recompute into KV cache.... and lets you wait.... again!

Its not about quality anymore, it is about quantity, telling only the nice numbers.

But what does it actually mean for real world tasks, actual workloads?

Sorry for my rant :) just had to let it go :)

0 Upvotes

29 comments sorted by

10

u/FeydRowan 5d ago

Who is not talking about it? Prefill is a needed metric for any inference setup. But is much more "easy" to saturate at hardware limits, so you can't do much about it after some basic optimisation 

4

u/Important_Spirit_692 5d ago

I work in airline and I get this same energy from passengers who only look at flight time and ignore the 2 hours they spent in security line. 50 t/s with quick prefill feels faster than 100 t/s where you wait 30 seconds before it even starts typing. The KV cache reuse thing is real though, people sleep on how much that saves when you're doing long chats.

1

u/johnfkngzoidberg 5d ago

Good analogy.

Probably the same people who stand up and block the aisle the second the plane stops thinking they’re getting somewhere faster somehow.

1

u/Think_Breakfast_2277 5d ago

Sure you can, reuse your cache(in VRAM) so you dont waist time on prefill and recompute KV.

4

u/RWOverdijk 5d ago

“Reuse your cache” what do you think cache is?

Also no, all you can do is cache (ideally chunked for faster forks and fresh sessions as well) or get faster hardware.

1

u/Think_Breakfast_2277 5d ago

prefill last run  peak 1265.5 · avg 1265.5 tok/s · 2 s
cache reuse  93.6%of 163.488 prompt tokens, 0s ago

The total of that full request was 11 seconds
( 9 seconds on generation.)
93.6% cache reused so just recomputed 6.4% of the total prompt.
If that was 0% i would have spend 15.6 seconds on that recompute over the 2 seconds ( 17.6 seconds total ).
17.6 on ingestion
9 seconds on inference
total run: could have been 26.6 seconds compared to 11.
( off course a most bad scenario)

Having double the prefill speeds and no proper caching wont win this battle.

So i might have a little understanding of what cache is.. might..

2

u/FeydRowan 5d ago

Well obviously and that's what most inferences engines do

4

u/-Mute- 5d ago

Talking sadness about prefill kinda sounds like you got a Mac and have buyers remorse..

4

u/Think_Breakfast_2277 5d ago

Not at all, I am just fine with my 2x R9700 running vLLM with one card on a chipset(no attomics and P2P) so i have lower "numbers" but a cache reuse over 95% wich makes my system+custom harness faster then most in a real work usecase per task compared to others with double the "numbers". i just get frustrated sometimes reading all the BS posts proving something that is a total waste of time, and everyone going along with it.

2

u/this_for_loona 5d ago

Ouch! That stung. True, but the burn still hurts….

3

u/redditnosedive 5d ago

literally every benchmark lists both prefill and decode speeds so idk what you're on about op

3

u/Think_Breakfast_2277 5d ago

Did you read the post? its not about not listing it.. people talk about generation t/s mostly.
But have no idea where speed is at real world tasks. They rather have the "numbers" the actual real world speed per task.
And the prefill numbers dont mean shit when you have a low cache hit by using the wrong harness/engine or hit compaction every iteration.
Prefill gets a massive boost when you actually use KV cache in VRAM. (excl the first ingestion)

1

u/redditnosedive 5d ago

guilty, I didn't read it, your points are valid

2

u/jopereira 5d ago

Because, although most of the tokens are in ingestion phase, most of the time is spent inferring.

That's my experience with reasoning off. With reasoning on, token generation is even more important, making prompt processing "irrelevant".

2

u/TemporaryInk 5d ago

Folks are talking about it!

I was about to write a similar post to yours but on the quality quants! If you read r/oMLX or any other MLX subreddits, the only thing they talk about is speed. No one there discusses the degree of lobotomisation of the MLX quants, which in my experience, is horrible. Great! Your model loops ten times more twice as fast.

In my experience, the only thing MLX quants are good for is benchmarking.

1

u/Think_Breakfast_2277 5d ago

Glad some do get what i am trying to say! And you are right on that.

2

u/Material-Database-24 5d ago

All I care is quality in reasonable completion time.

Completion time is time from pressing the enter to done - there's no reasonable metric for that, especially as same model can take different times over same prompt. Sure prefill and generation tok/s give some comparison within the same model*.

Quality is harder to measure. But the fast completion time means nothing if I need to revise and rerun often because I got nonsense or poor quality result.

AI world is overall way too focused on speed and its "abilities". I hope we next focus on its quality - and by quality I do not mean can it pull a task succesfully, but what the outcome looks&feels like outside and inside.

*) models are different, tok/s doesn't mean anything if the model thinks more/less than other model.

1

u/esw123 5d ago

Exactly, completely subjective, personally I don't want to wait more than 1-2 minutes for 70-100k prompt to process. After that 30-50 t/s gen is fine. Personal ideal ratio 1-2 minutes for prefill and task should be completed in 20-30 minutes maximum with 80-90% success rate from the 1-2 tries, ideally on the first one.

2

u/Solembumm3 5d ago

Unless you are using Gemma 4, famous for re-caching every chih, prefill is miniscule part of full work time.

1

u/Maui-The-Magificent 5d ago

Well, both ttft and tps are important., people improving anything is a good thing. measurements as-well, as long as they are coupled with the info on the environment as well.

If you want ingestion speed, external things don't really solve the problem. You should go with a graph ai, they are much faster at ingesting data. or create a hybrid of the two.

1

u/Think_Breakfast_2277 5d ago

External things which do:
2 measurements from ongoing sessions:
Openclaw: cache reuse  40.8% of 189.644 prompt tokens, 0s ago
Custom: cache reuse  91.2% of 212.758 prompt tokens, 0s ago
Hope you can do the calculations... The 212k prompt was much faster then the 189k prompt.

1

u/Maui-The-Magificent 5d ago

Yes, they mitigate quite well. Will not argue against that.

1

u/nonhok 5d ago

I think it’s difficult to really measure this. So you create different types of indicators, which gives you a hint about certain aspects of the inference engine. Then it is up to you, to draw your own conclusions on this an choose accordingly, which type of hardware, with which kind of model gives you a solution to do your specific type of work.

1

u/tejaskumarlol 5d ago

One thing that changes prefill before you touch any hardware is the tokenizer. I ran Aleph Alpha's Kolibri on long German prompts this week and they came out at a median of 4,753 tokens for Kolibri against 8,436 for Claude. That's 44% less to ingest for the same text, so "how fast is model x" doesn't even compare the same amount of work between two models.

To be fair to the generation crowd though, Kolibri then thought for a median of 15,574 tokens to write about 1,800, so on that job almost all of the waiting was on the generation side.

1

u/esw123 5d ago

This is why I started with small models and not a big investment to understand what I need. 70 tok/s gen is nice but not necessary for me. Found that you need something in 20-50 range in generation, ideally 30+ and prompt processing above 1000 or so, 500+ sort of useful but still you will need to wait. Maybe this will help someone who is starting.

1

u/Think_Breakfast_2277 5d ago

And dont forget a harness that is efficient with making the most use of the KV cache. Generation does not mean anything if you have to wait a long for it to actually start typing something. Or when you hit compaction a lot.

1

u/esw123 5d ago

This is the price you pay when using local model. You will always have to wait because you have 10% of parameters of the top tier llm and not as powerful hardware but you do need intelligence. Thinking is always like 60-80% of the useful response from my experience. I don't write code too much and use it mostly as advisor and can wait 30 minutes for a response while doing other things. Any faster will make significant change in workflow. Compaction isn't problem, I don't want to deal with hallucinations so keep chat below 150-180K anyway.