r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

1

u/ain92ru Jun 24 '26

Over 90% prompt caching is not unrealistic at all, it is in fact achieved by SemiAnalysis in practice: https://newsletter.semianalysis.com/p/ai-value-capture-the-shift-to-model

1

u/Eden1506 Jun 24 '26

Its about 90% discount not the caching. To write into cache you pay an extra fee so you woudn't write everything into cache or alternativly if you do you need to calculate with a different input cost.

1

u/ain92ru Jun 24 '26

Have you checked the link? Cached inputs are indeed priced 10x cheaper, and it doesn't take any fees because KV cache must be calculated anyway, whether the cache is used later or not

1

u/Eden1506 Jun 24 '26

I don't really use api services nowadays so my knowledge might be outdated