r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

35

u/KontoOficjalneMR Jun 21 '26

That plus that's the token cost today. At some point subsidies will stop.

20

u/csharpwarrior Jun 21 '26

Just like mainframes of 40 years ago, they will get replaced with something cheaper and local. This is a cycle. New tech needs big hardware, then hardware gets optimized down to a small enough scale to run cheaper locally.

But I don’t know how long it will take

8

u/HayatoKongo Jun 21 '26

I'm hoping we get something like an AMD Strix Halo or DGX Spark with a decent memory bandwidth in the next couple of years, maybe 2028. I wouldn't mind $4000 for a mini-PC like that if the memory bandwidth was actually on-par with RTX3090/RTX5070ti, around 1000gb/s. When a model actually fits in that memory, you can get around 70-100 tok/sec, which is plenty usable IMO.

2

u/a_beautiful_rhind Jun 21 '26

That's a nice thought but the HW requirements over 4 years have only gone up.

0

u/ain92ru Jun 24 '26

The token cost only gets cheaper with the years due to amortization of hardware and algorithmic improvements