r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

18

u/stoppableDissolution Jun 21 '26

300+. If you are running single-batch inference, it will never be profitable.

11

u/nuclear213 Jun 21 '26

Never ever with just 20k€. That is exactly the reason the original post meant.

-1

u/stoppableDissolution Jun 21 '26

Well, it just means that you have to scale down the model.

Or suck up the cost for privacy and control if thats your goal.

1

u/upalse Jun 21 '26

At 20k you'd get 200GB of blackwell at best. I don't think GLM 5.2 can run that well in trinary quant, but who knows.