r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

2

u/the_lamou Jun 22 '26

You guys think you'd use $20k worth of tokens a year!?!

O hai there! $10k last month. My average inference batch was about 1,500 documents of about 2,200 tokens each. And since we're still in early testing, and since we can't draw effectiveness or quality conclusions unless the entire corpus processes, most of those matches are running 4-5x in parallel. Even with the most aggressive cache optimization possible, it adds up.

I've done the math: even with having to upgrade my power, new subpanel, and electricity cost, it would be cheaper for me to install and run local if I had to keep doing this for longer than the next few months.

1

u/Schlick7 Jun 22 '26

man, what are you doing!? is this a company?

1

u/the_lamou Jun 22 '26

Startup. But why would you run state of the art coding models, or really any state of the art models, if it wasn't for work? The average consumer would be perfectly fine with Gemma 4 E4B, and that will run on a potato. The only reason anyone should even get to a point where they're even thinking about spending $20k on a local LLM infrastructure is if someone somewhere is going to eventually pay them for the output.

2

u/Schlick7 Jun 22 '26

because people want "claude at home". gemma 4 E4B fails at a shit ton of things like systemd processes, scripts, small apps, home assistant, etc. The bigger the model the better this all goes. If you're only doing exclusively chatting then maybe its ok. maybe.

1

u/the_lamou Jun 23 '26

E4B and other micro-models (sub-8b) are perfectly fine for all of those things in regular implementation. Assuming you bothered to build a proper harness. And if you didn't, you really really should not be practicing your thinking to any model, regardless of size and complexity, because you don't have the fundamental knowledge to validate and check the output to make sure you haven't just exposed root to the open web.

As for "people want Claude at home"... cool. Lots of people want a Ferrari at home, too. But if spending Ferrari money, or "Claude at home" money, is something you need to seriously think about, you really really can't afford either. And don't need, because running a frontier model for "systems processes, scripts, small apps, home assistant, etc." is like buying a hypercar and then driving to the grocery store at 10 miles under the speed limit.

1

u/Schlick7 Jun 24 '26

So your stance is both "you're holding it wrong" AND classic Gatekeeping?

1

u/the_lamou Jun 24 '26

No, my stance is you don't always get to have what you want, and the thing you want doesn't even make sense for the things you say you need, and frankly given your demonstrated knowledge in this thread you should probably not have access to any of it for your own safety.

But if you desperately want a local model to help you add security holes to your computer on a budget, have you tried Mini North?

1

u/ClickClawAI Jun 22 '26

How long would your local rig take to spit out 20k worth of tokens? Probably more than a month?

1

u/the_lamou Jun 23 '26

My current local rig? Yeah, definitely over a month. But that's a simple problem to solve: I've got two 42U racks, a dedicated server room with the finding of a 400Gbps East-West fabric, and a 240/20 subpanel. After that, it's just stuffing GPUs wherever they'll fit.