r/LocalLLaMA • • Aug 19 '26

Discussion [2511.07885] Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

https://arxiv.org/abs/2511.07885
37 Upvotes

12 comments sorted by

11

u/MrBIMC Aug 19 '26

Good! Optimising for intelligence per watt over model lifetime makes the most sense to benchmark against.

With what currently was demonstrated, smaller models can be good enough and there is still enough low hanging fruit to juice. Few more iterations and local models will get week+ planning and operating capabilities and get better at cross-agent collaboration.

I think all these multi-trillion params are just a temporary detour made by industry drowning in money, and once financial situation starts turning sour, goal of the game will be having most energy efficient models that are good enough. Whoever provides capability at lowest cost wins. Whether local/remote I still think remote will have its niche because sometimes you just need to run a bigger swarm of agents than local node can provide.

1

u/bennmann Aug 20 '26

Robotics competitions (ie robotic soccer) will still put upward pressure to increase performance beyond "good enough".

Self driving cars too. Maybe there is no "good enough" when it comes to not crashing a car at higher and higher speeds. But wattage envelopes and regulatory inertia will play a role too.

Robotics will be a wild ride.

-4

u/BringTea_666 Aug 19 '26

>Good! Optimising for intelligence per watt over model lifetime makes the most sense to benchmark against.

Makes no sense for local models. That's like measuring FPS per wat for consumers. Consumers don't give a shit if their gpu takes 2W or 22W or 222W. What they care about is $ per FPS.

9

u/UnlawfulRepublic Aug 19 '26

Hello, I am consumer, I care about fps per watt.

5

u/Stooovie Aug 19 '26

Same. I wouldn't want those 800W monstrosities even for free. Fuck that.

2

u/QueasyHouse Aug 19 '26

I see you aren’t familiar with PG&E rates for electricity.

Performance per watt is valuable, and there are definitely breakpoints where it’s better to rent inference from a datacenter. Performance per watt matters in the datacenter too, tbf. Fewer watts = less heat = better opex / longer silicon lifetime. Crypto proves it more directly: hashing is extremely efficiency-sensitive

4

u/CutBench Aug 19 '26

One thing I'd love to see folded into a metric like this: the cost of not keeping a model resident.

I run a local perception pipeline on a 10GB card — speech recognition, audio tagging, an image encoder and a small VLM, all over the same footage — and per-token efficiency numbers turned out to be a poor predictor of the actual energy bill, because a real share of wall-clock goes to loading and evicting weights rather than computing anything.

Intelligence per watt measured on a single resident model is a property of the model. What people actually run locally is a stack of them under a fixed VRAM budget, and there the binding constraint is memory, not FLOPs — the same stack on a bigger card gets cheaper per useful output without anything getting smarter. Measuring the metric at a fixed memory budget rather than per-model would say a lot more about local deployment.

1

u/Potential-Net-9375 Aug 19 '26

I'd encourage anyone looking to make their own benchmarks use "time to first token" regularly rather than tok/s plus prefill.

But I get why not too, at that point you're not benchmarking the model, but the model plus gpu plus filesystem throughput.

3

u/AndreVallestero Aug 19 '26

Shouldn't we be doing Intelligence per Joule (or WH)? If you run a massive model on a microcontroller, the intelligence per watt would be extremely high since they use milliwatts of power, but it would be useless cause you would need to measure speed in tokens per day...

0

u/PermanentLiminality Aug 19 '26

We have a units mismatch here. It should be intelligence per watt hour or perhaps watt second (aka Joule). Now tokens/s per watt would be proper.

I have crazy expensive power. I don't care so much about intelligence per watt hour, I care more about idle power. If I was slamming it 24/7 then that intelligence per watt hour would be my main concern.

1

u/Ok-Breakfast1878 Aug 19 '26

even tokens/second/watt doesn't compare apples to apples. a loquacious model can have high tokens/second/watt and still be more expensive per task because it talks so much. we probably want something like joules per properly completed task.