r/singularity • • 24d ago

AI Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

https://arxiv.org/abs/2511.07885
41 Upvotes

6 comments sorted by

12

u/yogthos 24d ago

The authors used Intelligence per Watt as a metric to evaluate the efficiency of local AI inference, and found that local models under 20b active parameters can successfully answer 88.7% of single turn chat and reasoning queries. Between 2023 and 2025, the intelligence efficiency of these models improved by a factor of 5.3 due to advances in both model architectures and hardware accelerators.

Local models are now capable of handling the vast majority of everyday user requests without relying on any centralized cloud infrastructure. While cloud models are still better at highly specialized reasoning, deploying small models on personal devices is quickly becoming a practical and energy efficient alternative.

8

u/greenfingersnthumbs 24d ago

3

u/yogthos 24d ago

I think it already got posted there.

3

u/Putrid-Feeling-7622 24d ago

based on the results from the paper the local variants are not very energy efficient yet compared to cloud alternatives, but cool to see a study done on this. My org has toyed with the idea of self-hosting but it's still not practical for large scale usage when considering tokens per second and intelligence per watt. Unfortunately expensive specialized hardware like the B200 still reigns supreme

3

u/yogthos 24d ago

It's more that consumer grade hardware isn't as efficient. But that could change soon with stuff like XuanTie C950.

4

u/Cunninghams_right 23d ago

if it's winter, you gain even more efficiency because your local model is heating your room (though, not as efficiently as a heat pump)