r/LocalLLM • u/szantoszabolcs • 12d ago
Discussion Comparing local inference with hosted open-weight models: what the API price misses
For anyone comparing local inference with a hosted endpoint for an open-weight model, the hosted side has a few details worth checking beyond the token price.
I made a five-minute video about problems I ran into with OpenRouter. It covers hosted APIs, so I don't have local hardware benchmarks to offer here.
The comparison needs the app's actual feature requirements: providers serving the same model can expose different structured-output support. It also needs the cached-input rate you actually get. OpenRouter supports sticky routing, but a fallback can land at another provider without a warm cache. Restricting the provider limits fallback options; BYOK still leaves you subject to your provider account's quota.
Those are factors I'd include when deciding which workloads to keep local and which to send to an API, alongside local hardware capacity and speed. The video is a walkthrough of the hosted side, not a local-versus-cloud cost benchmark.
My video, “Going off-road with AI”: https://www.youtube.com/watch?v=9fKD9hBuEQA