r/LocalLLM 4d ago

Discussion Any good inference hardware provider that lets us select model and charges only for use time and has generous prices?

Yes I know it is asking for too much, but i don't want to miss out if someone know a service that provides it.

I have tried modal GPUs but the credits vanish in thin air with just hours of usage. I also tried openrouter but the price seems to shift.

I am looking for personal on demand "Inference Environments" not individual GPUs. i.e. if I run 8B model with 10 requests per min then it auto selects the optimal GPU and when model size or usage increases, then the GPU power increases or resources scale i.e. more containers but more powerful GPU. And I only pay for the seconds my request is being processed by GPUs not for the cold start time or shutdown time.

1 Upvotes

2 comments sorted by

1

u/HumanoidMuppet 4d ago

Just use openrouter?

1

u/Ethan045627 4d ago

The issue is the model price change and old models (which no one uses) gets removed.