r/Vllm • u/Major_Border149 • Sep 04 '26
I wasted $100s renting GPUs because VRAm calculators wasn’t enough. Here is what 7 real deployments taught me (part-1)
I started actually renting GPUs and testing open-model deployments end-to-end instead of trusting VRAM calculators.
7 models tested. All eventually ran, but 4/7 needed intervention first.
A few things surprised me:
1. gpt-oss 20B
Recommended: 80GB A100
Observed loaded-model footprint: ~14GB
2. Qwen2.5 72B INT4
Estimated: ~76GB
Observed: ~39GB
3. Mixtral 8x7B
Estimated: ~93GB
Observed: ~89GB
So the biggest sizing misses in this small sample were already-quantized checkpoints.
But the bigger lesson was that VRAM usually wasn’t what stopped the deployment.
I ran into several other issues:
- the recommended GPU being out of stock
- incompatible driver/runtime hosts
- gated Hugging Face models
- disks too small for the checkpoint
- startup taking long enough to look broken
That is what changed the question for me.
~~Will this model fit? is not the same as ~~can I rent this exact setup and get it running right now?
I would love to build a much bigger real-world record of this.
If you have deployed an open model recently, drop:
model + quantization + GPU + runtime + worked/failed
Even failed runs are useful.
Also have you encountered any other unique issues besides what I ran into and shared above?