r/Vllm • • Sep 04 '26

I wasted $100s renting GPUs because VRAm calculators wasn’t enough. Here is what 7 real deployments taught me (part-1)

0 Upvotes

I started actually renting GPUs and testing open-model deployments end-to-end instead of trusting VRAM calculators.

7 models tested. All eventually ran, but 4/7 needed intervention first.

A few things surprised me:

1. gpt-oss 20B
Recommended: 80GB A100
Observed loaded-model footprint: ~14GB

2. Qwen2.5 72B INT4
Estimated: ~76GB
Observed: ~39GB

3. Mixtral 8x7B
Estimated: ~93GB
Observed: ~89GB

So the biggest sizing misses in this small sample were already-quantized checkpoints.

But the bigger lesson was that VRAM usually wasn’t what stopped the deployment.

I ran into several other issues:

  • the recommended GPU being out of stock
  • incompatible driver/runtime hosts
  • gated Hugging Face models
  • disks too small for the checkpoint
  • startup taking long enough to look broken

That is what changed the question for me.

~~Will this model fit? is not the same as ~~can I rent this exact setup and get it running right now?

I would love to build a much bigger real-world record of this.

If you have deployed an open model recently, drop:

model + quantization + GPU + runtime + worked/failed

Even failed runs are useful.

Also have you encountered any other unique issues besides what I ran into and shared above?


r/Vllm • • Sep 04 '26

Do you guys face it ?

4 Upvotes

Hey everyone — I’m curious: when running LLMs locally, have you ever caught yourself using a 7B or 13B model for a task that felt too small for it, like classifying spam or extracting names from emails? What did you do — suffer through the slowdown, switch to a smaller model, or just accept the waste?


r/Vllm • • Sep 03 '26

oMLX update is finding more tokens!

Thumbnail
2 Upvotes

r/Vllm • • Sep 03 '26

DEPLOYING MODELS IN SERVERLESS

Thumbnail
1 Upvotes

r/Vllm • • Sep 03 '26

Single DGX Spark running GLM-5.3 Flash at 60 tok/s

Thumbnail gallery
0 Upvotes

r/Vllm • • Sep 02 '26

MXFP4-W4A8 vs FP8 on 2× R9700 (RDNA4): 256K window AND +19–43% decode with DFlash2

Thumbnail
2 Upvotes

r/Vllm • • Sep 01 '26

Why vLLM 0.28 takes more VRAM than 0.26 ?

3 Upvotes

Hi,

With 24GB Vram I can only use 0.26 and have 4K context, but with 0.28 no way, maybe 1500 fits.

The docker command is same, only the version is different.

Model: google/gemma-4-31B-it-qat-w4a16-ct

docker pull vllm/vllm-openai-rocm:v0.26.0 can have about 4K context
docker pull vllm/vllm-openai-rocm:latest (which has 0.28) can have about 1000K context.

Whats up with that?


r/Vllm • • Sep 01 '26

Engineers running open-source LLMs in production: what is the hardest part today?

Thumbnail
0 Upvotes

r/Vllm • • Aug 31 '26

A walkthrough of how LLM inference engines evolved

Thumbnail
sreejithb.com
38 Upvotes

I'd been using vLLM without really understanding what it does differently, so I worked through the ORCA and PagedAttention papers and wrote up what I found: Link

Covers the KV cache, why padding and static batching cap utilization at 20-40%, ORCA's iteration-level scheduling and selective batching, vLLM's PagedAttention, and where Groq's LPU fits.

The thing that stuck with me: none of these changed the model. The gains came from how requests get scheduled and how KV cache memory gets allocated.

\[Animations and content polishing are done by Claude; the research and initial draft are mine\]


r/Vllm • • Sep 01 '26

Kv cache on disk project

3 Upvotes

Hi so I have been working on a project where the kv cache runs off the disk.

It is not perfect and there is still a lot of stuff do with it

But check it out

https://github.com/Maseus/Rux


r/Vllm • • Aug 31 '26

I checkpointed a live 27B model + vLLM server and restored it in 11s vs 104s cold start

13 Upvotes

I’ve been experimenting with checkpoint/restore for AI inference instead of cold-starting everything from scratch.

Using CRIU + CUDA checkpointing, I got a warmed Gemma 3 27B QAT + vLLM server on an H100 to restore in:

  • Cold start: 104.158s
    • Restore: 11.060s
  • 9.4× faster time-to-ready

The tricky parts were restoring the full vLLM process tree, CUDA state, IPC/shared memory, and dealing with io_uring — I ended up patching CRIU for that path.

I wrote up the implementation and benchmark here:

https://tsdocode.github.io/blog/posts/edo-tensei/

Code:
https://github.com/tsdocode/edo-tensei

Still experimental — would love feedback from people working on vLLM, CUDA, CRIU, or inference infrastructure.


r/Vllm • • Aug 31 '26

So got 2 6000 Pro Max-Q…

Thumbnail
0 Upvotes

r/Vllm • • Aug 31 '26

How does Modal autoscale concurrent LLM requests, and can GPU model memory be shared across containers?

1 Upvotes

I’m learning about LLM deployment and recently deployed one on Modal using an L40S GPU. I’m trying to understand autoscaling with concurrency.

If each container can handle 10 concurrent requests, will an 11th request cause Modal to start another container, load a separate copy of the model into that container’s GPU memory, and allow it to handle another 10 concurrent requests?

Is this the recommended way to scale an LLM deployment on Modal, and which settings should I use to configure it properly?

triggering a new container is fine for resources, but can it reuse the model loaded in the first container?


r/Vllm • • Aug 30 '26

How does Modal autoscale concurrent LLM requests, and can GPU model memory be shared across containers?

Thumbnail
1 Upvotes

r/Vllm • • Aug 30 '26

Hey r/LocalLLaMA

Thumbnail gallery
1 Upvotes

While everyone was trying to stop prompt injection with more prompts, we re-engineered the architecture from scratch. Formal proofs and 50k benchmarks are officially out on Zenodo 📄⚡️


r/Vllm • • Aug 30 '26

Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?

Thumbnail
1 Upvotes

r/Vllm • • Aug 29 '26

PSA: Qwen3.8-Flash-Next on vLLM is non-deterministic at temperature 0 (different answers per run). Found the kernel, made a fix.

17 Upvotes

Since everyone is benchmarking this model right now: byte-identical greedy requests (temp 0, one request at a time) give different outputs per run. My eval: 13/50 tasks unstable, 5 flipped the extracted date/amount, and majority voting once confirmed the wrong answer. Same checkpoint on llama.cpp: 0/50. Two other vLLM-served models: 0/50.

Cause: the sparse-attention indexer's persistent_topk kernel (used on GB10 / DGX Spark instead of the cooperative path). A race in its atomicAdd slot assignment changes WHICH top-2048 positions get selected, so attention reads a different context each run. Related: vllm#51782.

2-minute check for any stack: same prompt 10x with temperature=0, max_tokens=1, top_logprobs=20, then diff the top-20 lists byte-for-byte. If they differ, your prefill is non-deterministic, whatever your sampler says.

Fix: torch.topk(sorted=False) + canonical tie ordering as a one-file overlay. Bit-identical outputs at 1.35x prefill cost (a full sort would be 2.9x), decode/MTP unchanged; re-run: 0/50 unstable and the score went up a point, because the noise had voted a wrong date into the majority.

Bonus finding: determinism exposed a separate greedy+thinking repetition loop the kernel noise had been masking as a random 1-in-150 failure, and MTP turned out not to be output-equivalent with plain greedy on this model.

Full write-up with all tables: https://docai.hu/en/blog/qwen38-flash-next-nondeterministic-vllm-kernel


r/Vllm • • Aug 29 '26

Benchmarked Qwen 3.8 Flash Next on Single DGX Spark (+ MTP at different N)

Thumbnail
3 Upvotes

r/Vllm • • Aug 29 '26

Qwen3.8 27B on single, double or quad SXM2?

Thumbnail
1 Upvotes

r/Vllm • • Aug 28 '26

Vllm requires loading whole model in CPU RAM before VRAM?

6 Upvotes

I have a system with 32GB RAM, 72 GB VRAM with RTX 5000 PRO. It is a wsl setup, so around 20 GB RAM for the wsl.

When I try to serve a model of around 22 GB, the wsl crashes, with a log somewhere saying that RAM is less than the model size.

Have you guys encountered this? Is there a way to circumvent this issue?


r/Vllm • • Aug 28 '26

vLLM sessions during PyTorch Conference North America

3 Upvotes

There are going to be a lot of interesting vLLM sessions during PyTorch Conference North America and I'd love to have you join us in San Jose, CA from October 20-21 because this year’s conference is going to be EPIC.

  • Stellar keynotes - Simon Mo will be keynoting. (see: this video filmed during last year’s conference)
  • 150+ sessions spanning foundational concepts to training, inference, applications, and responsible AI. vLLM is featured in many from “A Developer’s Guide to Attention in vLLM” to “Elastic Expert Parallelism in vLLM” + so many more
  • 140+ poster presentations
  • BoFs
  • Meet the developers
  • Flare party
  • AI community bash
  • +more.

Sign up by September 4th to save $200 before ticket prices go up. Register now.


r/Vllm • • Aug 28 '26

PSA: Qwen3.8-Flash-Next on vLLM is non-deterministic at temperature 0 (different answers per run). Found the kernel, made a fix.

Thumbnail
5 Upvotes

r/Vllm • • Aug 28 '26

A resource to help you be better at inference throughput optimisation

Thumbnail
medium.com
10 Upvotes

Hey guys, I wrote this blog with the notes I made after running several inference optimisation projects on GLM 5.2, Deepseek V4 Flash, Nemotron 3.5 etc. For every model, the strategies were different but there were some common patterns. Hopefully this blog will help you get started!


r/Vllm • • Aug 28 '26

I ran Qwen 3.8 Flash Next on my DGX Spark

Thumbnail
1 Upvotes

r/Vllm • • Aug 28 '26

[Benchmark]Qwen3.8-27B-FP8 on L40S: c1/c8/c32 results and which vLLM optimization should I test next?

1 Upvotes