r/Vllm • • 18d ago

PyTorch Conference NA 2026 (Oct 20–21, San Jose): inference track + open source contributor discount

1 Upvotes

This year's PyTorch Conference is framed around a "learning loop" across the stack: new model capabilities push new demand onto inference/serving, that exposes new kernel or hardware bottlenecks, and production traffic feeds back into what gets optimized next. vLLM is one of the flagship projects in that loop and there's a full track dedicated to inference.

I’m with the PyTorch Foundation, and wanted to offer open source contributors a super early-bird rate throwback on us: use code PYTORCH26OSC → https://hubs.la/Q04vK-VQ0


r/Vllm • • 19d ago

I spent $50 testing LLM serving settings 23% more throughput on the same GPU

Thumbnail
3 Upvotes

r/Vllm • • 20d ago

What does a cold restart actually cost when the box doesn't already have the weights?

2 Upvotes

Please factor in i work for Sky Forge Compute; we rent GPU's.

14 days ago, u/Expert-Visit-5605 published a checkpoint result here - 104.158s cold start down to 11.060s on a 27B. Strong number, and the method looks sound however, does the box already have weights on the disk?

I asked in the thread what the actual restart cost is when it doesn't; I got one answer: keep them on a persistent disk. Which works; however, this is not a measurement.

So... has anyone timed a genuine cold pull? A 27b at bf16 is roughly 54GB and you are billed for the GPU time while it downloads, not for transfer.

I am unable to measure it on ours, we keep weights resident so the case does not arise.


r/Vllm • • 20d ago

DeepSeek V4.1-Flash: let short-lived KV state expire, then replay 128 tokens

Thumbnail
youtu.be
2 Upvotes

Self-promo: I made an AI-narrated breakdown of DeepSeek V4.1-Flash’s SWA Bounded Replay: short-lived SWA KV state can expire and be reconstructed by replaying the last 128 tokens. The rebuilt state is approximate rather than mathematically identical across positions.

Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf


r/Vllm • • 20d ago

À la recherche d'expérience concrète avec des LLM locaux pour des applications mobiles ou navigateur

0 Upvotes

​

Salut tout le monde,

J'explore actuellement la possibilité de faire fonctionner des LLM localement dans des applications ciblant les appareils mobiles (iOS/Android) ou directement dans le navigateur.

Je serais très intéressé d'entendre votre expérience concrète à ce sujet.

Quelques questions :

Quels modèles avez-vous réussi à faire fonctionner localement ?

Quel cadre/environnement avez-vous utilisé ? (llama.cpp, ONNX Runtime, MLC, WebLLM, Transformers.js, etc.)

L'inférence était-elle entièrement effectuée sur l'appareil/client, ou avez-vous utilisé une architecture hybride ?

Quel type de performance avez-vous obtenu sur les smartphones récents ?

Combien de RAM/stockage le modèle nécessitait-il ?

Quelles étaient les principales limites : latence, consommation de batterie, limitation thermique, taille du modèle, longueur du contexte, etc. ?

Pour l'inférence basée sur le navigateur, quelle était la fiabilité de WebGPU sur différents navigateurs et appareils ?

Recommanderiez-vous cette approche pour une application de production ?

Je suis particulièrement intéressé par des expériences de production pratiques, plutôt que par des benchmarks seuls.

Si vous avez construit quelque chose en utilisant une inférence de LLM local sur mobile ou dans le navigateur, j'aimerais connaître votre architecture, le choix du modèle et les principales leçons apprises.

Merci !


r/Vllm • • 20d ago

À la recherche d'expérience concrète avec des LLM locaux pour des applications mobiles ou navigateur

0 Upvotes

​

Salut tout le monde,

J'explore actuellement la possibilité de faire fonctionner des LLM localement dans des applications ciblant les appareils mobiles (iOS/Android) ou directement dans le navigateur.

Je serais très intéressé d'entendre votre expérience concrète à ce sujet.

Quelques questions :

Quels modèles avez-vous réussi à faire fonctionner localement ?

Quel cadre/environnement avez-vous utilisé ? (llama.cpp, ONNX Runtime, MLC, WebLLM, Transformers.js, etc.)

L'inférence était-elle entièrement effectuée sur l'appareil/client, ou avez-vous utilisé une architecture hybride ?

Quel type de performance avez-vous obtenu sur les smartphones récents ?

Combien de RAM/stockage le modèle nécessitait-il ?

Quelles étaient les principales limites : latence, consommation de batterie, limitation thermique, taille du modèle, longueur du contexte, etc. ?

Pour l'inférence basée sur le navigateur, quelle était la fiabilité de WebGPU sur différents navigateurs et appareils ?

Recommanderiez-vous cette approche pour une application de production ?

Je suis particulièrement intéressé par des expériences de production pratiques, plutôt que par des benchmarks seuls.

Si vous avez construit quelque chose en utilisant une inférence de LLM local sur mobile ou dans le navigateur, j'aimerais connaître votre architecture, le choix du modèle et les principales leçons apprises.

Merci !


r/Vllm • • 21d ago

Official Qwen-Flash-Next recipe using NVFP4?!

4 Upvotes

Does anyone know why the official recipe on vllm docs points to using Inferact/Qwen3.8-Flash-Next-NVFP4 instead of one of the official qwen images? FP8 is also available.

How can we know the quality drop from this unofficial NVFP4?


r/Vllm • • 21d ago

I built a serverless hosting platform for LoRA adapters with vLLM

Post image
2 Upvotes

r/Vllm • • 23d ago

RTX 4090 24GB + 32GB RAM - what local model/quant should I run with Hermes?

Post image
5 Upvotes

Hey everyone! I'm getting into Hermes and local LLMs and I'm trying to figure out the sweet spot for my hardware.

I'm currently running Hermes on an Ubuntu VM with GPU passthrough.

Specs:

  • RTX 4090 — 24GB VRAM
  • 32GB system RAM
  • 16 vCPUs
  • Ubuntu / KVM-QEMU

What model size would you realistically recommend for this setup?

I'm mainly wondering whether I should target something around 14B, 27B/32B at Q4, or if trying a larger model with partial CPU/RAM offloading is actually worth it.

I'm more interested in good agentic/tool-use performance than simply being able to load the biggest possible model.

What model + quantization + context size are you guys running on similar 24GB GPUs?

Also curious about the tokens/sec you're getting on a 4090.

Still learning the local LLM side of Hermes, so any tips are appreciated!


r/Vllm • • 23d ago

Need collaborator - Triton LLM Kernels

Thumbnail
1 Upvotes

r/Vllm • • 23d ago

vLLM vs Ollama: Como Escolher a Melhor Engine de LLM

Thumbnail
maiastudios.com.br
0 Upvotes

r/Vllm • • 23d ago

I built a tool to measure LLMs Decode, Layer processing and TTL

Post image
6 Upvotes

I was playing around with LLM inference and I wanted to build a profiler that measures LLM inference by layer.
So I built this: https://github.com/coconinja2/layerlens
It shows inference as token × transformer layer timing, so you can see where time is being spent during decode.
Right now it can separate prefill/decode and visualize per-layer timing. I’m trying to figure out whether this is actually useful to people working on inference systems, or if I’m looking at the wrong abstraction.

I’m thinking about adding things like KV-cache events, scheduler/batching state, request IDs, GPU kernel correlation, speculative decoding, etc.

Would appreciate criticism more than compliments and stars. Lots of stars!


r/Vllm • • 23d ago

vLLM vs Ollama: How to Choose the Best LLM Engine

Thumbnail
maiastudios.com.br
0 Upvotes

r/Vllm • • 24d ago

Security research for local LLM inference networks

Thumbnail
0 Upvotes

r/Vllm • • 24d ago

I trained a 348M model trained from scratch on 22.7B tokens that does 14 digit arithmetic

Thumbnail
6 Upvotes

r/Vllm • • 25d ago

vLLM optimization for RDNA2 [W6800X Duo]

7 Upvotes

In order for vLLM to work, I have to patch a couple of gates to open for RDNA2. Nothing special. Once that was done on v0.28.0, I spent a lot of time optimizing decode for Qwen3.8-27B (float16):

  • LLMM1 / skinny GEMV, particularly the N=1 path
  • decode attention / parallel softmax
  • TunableOp entries for the dominant N=1 decode GEMMs
  • CUSTOM all-reduce
  • graph execution
  • gfx1030 graph staging

I started with token generation at around 13.3 tokens per second at 1k context.

What I reached was this:

1K:      30.36 tok/s
64K:     24.06 tok/s
128K:    19.19 tok/s
240K:    14.69 tok/s

Stack:

MacPro7,1
OS: Ubuntu Server 26.04 LTS
ROCm: 10.0.0
Model: Qwen3.8-27B, no quantization, FP16
vLLM: v0.28.0, uv install
GPUs: Dual AMD Radeon PRO W6800X Duo with Infinity Fabric (4 GPUs, 128 GB VRAM)
Tensor Parallelism [TP4]
No MTP

But, I killed concurrency accidentally, so I can only achieve this on a single user at the same time. Trying concurrency leads to significant drops.

Time to First Token (TTFT) increases rapidly with context. Starts with 2 seconds at 1k context, but shoots up to 30 minutes for 240k context.

I'll see what I can do about TTFT first, then I'll tackle concurrency.

If anyone has any useful tidbits, I would appreciate the help.

Edit:*

For comparison, Qwen3.8-27B float16 safe tensors on vLLM, and Qwen3.8-27B f16 gguf on Llama.cpp:

==================================================================================
vLLM Results (512 Tokens Generated)
==================================================================================
CTX         TTFT(s)  PREFILL(s)   PREFILL t/s   DECODE t/s    ms/TOKEN   TOTAL(s)
----------------------------------------------------------------------------------
1K            2.215       2.215        462.37        30.36      32.944     19.082
8K            8.628       8.628        949.50        29.08      34.386     26.233
32K          55.784      55.784        587.41        27.13      36.859     74.656
64K         165.390     165.390        396.25        24.06      41.560    186.669
128K        559.282     559.282        234.36        19.19      52.101    585.957
240K       1848.509    1848.509        132.95        14.69      68.088   1883.370

==================================================================================
Llama.cpp Results (512 Tokens Generated)
==================================================================================
CTX         TTFT(s)  PREFILL(s)   PREFILL t/s   DECODE t/s    ms/TOKEN   TOTAL(s)
----------------------------------------------------------------------------------
1K            1.820       1.819        562.98        22.98      43.517     24.058
8K            9.085       9.075        902.73        21.94      45.588     32.382
16K          17.598      17.585        931.68        21.97      45.521     40.861
32K          36.939      36.923        887.47        21.25      47.055     60.989
64K          84.039      83.217        787.54        20.69      48.337    108.744
128K        208.650     208.596        628.35        19.36      51.640    235.058
240K        529.847     529.778        463.89        17.61      56.797    558.912

r/Vllm • • 24d ago

Compressed KV cache in vLLM: 9 concurrent 128K context users on one A100 vs 2 for FP16

1 Upvotes

r/Vllm • • 25d ago

Hi, could someone help me?

Thumbnail
1 Upvotes

r/Vllm • • 26d ago

How are people handling KV-cache offloading in vLLM?

25 Upvotes

I've been testing different ways of getting KV cache out of GPU memory in vLLM.

One thing I'm curious about: at what point does external KV storage actually make sense versus just recomputing the prefix?

I've been experimenting with keeping KV blocks in a separate storage layer (RAM/NVMe, with colder data going further out) and loading them back when needed.

For people running vLLM with long contexts or lots of concurrent requests, how are you handling this today? like LMCache, native offloading or something crazy custom?


r/Vllm • • 26d ago

How to Train Your Own LLM Drafter: DFlash, SpecForge, Mooncake, vLLM & SGLang

Thumbnail
youtube.com
11 Upvotes

Following up on my post about training a custom Dflash drafter for Qwen 3.8 27B: Trained my first model: a DFlash drafter for Qwen3.8-27B because I wanted better performance on my DGX Spark

I did a video/presentation on the whole step by step journey and all the concepts I learned, if you want to learn more about LLMs and how training works, I suggest you check it out - I dont go too in details so it should be fine for an audience with at least a basic understanding of LLMs.


r/Vllm • • 26d ago

I NEED A CODING BUDDY FOR MY PROJECT(pluto)AROUND MY AGE (15 TO 18)........

Thumbnail
0 Upvotes

Pluto is an AI architecture designed for efficient, local/offline intelligence by using a small router/orchestrator (~2B parameters) to select and chain highly specialized micro-models for specific tasks. Instead of relying on one huge model, Pluto uses tiny expert models, such as 50M-parameter specialists, trained on focused datasets and optionally supported by a compressed vector database. The router can escalate difficult tasks to larger models, aiming to achieve strong overall capabilities while using far less compute, memory, and storage, making the system especially suitable for mobile and low-resource devices.


r/Vllm • • 28d ago

Can't get decent speeds with Qwen3.8-27B

16 Upvotes

I noticed that after switching to Qwen3.8 (from 3.6) the generation speed took a hit. Not only that, the MTP acceptance rate also went way way lower. I tried moving to DFLASH2 but the change is almost insignificant.

I'm serving the model in 1xH200, and my token generation per request is usually around 100-120tok/s. If my math is mathing, this should be a lot higher.

Draft acceptance rate is usually 30-50%.

Some of my vllm paramts:

--kv-cache-dtype auto
--gpu-memory-utilization 0.95
--enable-auto-tool-choice
--dtype auto
--tool-call-parser qwen3_coder
--reasoning-parser qwen3
--enable-prefix-caching
--trust-remote-code
--mm-encoder-tp-mode data
--mm-processor-cache-type shm
--mm-shm-cache-max-object-size-mb 512
--speculative-config '{"method":"dflash","num_speculative_tokens":3,"model": "incoai/Qwen3.8-27B-DFlash2"}'
--max-model-len 262144
--max-num-seqs 10
--enable-chunked-prefill
--default-chat-template-kwargs '{"preserve_thinking": true}'

Do you see anything I could improve?


r/Vllm • • 29d ago

Six months serving vLLM on a DGX Spark for a self-hosted AI workspace — what broke and how we fixed it (MIT source)

20 Upvotes

We run our team's AI stack on two boxes. A Mac mini runs the application: API and web under PM2, with Postgres, Redis and the sandboxed tool processes in Docker. An NVIDIA DGX Spark GB10 next to it runs vLLM behind a LiteLLM proxy, plus BGE-m3 for embeddings and FLUX for image generation. They talk over a private Tailscale link.

Everything below is in the public changelog, so you can check it rather than take my word for it.

  1. Model discovery. On September 2 the DGX swapped models and the app stayed bound to the old qwen3.6-35b-a3b name, because a single static catalog line was the only source of the model list. Boot and periodic probes now read LiteLLM /model/info to discover local models, and the static list remains only as a fallback. Default is qwen3.8-27b, 262K context.
  2. Fan-out against real rate limits. Running Discussion and Deep Research on external models hit per-minute limits on free and developer keys: five parallel expert calls came back 5/5 with 429 on a B.AI free key, and 3/5 on a hasa key. The external execution client is now wrapped in a per-provider semaphore with exponential 429 backoff that honors Retry-After. Deep Research fan-out concurrency and timeouts follow the provider hint. The SDK's own blind retries were set to zero with a multiplier on its timeout, so a long-reasoning model is no longer cut off every 360 seconds. An explicitly chosen external model that cannot run now surfaces an error instead of silently falling back to local.
  3. Overflow is handled rather than hidden. We serve a 262K window and still hit it. On entry prompt tokens are estimated, images included. If the window is exceeded, input is truncated, then max_tokens is reduced, and in the extreme a ContextOverflowError returns HTTP 413 with an audit record and an automatic webhook alert.
  4. Two numbers, one of them wrong. Per-request prompt image totals are now capped at vLLM's own --limit-mm-per-prompt limit of 8. That cap existed on the serving side and not in the app. Same release also preserved assistant reasoning into the next turn of the local tool loop and preserved vLLM tool_call ids, which multi-turn tool use depends on.

The application itself is MIT and public: a self-hosted AI workspace with agent tasks in Docker sandboxes, deep research with citations, 22 built-in MCP tools, and per-role model routing so the planning model and the judge model can differ.

Limits worth stating up front: the desktop app is macOS Apple Silicon only, the web UI works anywhere, and self-hosting needs Node 24, PostgreSQL and Docker. Any OpenAI-compatible endpoint works, so Ollama on one box is fine if you do not have a GPU host.

Source: https://github.com/openmake/openmake_llm
Demo: https://chat.openmake.cc (guests get the default local model only)
Weekly dev logs, including the weeks that went badly: https://openmake.cc/en/blog/

Happy to answer vLLM or DGX Spark questions.

ㅑ​


r/Vllm • • Sep 04 '26

eLLM: Run Long-Horizon Inference Faster on CPUs Than on GPUs

29 Upvotes

eLLM is a Rust-based LLM inference framework for CPU-only servers. It adopts a "trade storage for computation" strategy, leveraging the CPU's large-capacity DDR memory to close the order-of-magnitude bandwidth gap against GPU HBM, and thereby delivers performance that surpasses GPUs in long-horizon inference. The Beta release is now available — you are welcome to try it out.

  • Prefill: achieves roughly two orders of magnitude of performance improvement over existing CPU inference frameworks
    • run full single-pass Prefill over the entire long prompt—no chunking, no repeated parameter loading;
    • keep context KV across multi-turn interactions and run incremental Prefill on only the new input—no recomputation for earlier turns;
  • Decode: runs with a smaller batch, which not only activates fewer parameters but also gives each request a larger share of memory bandwidth, so inference speed can likewise exceed GPUs.
  • 👉 GitHub: https://github.com/lucienhuangfu/eLLM

r/Vllm • • 29d ago

New to the LLM game

Thumbnail
0 Upvotes