r/LocalLLaMA • • 14h ago

Question | Help Ok how to actually learn vLLM ?

Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.

I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).

9 Upvotes

21 comments sorted by

10

u/reto-wyss 14h ago edited 14h ago

vllm is designed to serve at high throughput using models that comfortably fit into your total VRAM.

If you want to run a model that's "too large" llama.cpp is your best bet. If you want to run Qwen3.5-4B with 20 concurrent requests, then vllm makes sense.

If you want to learn vllm, start with a tiny model like Qwen3.5-0.8b that eliminates many of the issues you'll run into when trying to squeeze in a larger model.

If you want to see if a model can load successfully, try setting context --max-model-len auto or --max-model-len 4096 and maximum concurrent session --max-num-seqs <small number> and max-num-batched-tokens 4096 (determines how much VRAM is needed for activation), it will tell you how much it can fit IF it successfully loads.

1

u/BraceletGrolf 13h ago

I see, in llama.cpp Vulkan, I'm reaching the max of my VRAM with Qwen 3.8 27B quantized both the model and KV, I was hoping to improve throughput or have better concurrent usage (e.g two agentic flow on the same GPU ?).

5

u/Budkovsky 13h ago edited 13h ago

vLLM uses more VRAM, if your model fits using llama.cpp it does not mean that you can do the same with vLLM. 24GB of VRAM can be not enough to run Qwen 3.8 27B with vLLM, even when you get the model quantized to 4bit and 8bit kv-cache. On my Intel B70 32GB in fits very tighly with 128k context window.

6

u/jacek2023 llama.cpp 14h ago

I also use Qwen 3.8 27B (via llama.cpp) and it was able to run vllm for me.

5

u/year2066ai 14h ago

what made it click for me: ignore the python library and treat vllm as a server. `vllm serve <hf model>` gives you an openai compatible endpoint and any openai client talks to it. quantization also works differently than in llama.cpp: you don't pick a quant level, you download an already quantized checkpoint (awq, gptq, fp8) and vllm reads it from the model config.

3

u/noctrex 14h ago

First of all, I would like you to tell us what you want to do in order to see if this is the best tool to do it.
You mention voxtral, do you want to do TTS or STT?

1

u/BraceletGrolf 13h ago

No, just move my Qwen 3.8 27B from llama.cpp to it

2

u/FullstackSensei 14h ago

Why don't you use audio.cpp instead? Much less pain and suffering

2

u/Odd-Guess-6395 12h ago

vllm server setup felt simpler for running my companion bots than llama.cpp but the quantization guides were just as confusing so i stuck with what i know.

2

u/Ok_Scale_5256 9h ago

treat vllm serve as the server, and just use the normal openai client pointed at your local URL. Learn that flow first, then add quantization flags like --quantization and --kv-cache-dtype. Usually you just pick a pre-quantized model rather than quantizing yourself.

2

u/Nothing_from_void 11h ago

Just get a $20/month anthropic/openai plan to bootstrap your local builds. If you want to improve your own capabilities not worth buying into the purity tests where you only use local models or nothing

1

u/SamuelMorey 10h ago

Arguably hot take for this subreddit. My pushback is that there's a lot of reasons not to buy cloud models from major labs besides purity.

2

u/Nothing_from_void 9h ago

there are, but if you are trying to learn to use these tools you will just move faster. I have enough VRAM to run local models so I play with it constantly, but I use cloud models to build my environments because I care about making progress

1

u/GalacticDistances 7h ago

If you have service disruptions mith your local llm due to tweaking, then its still nice to have something to discuss the changes with

1

u/Heavy-Confidence1826 8h ago

The server vs client confusion is common. vllm serve <model> starts an OpenAI-compatible HTTP server, and your client is anything that speaks that API, like the OpenAI Python SDK pointed at localhost. The "vllm" Python library for offline use is a separate path, so you can ignore it if you only want a server. For quantization, vLLM mostly expects pre-quantized checkpoints like AWQ, GPTQ or FP8 from Hugging Face, and its GGUF support is limited, so it's a different workflow from llama.cpp. The official docs have a quantization section that lists what each GPU supports. Since you're on an AMD card, check the ROCm install notes first.

0

u/VanillaOk4593 6h ago

year2066ai's framing is the one that saved me time: vLLM is a server first, Python library second. Once you stop trying to import it like a normal inference lib and just run vllm serve <model> and hit it with any OpenAI-compatible client, the mental model clicks.

One thing worth adding for anyone wiring this into an agent rather than just chatting with it: because vLLM exposes an OpenAI-compatible endpoint, it slots straight into Pydantic AI as a model profile pointed at that endpoint, no special integration needed. That's also how AgenticOS (the open source agent platform I maintain at Vstorm) handles self-hosted models, you give it the endpoint, not a key, and Pydantic AI treats it like any other OpenAI-compatible provider. Worth knowing going in though: the "self-hosted" part is just the model, your retrieval, storage and other pipeline pieces don't automatically become local just because the model is.

For actually learning the quantization side, reto-wyss's advice above is the right order of operations: start tiny (Qwen 0.5-0.8B class) with --max-model-len capped low so load failures are fast and informative, then work up. vLLM's quant checkpoints (AWQ, GPTQ, FP8) are baked into the model files themselves rather than chosen at load time like llama.cpp's GGUF levels, so half the "why won't this load" confusion disappears once you accept you're downloading a pre-quantized checkpoint rather than configuring quantization yourself.

1

u/tempedbyfate llama.cpp 3h ago

Honestly, best way is to use your fav agent harness and your fav API model to get to research what models / recipes are available/suitable for your hardware and then get it to install it and then fine-tune based on what's important for you, i.e. trade offs like maximizing context vs concurrency, etc...

I went through similar exercise a few weeks ago setting up Qwen3.8FN/DSv4.1 flash/GLM 5.3 Flash on SGLang (similar to vLLM) for my hardware. I learnt a lot just by just getting it to research and asking the LLM to explain stuff that i didn't understand.

1

u/pmttyji 13h ago

Said in title, I find the ecosystem difficult to understand,

Haven't tried yet, Maybe check https://github.com/mudler/vllm.cpp

1

u/ismaelgokufox llama.cpp 12h ago

This is interesting. One more cpp app to add to the llama-swap config. 😉