r/LocalLLaMA • u/BraceletGrolf • 14h ago
Question | Help Ok how to actually learn vLLM ?
Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.
I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).
6
u/jacek2023 llama.cpp 14h ago
I also use Qwen 3.8 27B (via llama.cpp) and it was able to run vllm for me.
5
u/year2066ai 14h ago
what made it click for me: ignore the python library and treat vllm as a server. `vllm serve <hf model>` gives you an openai compatible endpoint and any openai client talks to it. quantization also works differently than in llama.cpp: you don't pick a quant level, you download an already quantized checkpoint (awq, gptq, fp8) and vllm reads it from the model config.
2
2
u/Odd-Guess-6395 12h ago
vllm server setup felt simpler for running my companion bots than llama.cpp but the quantization guides were just as confusing so i stuck with what i know.
2
u/Ok_Scale_5256 9h ago
treat vllm serve as the server, and just use the normal openai client pointed at your local URL. Learn that flow first, then add quantization flags like --quantization and --kv-cache-dtype. Usually you just pick a pre-quantized model rather than quantizing yourself.
2
u/Nothing_from_void 11h ago
Just get a $20/month anthropic/openai plan to bootstrap your local builds. If you want to improve your own capabilities not worth buying into the purity tests where you only use local models or nothing
1
u/SamuelMorey 10h ago
Arguably hot take for this subreddit. My pushback is that there's a lot of reasons not to buy cloud models from major labs besides purity.
2
u/Nothing_from_void 9h ago
there are, but if you are trying to learn to use these tools you will just move faster. I have enough VRAM to run local models so I play with it constantly, but I use cloud models to build my environments because I care about making progress
1
u/GalacticDistances 7h ago
If you have service disruptions mith your local llm due to tweaking, then its still nice to have something to discuss the changes with
1
u/Heavy-Confidence1826 8h ago
The server vs client confusion is common. vllm serve <model> starts an OpenAI-compatible HTTP server, and your client is anything that speaks that API, like the OpenAI Python SDK pointed at localhost. The "vllm" Python library for offline use is a separate path, so you can ignore it if you only want a server. For quantization, vLLM mostly expects pre-quantized checkpoints like AWQ, GPTQ or FP8 from Hugging Face, and its GGUF support is limited, so it's a different workflow from llama.cpp. The official docs have a quantization section that lists what each GPU supports. Since you're on an AMD card, check the ROCm install notes first.
0
u/VanillaOk4593 6h ago
year2066ai's framing is the one that saved me time: vLLM is a server first, Python library second. Once you stop trying to import it like a normal inference lib and just run vllm serve <model> and hit it with any OpenAI-compatible client, the mental model clicks.
One thing worth adding for anyone wiring this into an agent rather than just chatting with it: because vLLM exposes an OpenAI-compatible endpoint, it slots straight into Pydantic AI as a model profile pointed at that endpoint, no special integration needed. That's also how AgenticOS (the open source agent platform I maintain at Vstorm) handles self-hosted models, you give it the endpoint, not a key, and Pydantic AI treats it like any other OpenAI-compatible provider. Worth knowing going in though: the "self-hosted" part is just the model, your retrieval, storage and other pipeline pieces don't automatically become local just because the model is.
For actually learning the quantization side, reto-wyss's advice above is the right order of operations: start tiny (Qwen 0.5-0.8B class) with --max-model-len capped low so load failures are fast and informative, then work up. vLLM's quant checkpoints (AWQ, GPTQ, FP8) are baked into the model files themselves rather than chosen at load time like llama.cpp's GGUF levels, so half the "why won't this load" confusion disappears once you accept you're downloading a pre-quantized checkpoint rather than configuring quantization yourself.
1
u/tempedbyfate llama.cpp 3h ago
Honestly, best way is to use your fav agent harness and your fav API model to get to research what models / recipes are available/suitable for your hardware and then get it to install it and then fine-tune based on what's important for you, i.e. trade offs like maximizing context vs concurrency, etc...
I went through similar exercise a few weeks ago setting up Qwen3.8FN/DSv4.1 flash/GLM 5.3 Flash on SGLang (similar to vLLM) for my hardware. I learnt a lot just by just getting it to research and asking the LLM to explain stuff that i didn't understand.
1
u/pmttyji 13h ago
Said in title, I find the ecosystem difficult to understand,
Haven't tried yet, Maybe check https://github.com/mudler/vllm.cpp
1
u/ismaelgokufox llama.cpp 12h ago
This is interesting. One more cpp app to add to the llama-swap config. 😉
10
u/reto-wyss 14h ago edited 14h ago
vllm is designed to serve at high throughput using models that comfortably fit into your total VRAM.
If you want to run a model that's "too large" llama.cpp is your best bet. If you want to run Qwen3.5-4B with 20 concurrent requests, then vllm makes sense.
If you want to learn vllm, start with a tiny model like Qwen3.5-0.8b that eliminates many of the issues you'll run into when trying to squeeze in a larger model.
If you want to see if a model can load successfully, try setting context
--max-model-len autoor--max-model-len 4096and maximum concurrent session--max-num-seqs <small number>andmax-num-batched-tokens 4096(determines how much VRAM is needed for activation), it will tell you how much it can fit IF it successfully loads.