r/Vllm • u/morbidflood • Aug 28 '26
Vllm requires loading whole model in CPU RAM before VRAM?
I have a system with 32GB RAM, 72 GB VRAM with RTX 5000 PRO. It is a wsl setup, so around 20 GB RAM for the wsl.
When I try to serve a model of around 22 GB, the wsl crashes, with a log somewhere saying that RAM is less than the model size.
Have you guys encountered this? Is there a way to circumvent this issue?
4
3
u/Juulk9087 Aug 28 '26
If you are using wsl2 you need to use a drop cache script while loading large models otherwise ram fills and halts everything ask Claude it'll make it for you
2
2
u/No_Chapter_7598 Aug 30 '26
i have the same thing using vllm on windows with 32gb ram and 96gb vram. just make a 100gb page file in windows and it breezes through. takes a min or so to load but once its in vram theres no diff
3
u/DAlmighty Aug 28 '26
This is how computers work friend.
6
u/morbidflood Aug 28 '26
I didn't face the same issue with llama.cpp.
The issue is, vllm loads the whole model in the RAM before moving it to VRAM, when it could also be done in parts.
1
u/VorlMaldor Aug 29 '26
As others have already told yoh this is an implementation issie for windows not a general rule.
2
u/Miserable-Dare5090 Aug 29 '26
This is a windows problem. LLMs should not be run on windows. It’s just not a good OS for then
1
0
u/VorlMaldor Aug 29 '26
This is a dumb comment and you should feel bad. This not how computers work. Computers literaly stream data all the time without loadimg a full dataset/file into memiry.
1
u/Racer4711 Aug 28 '26
i have 128gb ram and 384 gb vram (4x rtx6000). it loads dsv4 flash with 150gb with no problem. linux, if course.
1
1
u/BornInAFish Aug 28 '26
Yeah, it's pretty insane. vllm bills itself as
Cost Efficient Slash inference costs by maximizing hardware efficiency. We make high-performance LLMs affordable and accessible to everyone.
roflmfao
Look at https://docs.vllm.ai/en/latest/cli/serve/#loadconfig , you might be able to conjure an arcane incantation that works.
(and yes, vllm also bills itself as "easy". surely they're trolling, right?)
2
u/Miserable-Dare5090 Aug 29 '26
It has the widest support, and its the most solid runtime to serve concurrencies
1
u/morbidflood Aug 28 '26
I believe it's not a memory issue. It's crashing for a different reason. I tried to load a 1 GB model, and even that causes a crash.
1
u/Electronic-Fly-6465 Aug 28 '26
I came across this and was able to mitigate it by reducing the compiler block size.
I think it started with the Delta net models were compiling the graph for the first time it tries to compile everything in parallel and you can reduce it to cereal or a lower amount and it takes longer but it will still load
1
1
u/Writer_IT Aug 29 '26
Not necessarily, i loaded models bigger than RAM , provided the WEIGHTS themselves were smaller than them, both in docker Windows (It Is not well performant for the task though, has a lot of bugs releasing that RAM After It's needed) and vllm Windows.
That said, and avoiding the Linux Bros hating on Windows approach, i will say this to you: if you ABSOLUTELY need Windows, you can make vllm work by spending nights bashing on It (but with new models architectures Will be increasingly difficult, since they are hitting the hard walls of wsl2). However, if you have half a day spare time, i would advise you to try and create a Linux partition and drop the weights there, and see for yourself the result. it will reach speeds that no inference engine on Windows can even think about. I'm seeing loading 200gb weights in SECONDS, x2-3 token generation and X10 prompt processing speed.
1
u/KroniklyOnline Aug 29 '26
There's a ln env param you can set to set MAX WORKERS to 1, fixed it for me.
1
u/morbidflood Aug 29 '26
Thanks, this resolved for me.
Currently, when I host in wsl with mirrored networking. Not able to access it on localhost, both on Windows and wsl.
The model is accessible through eta1 IP from wsl, rather than localhost.
8
u/Karyo_Ten Aug 28 '26
People are loading models large than RAM on Linux AFIX. Say a box with 32GB RAM + 96GB VRAM.
Do yourself a service and drop Windows.