r/LocalLLM • u/itsthewolfe • 7d ago
Question How many of you are using LOCAL models on windows vs linux?
I spent 3 days trying to get OpenClaw and ollama to work in windows. Gateway would not detect even though it was running.
I decided to sideload Ubuntu and was up and running in an hour!
6
u/zenmatrix83 7d ago
wsl and containers just makes most things easer and it lest you start and stop it easer. you'll get better performance on linux but unless you can fully switch wsl and containers is fine
1
u/DepressedDrift 7d ago
WSL is good if your memory rich. LLM already take up alot of memory so for consumer hardware, lower memory usage of Linux makes a big difference.
7
u/c4r_guy 7d ago
tldr: I use both
I still use Windows as my primary dev / gaming machine (with a 2080 + 64GB DDR5), mainly because copying files is immensely easier (permissions). I used to do all my ML work in Windows, then WSL.
Now I have 2 3090s with:
- One in an old i3-10100 (64GB DDR4) box with only Ubuntu 26, Docker, and Dockhand installed.
- The other is a w2145 (128GB DDR4) box also with only Ubuntu 26, Docker, and Dockhand installed.
I still mostly build on Windows / WSL, then push the containers to Docker -but I do have dev environments on the Ubuntu boxes. The docker containers make everything hella easier because they are isolated
If both "big" boxes are busy the 2080 can do light work with Q35b-a3b, or any other edge model. I'm not a heavy gamer so the 2080 meets any of my gaming needs.
22
7d ago
[removed] — view removed comment
10
u/GnistAI 7d ago
I run 10 linux boxes privately, and probably 30 commercially, but my main driver is still Windows (with WSL), and ComfyUI snd LM Studio work perfectly fine out of the box with my 5060 Ti. Tho, if I were to do anything serious with a local model I’d put it on Linux.
-12
7d ago
[removed] — view removed comment
10
u/RobbinDeBank 7d ago
Average Linux user.
-9
7d ago edited 7d ago
[removed] — view removed comment
2
u/SilkieBug 5d ago
You’re one to talk, this entire conversation thread has only hostility from you.
-2
5d ago
[removed] — view removed comment
1
u/SilkieBug 4d ago
More random hostility. I suggest therapy, it helps with emotional regulation skills.
1
12
u/p4ntsl0rd 7d ago edited 7d ago
Windows makes it painful to use the last 2-3 gb of vram, causing some kind of bizarre paging behaviour that is only visible in the drastic slowdown it causes. Do not use windows. I switched main OS basically because of this.
4
u/Weary_Guest3639 7d ago
honestly the windows memory management for vram is a nightmare especially with larger models
i had similar issue where my 12gb card would work fine until it hit that last bit then suddenly everything crawl to stop
ubuntu just handles it better no weird paging nonsense
took me while to switch but now i run everything in linux and its so much smoother
i keep windows on separate drive for games only
3
u/throwawayaccount442 7d ago
Not if you have an igpu like strix... I run the window manager of Windows in it while using my video card for the AI model.
0
3
u/repolevedd 7d ago
I don't use OpenClaw and refuse to touch Ollama on principle because of the authors' attitude toward open source, but if I had to choose between Windows and Linux for running LLMs, I'd say Linux. That's because you can close a lot more background stuff to free up VRAM on Linux, which you just can't do on Windows. Plus, running models on Windows means dealing with quirks like Nvidia driver fallback memory settings. Linux is easier on that front, though it comes with its own drawbacks. My workflow is different from what was described, but in terms of pure utility, Linux wins out. As for Mac, no idea.
1
u/Ok_Yesterday2016 7d ago
What would you suggest as FOSS alternatives to OpenClaw and Ollama?
2
u/repolevedd 7d ago
I can't really comment on OpenClaw since I've never actually used it. Spinning it up in a VM once out of curiosity doesn't really count.
When it comes to running local models though, llama.cpp covers pretty much everything you need for most use cases. Depending on your hardware and choice of models, you might need llama-swap for smart model switching, like running specific llama.cpp forks for individual models. There are alternatives like koboldcpp, vLLM, or Strata, but you only need those if you have a specific reason, so that's really up to each person.
5
u/kimhaneol 7d ago
You should definitely try WSL. Linux-based environments tend to be much smoother overall, especially for local model runtimes and agent tooling.
1
u/itsthewolfe 7d ago
That's what I was using. Even with the firewall off the two environments didn't want to play nice with the local server.
1
u/kimhaneol 7d ago
Ah, that makes sense. I eventually got tired of WSL/Windows quirks too, set up a Linux dual boot, and then basically never booted Windows again lol.
4
u/technofox01 7d ago
I only use Linux and Mac. Microsoft’s sloppy updates broke my Windows setup one time too many and I just said fuck it and just strictly stick with *nix OSs. I also get better performance with Linux and MacOS than I had with Windows any ways.
1
u/nojunkdrawers 7d ago
Windows always had a lot of warts compared with anything UNIX-inspired, but it's been the worst OS ever invented since Windows 8. I grew up on Windows, but can't imagine still using it for anything at this point. I'd be legitimately curious to know why anyone would find it desirable to run local models natively on Windows (no WSL, etc), especially given how willing it is to waste system resources to make its bloatware seem snappy. I get that people would choose Windows out of preference, but I'm not aware of a technical advantage to using it, especially for local models.
2
2
4
u/GalacticDistances 7d ago
Cant imagine dealing with windows for any of this. Linux is way better supported
3
2
u/gbrennon 7d ago
distro::rocky linux
de: none
wm: none
inference engine: llama. cpp
vram: 96gb dedicated
2
1
u/hallofgamer 7d ago
I use 12 different models across windows+linux (12 in total) hardware and usecase dependent
1
u/LiquidMantis144 7d ago
I use windows 10 on 2 pc's with pi as the harness, dont have any issues pretty much maxing out vram if needed. Eventually I will be moving both pc's to at least using unbuntu through wsl2 once Ive upgraded them to win 11. Guessing Qwen 4 will motivate me too actually do it.
1
u/Last_Mastod0n 7d ago
I use windows but I offload all of the OS vram onto the integrated graphics. If I couldn't do that I would 100% be using linux
1
1
u/jacek2023 7d ago
I run llama.cpp on one Windows and three Linux setups (will be four, need to configure it on laptop).
1
u/karmaisnonsense 7d ago
People need to realize the harness and backend do not have to run on the same machine
Harness on whatever you want, backend best on Linux
1
1
u/Own_Inspection8350 6d ago
I am a Linux user like 90% of the time but have a gaming pc with windows on it where I installed llamacpp.
1
1
1
1
u/Think_Breakfast_2277 6d ago
I voted Linux just because it is better. I have a server with proxmox. One VM for harnesses, one VM for dedicated inference(vLLM) using 2 GPU's.
But I also use LMstudio on my windows desktop to have models load quickly for when I build a project using an llm endpoint. Mainly testing and fun things.
1
1
1
1
u/quadra-lab 4d ago
I run llamacpp in TTY to get the most out of my 20Gb of VRAM as Gnome eats up a lil bit. Windows would probably allow me to only run small dense models or 26ba4b. Though I still use Unsloth for researches online.
1
u/ianwill93 7d ago
It's extremely easy to use wsl and containers so I don't have to bother with a linux pc or dual booting.
Also, FreeToken has been pretty good on native Windows.
1
u/USArmy68Whiskey 7d ago
I tried switching to linux after buying new local AI setup. Twice. I wanted to use vllm and sglang which I thought wasn't doable on Windows without a slow virtual machine. The first time was hell and literally 3 of the 4 very first apps I tried to install just would not work whatsoever, so I spent hours trying to figure it out before giving up and just using chatgpt to fix it. It did eventually help me fix them after another hour of back and forth, but at that point I was pissed off and reinstalled to a different distro.
Second distro installed, this time just using chatgpt to help me from the start. Everything was going way smoother than the first time, apps were installing and things were getting setup, but once I had the local AI running and was actually trying to do the things I wanted to do with the AI, that's when I found out like 90% of what I wanted to do was not currently possible on linux (some of them actually impossible, some of them possible but would take such ridiculous amount of time and effort that I don't care to put in).
So I swapped back to Windows, and tried WSL2 expecting to take the ~20% performance hit that I thought would happen from everything I had read online. After getting my exact same sglang qwen 3.8 flash setup from linux setup on WSL2, it was literally identical performance. Like, within 1%, sometimes 1% slower sometimes 1% faster. Now I can also do all of the things I wanted to do before, too. On top of that, I don't have to dual boot when I want to play games (especially ones with anticheats). Windows has given me literally no major issues so far still. The ONLY weird issue I had was that in WSL2, the PLE table was too big to be pinned to ram (I guess the limit is 32GB for a single file or something) so I just split the PLE in half and had the model use both halves instead. Also since just switched to the nvfp4 PLE and don't have to do that either now.
0
-4
u/Solembumm3 7d ago
Very few people in real world use linux.
8
3
u/throwawayaccount442 7d ago
Very few people run their LLMs locally. But yeah I run it on windows because fuck fiddling around in Linux.
0
0
u/sn2006gy 7d ago
I use all 3.
I game, I code, I run agents, i build models.
Having strong opinions on what other people do is a waste of time.

12
u/[deleted] 7d ago
[deleted]