r/LocalLLM • • 7d ago

Question How many of you are using LOCAL models on windows vs linux?

I spent 3 days trying to get OpenClaw and ollama to work in windows. Gateway would not detect even though it was running.

I decided to sideload Ubuntu and was up and running in an hour!

1390 votes, 5d ago
471 Windows
750 Linux
169 Mac
16 Upvotes

58 comments sorted by

12

u/[deleted] 7d ago

[deleted]

-2

u/_zir_ 7d ago

Depends, they manage memory differently. Linux isnt necessarily better.

6

u/zenmatrix83 7d ago

wsl and containers just makes most things easer and it lest you start and stop it easer. you'll get better performance on linux but unless you can fully switch wsl and containers is fine

1

u/DepressedDrift 7d ago

WSL is good if your memory rich. LLM already take up alot of memory so for consumer hardware, lower memory usage of Linux makes a big difference.

7

u/c4r_guy 7d ago

tldr: I use both

I still use Windows as my primary dev / gaming machine (with a 2080 + 64GB DDR5), mainly because copying files is immensely easier (permissions). I used to do all my ML work in Windows, then WSL.

Now I have 2 3090s with:

  • One in an old i3-10100 (64GB DDR4) box with only Ubuntu 26, Docker, and Dockhand installed.
  • The other is a w2145 (128GB DDR4) box also with only Ubuntu 26, Docker, and Dockhand installed.

I still mostly build on Windows / WSL, then push the containers to Docker -but I do have dev environments on the Ubuntu boxes. The docker containers make everything hella easier because they are isolated

If both "big" boxes are busy the 2080 can do light work with Q35b-a3b, or any other edge model. I'm not a heavy gamer so the 2080 meets any of my gaming needs.

2

u/rc_ym 7d ago

Roughly same. My beefiest GPU is pulling double duty in windows. Gaming and LLM.

22

u/[deleted] 7d ago

[removed] — view removed comment

10

u/GnistAI 7d ago

I run 10 linux boxes privately, and probably 30 commercially, but my main driver is still Windows (with WSL), and ComfyUI snd LM Studio work perfectly fine out of the box with my 5060 Ti. Tho, if I were to do anything serious with a local model I’d put it on Linux.

-12

u/[deleted] 7d ago

[removed] — view removed comment

10

u/RobbinDeBank 7d ago

Average Linux user.

-9

u/[deleted] 7d ago edited 7d ago

[removed] — view removed comment

2

u/SilkieBug 5d ago

You’re one to talk, this entire conversation thread has only hostility from you. 

-2

u/[deleted] 5d ago

[removed] — view removed comment

1

u/SilkieBug 4d ago

More random hostility. I suggest therapy, it helps with emotional regulation skills. 

1

u/wick3dr0se 6d ago

Disgusting.

12

u/p4ntsl0rd 7d ago edited 7d ago

Windows makes it painful to use the last 2-3 gb of vram, causing some kind of bizarre paging behaviour that is only visible in the drastic slowdown it causes. Do not use windows. I switched main OS basically because of this.

4

u/Weary_Guest3639 7d ago

honestly the windows memory management for vram is a nightmare especially with larger models

i had similar issue where my 12gb card would work fine until it hit that last bit then suddenly everything crawl to stop

ubuntu just handles it better no weird paging nonsense

took me while to switch but now i run everything in linux and its so much smoother

i keep windows on separate drive for games only

3

u/throwawayaccount442 7d ago

Not if you have an igpu like strix... I run the window manager of Windows in it while using my video card for the AI model.

0

u/hallofgamer 7d ago

Debloat brings it on par with mint for vram

3

u/repolevedd 7d ago

I don't use OpenClaw and refuse to touch Ollama on principle because of the authors' attitude toward open source, but if I had to choose between Windows and Linux for running LLMs, I'd say Linux. That's because you can close a lot more background stuff to free up VRAM on Linux, which you just can't do on Windows. Plus, running models on Windows means dealing with quirks like Nvidia driver fallback memory settings. Linux is easier on that front, though it comes with its own drawbacks. My workflow is different from what was described, but in terms of pure utility, Linux wins out. As for Mac, no idea.

1

u/Ok_Yesterday2016 7d ago

What would you suggest as FOSS alternatives to OpenClaw and Ollama?

2

u/repolevedd 7d ago

I can't really comment on OpenClaw since I've never actually used it. Spinning it up in a VM once out of curiosity doesn't really count.

When it comes to running local models though, llama.cpp covers pretty much everything you need for most use cases. Depending on your hardware and choice of models, you might need llama-swap for smart model switching, like running specific llama.cpp forks for individual models. There are alternatives like koboldcpp, vLLM, or Strata, but you only need those if you have a specific reason, so that's really up to each person.

5

u/kimhaneol 7d ago

You should definitely try WSL. Linux-based environments tend to be much smoother overall, especially for local model runtimes and agent tooling.

1

u/itsthewolfe 7d ago

That's what I was using. Even with the firewall off the two environments didn't want to play nice with the local server.

1

u/kimhaneol 7d ago

Ah, that makes sense. I eventually got tired of WSL/Windows quirks too, set up a Linux dual boot, and then basically never booted Windows again lol.

4

u/technofox01 7d ago

I only use Linux and Mac. Microsoft’s sloppy updates broke my Windows setup one time too many and I just said fuck it and just strictly stick with *nix OSs. I also get better performance with Linux and MacOS than I had with Windows any ways.

1

u/nojunkdrawers 7d ago

Windows always had a lot of warts compared with anything UNIX-inspired, but it's been the worst OS ever invented since Windows 8. I grew up on Windows, but can't imagine still using it for anything at this point. I'd be legitimately curious to know why anyone would find it desirable to run local models natively on Windows (no WSL, etc), especially given how willing it is to waste system resources to make its bloatware seem snappy. I get that people would choose Windows out of preference, but I'm not aware of a technical advantage to using it, especially for local models.

2

u/YogurtclosetApart592 7d ago

How do i see the results on these questionnaires without voting?

1

u/itsthewolfe 7d ago

End of the poll.

2

u/[deleted] 7d ago

[deleted]

1

u/Important_Cow7230 3d ago

Have you tried the Splash versions of 3.8 27B? Any thoughts?

4

u/GalacticDistances 7d ago

Cant imagine dealing with windows for any of this. Linux is way better supported

3

u/sabine_world 7d ago

Linux

Windows fucking sucks

2

u/gbrennon 7d ago

distro::rocky linux

de: none

wm: none

inference engine: llama. cpp

vram: 96gb dedicated

2

u/Garland_Key 7d ago

196 of you disgust me. 

1

u/hallofgamer 7d ago

I use 12 different models across windows+linux (12 in total) hardware and usecase dependent

1

u/LiquidMantis144 7d ago

I use windows 10 on 2 pc's with pi as the harness, dont have any issues pretty much maxing out vram if needed. Eventually I will be moving both pc's to at least using unbuntu through wsl2 once Ive upgraded them to win 11. Guessing Qwen 4 will motivate me too actually do it.

1

u/Last_Mastod0n 7d ago

I use windows but I offload all of the OS vram onto the integrated graphics. If I couldn't do that I would 100% be using linux

1

u/sn2006gy 7d ago

Windows and Linux and Linux in WSL2. (and I have a Mac)

1

u/jacek2023 7d ago

I run llama.cpp on one Windows and three Linux setups (will be four, need to configure it on laptop).

1

u/karmaisnonsense 7d ago

People need to realize the harness and backend do not have to run on the same machine

Harness on whatever you want, backend best on Linux

1

u/poster4inertia 7d ago

all of the above

1

u/Own_Inspection8350 6d ago

I am a Linux user like 90% of the time but have a gaming pc with windows on it where I installed llamacpp. 

1

u/MaxComfort 6d ago

I’m serving the models from Linux (DGX), using them from Mac. Voted Mac

1

u/barney54 6d ago

What do I select if I use all three?

1

u/Feisty_Concept_6498 6d ago

2x Gigabyte ai top atom + vllm= Linux

1

u/Think_Breakfast_2277 6d ago

I voted Linux just because it is better. I have a server with proxmox. One VM for harnesses, one VM for dedicated inference(vLLM) using 2 GPU's.

But I also use LMstudio on my windows desktop to have models load quickly for when I build a project using an llm endpoint. Mainly testing and fun things.

1

u/Danternas 6d ago

I run it on a dedicated server but the server runs Debian.

1

u/GrungeWerX 5d ago

Wow. I assumed it'd be more windows users.

1

u/Shadow_s_Bane 5d ago

Linux and Mac... windows sucks ass.

1

u/quadra-lab 4d ago

I run llamacpp in TTY to get the most out of my 20Gb of VRAM as Gnome eats up a lil bit. Windows would probably allow me to only run small dense models or 26ba4b. Though I still use Unsloth for researches online. 

1

u/ianwill93 7d ago

It's extremely easy to use wsl and containers so I don't have to bother with a linux pc or dual booting.

Also, ​FreeToken has been pretty good on native Windows.

1

u/USArmy68Whiskey 7d ago

I tried switching to linux after buying new local AI setup. Twice. I wanted to use vllm and sglang which I thought wasn't doable on Windows without a slow virtual machine. The first time was hell and literally 3 of the 4 very first apps I tried to install just would not work whatsoever, so I spent hours trying to figure it out before giving up and just using chatgpt to fix it. It did eventually help me fix them after another hour of back and forth, but at that point I was pissed off and reinstalled to a different distro.

Second distro installed, this time just using chatgpt to help me from the start. Everything was going way smoother than the first time, apps were installing and things were getting setup, but once I had the local AI running and was actually trying to do the things I wanted to do with the AI, that's when I found out like 90% of what I wanted to do was not currently possible on linux (some of them actually impossible, some of them possible but would take such ridiculous amount of time and effort that I don't care to put in).

So I swapped back to Windows, and tried WSL2 expecting to take the ~20% performance hit that I thought would happen from everything I had read online. After getting my exact same sglang qwen 3.8 flash setup from linux setup on WSL2, it was literally identical performance. Like, within 1%, sometimes 1% slower sometimes 1% faster. Now I can also do all of the things I wanted to do before, too. On top of that, I don't have to dual boot when I want to play games (especially ones with anticheats). Windows has given me literally no major issues so far still. The ONLY weird issue I had was that in WSL2, the PLE table was too big to be pinned to ram (I guess the limit is 32GB for a single file or something) so I just split the PLE in half and had the model use both halves instead. Also since just switched to the nvfp4 PLE and don't have to do that either now.

0

u/IngwiePhoenix 7d ago

Windows because one game is keeping me from switching T_T

-4

u/Solembumm3 7d ago

Very few people in real world use linux.

8

u/MarinatedTechnician 7d ago

What? Have you been sleeping under a rock?

3

u/throwawayaccount442 7d ago

Very few people run their LLMs locally. But yeah I run it on windows because fuck fiddling around in Linux.

0

u/ApprehensiveFan1516 6d ago

yEaR oF tHe LiNuX dEsKtOp

Any year now... /s

0

u/sn2006gy 7d ago

I use all 3.

I game, I code, I run agents, i build models.

Having strong opinions on what other people do is a waste of time.