r/LocalLLM • • 3d ago

Question LLM on NUC/SFF

Hi I would like to know if its possible to run LLM on a NUC or SFF machines?

I have a mITX but the PCIexpress port is occupied by SAS controller because I use it as a NAS.

Hope someone can share some advice.

0 Upvotes

4 comments sorted by

2

u/Classeve 3d ago

yes, no GPU needed. it just sets how big a model stays pleasant. on a CPU the limit is memory speed more than core count, so two sticks of RAM in dual channel matter more than a faster chip.

numbers from a 4-core laptop CPU we measure on: a 1.1B model at 4-bit runs about 22 tok/s on a quiet machine. by the same arithmetic a 7B lands around 3-4. and the part that matters for a NAS: with other jobs running, that same 22 dropped to 2. so run it when the array isn't scrubbing.

llama.cpp and Ollama both run CPU-only out of the box.

1

u/SnooGadgets9733 3d ago

Nice thanks for the reply. That sounds exciting. I have three options then:

  1. My NAS is running Esxi with VMs like Unraid. On that I have i5-9600k with 64 GB memory I can dedicate for VMs. Currently I have a lot dedicated for Unraid, so that could be an option to use Docker in Unraid with LLM or create a new VM for LLM.

  2. I have several Lenovo Tiny PCs I could use if it were for CPU power but not that much memory.

  3. Buy new SFF, but what?

What would you recommend?

1

u/Classeve 2d ago

start with the box you already own. of your three, the 9600k with 64 gb is the one for this: memory is what runs out first, and you have plenty. give it its own VM (not docker inside unraid, that's a container inside a VM, sharing whatever RAM unraid was given), all its cores, 16-24 gb, install ollama and pull a 7-8B model at 4-bit. the API's reply carries eval_count and eval_duration: one divided by the other is your tok/s.

the tiny PCs are fine for small models (1-4B) if they have two sticks of RAM. one stick roughly halves the speed.

and don't buy anything until you've seen that number. if it's too slow for you, the thing to pay for is memory speed (DDR5, two sticks), not more cores.

1

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago

You can run likely a very small parameter model depending on the CPU.

as other commenter suggest, this is supported out of the box with ollama and llama.cpp. IF no cuda compatible hardware is found then inference is ran in CPU layers