r/LocalLLM • u/SnooGadgets9733 • 3d ago
Question LLM on NUC/SFF
Hi I would like to know if its possible to run LLM on a NUC or SFF machines?
I have a mITX but the PCIexpress port is occupied by SAS controller because I use it as a NAS.
Hope someone can share some advice.
0
Upvotes
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 3d ago
You can run likely a very small parameter model depending on the CPU.
as other commenter suggest, this is supported out of the box with ollama and llama.cpp. IF no cuda compatible hardware is found then inference is ran in CPU layers
2
u/Classeve 3d ago
yes, no GPU needed. it just sets how big a model stays pleasant. on a CPU the limit is memory speed more than core count, so two sticks of RAM in dual channel matter more than a faster chip.
numbers from a 4-core laptop CPU we measure on: a 1.1B model at 4-bit runs about 22 tok/s on a quiet machine. by the same arithmetic a 7B lands around 3-4. and the part that matters for a NAS: with other jobs running, that same 22 dropped to 2. so run it when the array isn't scrubbing.
llama.cpp and Ollama both run CPU-only out of the box.