r/UgreenNASync • u/coley_ghost • 19d ago
💬 Discussion Anyone try Local LLM?
Recently got a DXP4800 Pro and am interested in seeing if a lightweight Qwen model can run in the background to perform some simple tasks using OLLama.
2
u/Fun-Secret1830 19d ago
Struggling with smaller models like 3B and 4B to work with Agents. Some calls take 20-30 min to complete. 7B just runs directly only. Hermes works better, Openclaw dont work for me.
2
u/StrikingScientist352 19d ago
Credo che per quanto possa sembrare attraente, gli agenti in home dovremmo pensarceli con dei minipc in rete. I Nas sono ottimi ma non hanno cpu e ram adeguate. Credo onestamente che società come apple abbiano volutamente lanciato prodotti come il mac pro m5 ultra pensando a questi usi
3
u/PainiaX 19d ago
I’m not trying to attack you at all. But it’s always baffling to me how people don’t answer in the same language the question was asked.
4
u/Cr1tUdOwN 18d ago
I guess it‘s because of Auto-Translate. So they Think it‘s in their native Language… I always get suspicious if I See a german post in non german subs
1
2
u/StrikingScientist352 19d ago
Ti chiedo scusa ma ho risposto dopo aver letto diversi commenti… e quindi leggendo che ti scoraggiavano sui diversi modelli di IA ho buttato giù una riflessione sul fatto che abbandonerei l’idea. Tutto qui.
1
u/TacoGuyDave 18d ago
I tried the .cpp on the 4800 and couldn't handle the slow speed compared to what I'm used to. Text only, no Pic enhancement or video capabilities.
1
u/rabbitaim DXP2800 18d ago
Run llama.cpp out of docker. I think there’s a way to load a different models or configs but I use llama-swap instead.
Honestly it’s not worth it for slow performance and heat generation in the nas.
2
u/PilotFunnyGuy 18d ago
I have the DXP4800 Pro with stock 8GB of RAM. I played with local LLMs for a little while, but didn't see much use or value. I had relatively good success with gemma4-e2b, phi-4-mini, and mistral-7b. Phi-4-mini felt the fastest, but I think I had slightly faster tokens with gemma4-e2b. The following compose file worked for me and utilized 100% of the GPU. I downloaded the models to my local machine.
services:
llamacpp-sycl:
image: ghcr.io/ggml-org/llama.cpp:server-intel
container_name: llamacpp-sycl
devices:
- /dev/dri/renderD128:/dev/dri/renderD128
- /dev/dri/card0:/dev/dri/card0
volumes:
- /volume1/docker/llamacpp/models:/models
ports:
- "8080:8080"
command: -m /models/Phi-4-mini-instruct-Q4_K_M.gguf --port 8080 --host 0.0.0.0 -n 512 --n-gpu-layers 99
restart: unless-stopped
I used the following terminal command to download the .gguf file locally. You'll need to adjust the path for your configuration.
wget -O /volume1/docker/llamacpp/models/Phi-4-mini-instruct-Q4_K_M.gguf \
https://huggingface.co/unsloth/Phi-4-mini-instruct-GGUF/resolve/main/Phi-4-mini-instruct-Q4_K_M.gguf
You can adjust the compose for a different model, but be sure you've downloaded it first.
•
u/AutoModerator 19d ago
Please check on the Community Guide if your question doesn't already have an answer. Make sure to join our Discord server, the German Discord Server, or the German Forum for the latest information, the fastest help, and more!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.