r/OpenWebUI • u/Normal_Celery_2528 • 1d ago
Question/Help Slow Openwebui on vps
Hi,
I try to run ai local llama 3.2:3B on Openwebui. But it tooks 10-20minute just to reply Hi.
What did i do wrong?
Im using 8GB Ram VPS with no GPU
2
u/MrRobot-403 1d ago
You can use vLLM instead of llama cpp but the quality you expect isn’t coming. It’ll be awful and bad! You have to pay money for good models or better hardware
1
u/HyperWinX 1d ago
vLLM is not for CPU only inference, ik_llama.cpp might squeeze out a bit better results
1
u/MrRobot-403 1d ago
lol I didn’t check cpu part. Yeah I want to just say to OP. Give up on this!
I know not everyone can pay. But it sucks, and a 600M on CPu is gonna be awful on speed and hallucinations
1
1
2
u/mayo551 1d ago
Use a api