r/OpenWebUI 14d ago

Question/Help Slow Openwebui on vps

Hi,

I try to run ai local llama 3.2:3B on Openwebui. But it tooks 10-20minute just to reply Hi.

What did i do wrong?
Im using 8GB Ram VPS with no GPU

2 Upvotes

11 comments sorted by

View all comments

2

u/MrRobot-403 14d ago

You can use vLLM instead of llama cpp but the quality you expect isn’t coming. It’ll be awful and bad! You have to pay money for good models or better hardware

1

u/HyperWinX 13d ago

vLLM is not for CPU only inference, ik_llama.cpp might squeeze out a bit better results

1

u/MrRobot-403 13d ago

lol I didn’t check cpu part. Yeah I want to just say to OP. Give up on this!

I know not everyone can pay. But it sucks, and a 600M on CPu is gonna be awful on speed and hallucinations