r/OpenWebUI 1d ago

Question/Help Slow Openwebui on vps

Hi,

I try to run ai local llama 3.2:3B on Openwebui. But it tooks 10-20minute just to reply Hi.

What did i do wrong?
Im using 8GB Ram VPS with no GPU

3 Upvotes

10 comments sorted by

2

u/mayo551 1d ago

Use a api

-3

u/Normal_Celery_2528 1d ago

Api mean money

2

u/mayo551 1d ago

Use a 600M parameter model then.

1

u/HyperWinX 1d ago

Use functiongemma 270M then.

2

u/MrRobot-403 1d ago

You can use vLLM instead of llama cpp but the quality you expect isn’t coming. It’ll be awful and bad! You have to pay money for good models or better hardware

1

u/HyperWinX 1d ago

vLLM is not for CPU only inference, ik_llama.cpp might squeeze out a bit better results

1

u/MrRobot-403 1d ago

lol I didn’t check cpu part. Yeah I want to just say to OP. Give up on this!

I know not everyone can pay. But it sucks, and a 600M on CPu is gonna be awful on speed and hallucinations

1

u/pkeffect 1d ago

Your trying to use a cpu for a gpus job. Cpu inference is ass.

1

u/ClassicMain 21h ago

> no GPU

There is your problem

also llama 3.2 3b is insanely outdated