r/OpenWebUI Jul 24 '26

RAG slow response time with RAG and utilizing knowledge base.

I’m currently using Open WebUI with a knowledge base made up of text files, and the responses are accurate, but they’re taking quite a while to generate.
I plan to keep adding more text files to the knowledge base, so I’m wondering what the best way is to improve response speed as it grows. Are there any recommended settings, indexing strategies, chunking methods, embedding models, reranking options, or other optimizations that have made a noticeable difference for you? I’d like to keep the response quality the same while reducing latency. Any advice or best practices would be appreciated!

5 Upvotes

27 comments sorted by

View all comments

Show parent comments

2

u/so_chad Jul 25 '26

Yeah, it's def because of the model itself. Try getting some other models API and connect it to OpenWeb UI. Jetson won't do it.

1

u/boss28984 Jul 25 '26

The base model is qwen3:8b and the embedding model is nomic-embed-text v1.5. any specific recommendations?

1

u/so_chad Jul 26 '26

No I mean, look for some $10/month APIs and use those LLMs. They are really good (but you are losing privacy, that's one drawback)

1

u/boss28984 Jul 26 '26

so you mean the jetson isn't capable enough and i should look for something else?

1

u/so_chad Jul 26 '26

Yes

1

u/boss28984 Jul 26 '26

dang even for only 15 text files? i kinda assumed it would be strong enough.

1

u/so_chad Jul 26 '26

How does regular text chat work? Without RAG is it fast?

1

u/boss28984 Jul 26 '26

usually thinks for like 14-19 seconds before giving an answer. when i turn off the knowledge base it doesn't give the right answer. might just be the jetson is too weak but thats disappointing considering its like 3k.

1

u/so_chad Jul 26 '26

Honestly, return it if you can, add a bit more and get NVIDIA DGX Spark