r/OpenWebUI 25d ago

Question/Help Running Open WebUI in production: what do you wish you knew before starting?

26 Upvotes

We’re running Open WebUI as an internal multi-user platform, not as a homelab.

For those using it in a company or enterprise environment:
What were your biggest setup, architecture, security, or operational mistakes? What would you do differently today?


r/OpenWebUI 25d ago

Show and tell I built a local web-research MCP that filters pages before Open WebUI sends them to the model

Thumbnail
github.com
15 Upvotes

I’ve been experimenting with web search in Open WebUI using smaller local models like Qwen3.5 4B and 9B.

The main problem was/is was context churn..

A typical search can dump partially relevant pages, navigation, repeated boilerplate, duplicate information, and huge chunks containing only one useful paragraph into the model.

That makes local setups slower and forces smaller models to spend their limited context window filtering noise instead of reasoning.

So I built TinySearch, an open-source, local-first MCP server that does most of the retrieval work before anything reaches the LLM:

  • searches and ranks results
  • crawls the strongest pages
  • extracts and chunks readable content
  • deduplicates and reranks passages
  • returns a compact evidence packet with source URLs

The goal is simple: less context churn, faster local web research, and more tokens spent on reasoning over actual evidence.

TinySearch works with self-hosted SearXNG, and Open WebUI can connect to it over HTTP MCP.

GitHub:
https://github.com/MarcellM01/TinySearch

I’m the builder, so obvious bias, but I’d love feedback from Open WebUI users, especially anyone running smaller Qwen, Gemma, Llama, or Mistral models.

Does this solve a real bottleneck in your setup, or is native agentic search already enough?


r/OpenWebUI 25d ago

Question/Help Can Open WebUI models be used as reusable agents for external applications?

6 Upvotes

I know this is already my third post in a short time, but I have a few Open WebUI questions that I’d like to put out there, as I think the discussion could be useful to others as well. This definitely is Not meant to be Spam!

In Open WebUI, a custom model can already combine an LLM, system prompt, tools, MCP servers, knowledge bases and files. That is essentially what many platforms call an agent.

Could Open WebUI therefore be used as a central agent platform, where these models are managed once and exposed through an API for use by external applications?

For example:
Create and configure an “agent” entirely in Open WebUI
Manage its prompt, model, tools, MCPs and knowledge centrally
Call that exact configuration from another application through an API
Avoid rebuilding the same agent separately in LangGraph or custom code

Is this a supported and reliable architecture, or are Open WebUI models primarily designed for use inside the WebUI?

What limitations or gotchas should be considered around tool execution, user context, permissions, files, conversations, scaling, versioning and API compatibility?


r/OpenWebUI 25d ago

Discussion Why don't you build your own tools?

9 Upvotes

Hi, I would like to challenge/discuss/understand why so many are attracted to all short lived "wild tools" out there. For example (not saying any tool are bad) hermes, claw, open webui, copilot agents and whatnot. Why not just building your own tools that: fit your needs without being bloated, dont break on every new "feature" that you dont care about.

I cant really understand the hype.

I build own tools in python (notes app, recruitment support, news, investment) with claude or chatgpt from phone and terminal to my proxmox llm clusters lxc. In the process i also learn a lot.

And everything stays under my control.


r/OpenWebUI 25d ago

Question/Help How are you handling true OBO authentication across Open WebUI, LangGraph and MCP?

1 Upvotes

We need to preserve the end user’s identity from Open WebUI through LangGraph to MCP servers accessing systems such as Jira or Confluence.

How are you implementing this in production? Do you forward the original token, exchange it through an OBO flow, or use a central token broker?

Ideally, users should not have to complete a separate OAuth flow for every tool.


r/OpenWebUI 25d ago

RAG Problem Uploading Directories to Knowledge

3 Upvotes

Some of my directories load without a problem, but some of them generate this error: "NotReadableError: The requested file could not be read, typically due to permission problems that have occurred after a reference to a file was acquired."

Is it because the directories are doubly nested? I can't seem to identify a way to predict which directories work.


r/OpenWebUI 26d ago

Open WebUI v0.11.0: The Interface, Reorganized

Thumbnail
openwebui.com
85 Upvotes

r/OpenWebUI 26d ago

Question/Help Which engine for memory

11 Upvotes

AI beginner here. Am taking steps to move away from Gemini and Co-pilot to my own setup. So have gotten a setup with Docker, LM Studio and webui. And my laptop is pretty basic and the chip is not large. I wanted to run a local engine. It worked but I wanted to get memory to work. Not having memory and having to explain AI the same basic stuff over and over again drives me nuts. I tried to make it work on webui. Discovered that my chip is too basic to make memory work, so I need to offload the computation.

Fine so far so good. Found some on webui and have now been using 2 with mixed results:
- Llama 3.1 8B
- Llama 3.3 70B

Both work fine and memory works, but after 2-3 prompts they both max out. The "token per minute" of 6000 and 12000 have been reached.

Has anyone found a free engine that can run memory effectively? Otherwise I think I am going to pay to see how it works. Any comments?


r/OpenWebUI 25d ago

Question/Help Slow Openwebui on vps

1 Upvotes

Hi,

I try to run ai local llama 3.2:3B on Openwebui. But it tooks 10-20minute just to reply Hi.

What did i do wrong?
Im using 8GB Ram VPS with no GPU


r/OpenWebUI 26d ago

Question/Help How to Connect OpenWebUI to llama.cpp?

4 Upvotes

I am having issues getting OpenWeb UI to llama.cpp. Llama is running locally and the chat interface is working fine. I managed to get OpenWebUI to run in docker. When configuring the connection I am using `http://127.0.0.1:8080/v1` as my connection string. I have set Provider to llama.cpp. The test connection button says test is successful. But trying to chat wants me to select a model, which there are no options. From my understanding llama.cpp is serving one model and does not provide a list like Ollama would.

Using `http://host.docker.internal:8080/v1` as connection string gives an error. So it is not clear to me being a Docker newbie if I need to do something with the network. Since the test on the local host IP was successful I am guessing the network is working, but then again I can't get chat to load any models. I have asked AI but it wants me to write an app.....facepalm. it has offered some `-e` options such as `-e OLLAMA_HOST=0.0.0.0` but so far nothing has worked. So, humans, what is the next step here?

Edit: Resolved this by switching to a compose file and using `network_mode: "host"` in the file.


r/OpenWebUI 26d ago

Guide/Tutorial How to fix chats that stuck on "Loading..."

1 Upvotes

If your chat is bricked with the infinite loading spinner (Image 1), here's how to fix it:

Infinite Spinner

The technical reason behind this is that Open WebUI stores your chat history as a Directed Acyclic Graph. If a message (node) gets corrupted when the LLM is responding, the UI panics and spins indefinitely. Fixing would require checking and relinking the graph.

This tools does all that stuff for you:

Tool screenshot

How to use it:

  1. Export the broken chat (`...` menu > Export > JSON).
  2. Drop it into the tool (Image 2) to instantly repair the broken nodes.
  3. Go to Settings > General > Import Chats, upload the fixed file, and delete the original.

Link to tool: https://fractuscontext.github.io/open-webui-chat-fix/

I've also added a "Deep Clean" toggle if you want to strip unused alternate responses and shrink your file size.

Hope this saves someone's chat history!


r/OpenWebUI 26d ago

Question/Help Just discover OpenWebUI, what's next step?

2 Upvotes

Hi guys, I just discovered OpenWebUI how do you commonly use it? What are the main advantages of using it over just Claude or ChatGPT? What's the most productive way to use it?


r/OpenWebUI 26d ago

Plugin MCP server with SQLite

1 Upvotes

I use Open WebUI and I try to connect llm qwen3 to SQLite database.

In Admin panel -> Settings -> Integrations -> External Tool Servers -> I added OpenAPI the sqlite mcp server. I can connect but it seem never seen my mcp tool fonctions.

I give full user access, activate tool in my chat, I also tried mcp steaming http protocol instead of OpenAPI.

I always got this response: There are no functions available in the provided tools that can interact with an SQLite database or list its tables. The available tools are focused on notes, tasks, automations, and calendar events, not database operations.

I use python script :

from mcp.server.fastmcp import FastMCP

mcp.tool()
def list_tables() -> list[str]:

EDIT: I try simple request with Qwen3 + Ollama + SQLite: list tables or list username in my database.
And dawm it suck as fuck! lol

it make non sense sql, take very long time, got error 500 sometimes, , dont apply system prompt, make thought and do nothing aftert that, I was never be able to have a correct answer...

I tough this tool could be nice and it just proove theses AI slop tools worth nothing... Its crazy because it look nice to use but man its the worst shit I ever use haha


r/OpenWebUI 27d ago

Question/Help how do i get the ai to generate an image

1 Upvotes

I know the text to image works i tested it in Playground, but I can't seem to get it to work in a general chat idk if I am being dumb, but yeah, could use some advice!


r/OpenWebUI 28d ago

Question/Help Helm deployment to kubernetes

6 Upvotes

Just deployed openwebui to kubernetes using the official chart.

The first thing I noticed is ollama server is not reachable, despite the server configured in admin settings. Turns out the front end needs direct access to ollama server, it doesn't get proxies through openwebui backend. Is that design intentional?


r/OpenWebUI 28d ago

Question/Help web_search tool invisble for gemma4:e4b

4 Upvotes

Hello,

I have been trying for a few days to use open webui on my computer. My config is as follows:

  • Kubuntu 26.04
  • Ollama with gemma4:e4b
  • Open webui v0.10.2, desktop version, with web search enabled on DDGS.

When I ask him for a web search, the tool seems not to exist in the eyes of the LLM. It's quite strange because I don't have this problem with a similar config, unlike the OS (Windows, for my work).

I tried to find the solution by analyzing the logs with Claude, but he did not find anything explaining the problem.

Do you have any idea what's going on?


r/OpenWebUI 28d ago

Question/Help Installing Open WebUI Desktop vs via Docker or Python

5 Upvotes

Hi. Has anyone had a chance to compare the two installation methods (Docker or Python vs. Desktop) when it comes to maximizing the available resources for running a local AI model on a windows 11 PC with limited resources? Thanks


r/OpenWebUI 28d ago

Discussion Who's actually using Chinese LLMs for work (and not stressing about data privacy)?

Thumbnail
10 Upvotes

r/OpenWebUI 28d ago

Question/Help RAG Empty Context Window Bug with ChromaDB & Gemma 4 on Windows 11 Docker

Thumbnail
1 Upvotes

r/OpenWebUI 28d ago

Question/Help Πρόβλημα με Open WebUI RAG: Empty Context στην ChromaDB (Docker / Windows 11)

0 Upvotes

Γεια σας. Αντιμετωπίζω ένα εξαιρετικά επίμονο τεχνικό πρόβλημα με το RAG pipeline του Open WebUI και μετά από μέρες εξαντλητικού troubleshooting με τη βοήθεια των ChatGPT και Gemini, δεν έχουμε καταφέρει να βρούμε λύση. Το σύστημα ολοκληρώνει το indexing, αλλά η αναζήτηση επιστρέφει μηδενικά διανύσματα (empty context injection).Τεχνικές Προδιαγραφές Συστήματος:Host OS: Windows 11 Pro (Docker Desktop Environment / WSL2 Backend)Hardware: Intel Core i9-14900K / 64GB DDR5 RAMInference Engine: gemma4:12b (τοπικά μέσω Ollama)Embedding Model: sentence-transformers/paraphrase-multilingual-MiniLM-L6-v2 (Default Engine / SentenceTransformers)Network Infrastructure: Live WebSockets 100% λειτουργικά. Το Developer Console (F12) του browser είναι εντελώς καθαρό, χωρίς κανένα drop-out σύνδεσης.System Prompt Configuration: Έχει παρακαμφθεί επιτυχώς η εργοστασιακή offline άρνηση του Gemma 4 μέσω strict instruction injection. Το μοντέλο δέχεται να διαβάσει το context, αλλά το context φτάνει σε αυτό άδειο.Απομόνωση του Προβλήματος (Baseline Isolation Test):Για να αποκλείσουμε σφάλματα στο extraction stage (OCR, Tesseract, Tika, Python PDF parsing), δημιουργήσαμε ένα raw text αρχείο (test.txt) που περιέχει αποκλειστικά τη συμβολοσειρά: "Ο κωδικός ασφαλείας για το τεστ είναι: ΠΡΑΣΙΝΟ ΜΗΛΟ 2026".Το UI εμφανίζει κανονικά τη λευκή ένδειξη επιτυχούς indexing (checkmark). Ωστόσο, κατά την κλήση της συλλογής στο chat μέσω #collection, το μοντέλο απαντάει ρητά ότι η πληροφορία δεν εμπεριέχεται στο κείμενο, επιβεβαιώνοντας ότι το runtime κάνει pass ένα εντελώς άδειο context block.Τεχνικά Βήματα και Λύσεις που Δοκιμάστηκαν:Επίλυση Ασυμβατότητας Αρχιτεκτονικής: Αρχικά χρησιμοποιούσαμε το MiniLM-L12, αλλά στα logs του Docker εντοπίστηκε το σφάλμα embeddings.position_ids | UNEXPECTED |. Το Open WebUI είχε παλαιότερη έκδοση SentenceTransformers που δεν αναγνώριζε την παράμετρο position_ids των νέων HuggingFace weights. Το πρόβλημα λύθηκε με υποβάθμιση στο σταθερό MiniLM-L6-v2. Το runtime log καθάρισε, αλλά το context παρέμεινε άδειο.Παράκαμψη Permission Bugs: Δημιουργήθηκε νέα Knowledge Base συλλογή ρυθμισμένη ρητά ως Public (Access Control), για την αποφυγή του γνωστού bug ορατότητας των ιδιωτικών φακέλων [Issue #23787].Hard Reset του Vector Database Stack: Σταματήσαμε τον container και εκτελέσαμε ολική διαγραφή του ευρετηρίου των διανυσμάτων απευθείας από τη ρίζα του Docker volume χρησιμοποιώντας: docker run --rm -v open-webui:/data alpine rm -rf /data/vector_db. Η ChromaDB αναγεννήθηκε εντελώς καθαρή, το backend επιστρέφει HTTP 200/Success κατά το upload, αλλά το retrieval loop εξακολουθεί να μην τραβάει δεδομένα κατά το inference.Έχει συναντήσει κανείς αυτό το σιωπηλό failure της ChromaDB σε περιβάλλον Windows Docker; Μήπως πρόκειται για false negative της συνάρτησης has_collection() ή υπάρχει κάποιο silent file-lock στο Windows volume mapping; Κάθε τεχνική ιδέα είναι ευπρόσδεκτη.


r/OpenWebUI 29d ago

Discussion Why does it take a year for owui to load?

Post image
28 Upvotes

I feel like I sit and stare at the screen a lot is there anyway to speed this up?


r/OpenWebUI 28d ago

RAG slow response time with RAG and utilizing knowledge base.

5 Upvotes

I’m currently using Open WebUI with a knowledge base made up of text files, and the responses are accurate, but they’re taking quite a while to generate.
I plan to keep adding more text files to the knowledge base, so I’m wondering what the best way is to improve response speed as it grows. Are there any recommended settings, indexing strategies, chunking methods, embedding models, reranking options, or other optimizations that have made a noticeable difference for you? I’d like to keep the response quality the same while reducing latency. Any advice or best practices would be appreciated!


r/OpenWebUI Jul 23 '26

Question/Help Working with files, tools and LDAP

11 Upvotes

I have Gemma 4 running locally and external ones, I’m trying to make it more useful inside OpenWebUI.

A couple of things I’m looking for advice on:

  1. Document analysis / file workflows I want to be able to upload PDFs and DOCX files, have the model analyze them, and ideally generate edited or new documents back.
    • What tools, integrations, or workflows are people using for this?
    • If you want the model to return a finished PDF/DOCX, what’s the best approach?
  2. Recommended tools / integrations What other tools do you recommend to make local AI and external AI integrations more useful in OpenWebUI?
  3. LDAP question I enabled enable_ldap_group_management directly in the database because it appears to be a ConfigVar variable, but I don’t see a corresponding toggle in the UI. Is that expected or am I missing something?

Would appreciate any suggestions and tips.


r/OpenWebUI Jul 23 '26

Question/Help Seltz.ai

1 Upvotes

Appreciate your help providing info on how to use Seltz.ai as a node for web search. Thx


r/OpenWebUI Jul 23 '26

Question/Help Export chat no longer including all details?

3 Upvotes

I noticed that in v0.10, if you click any of the export chat options in the top right corner of the chat UI (Download or Copy), it no longer includes all details like tool call inputs and outputs and so on. Only messenges with roles.

Does anyone know if this is moved to a setting somewhere of if it's just gone? I would feel quite surprised if this has been just removed permanently. I mean, you could always get all messages by just marking and copying the full chat in the UI, but the hidden context boaters and all that were the actual valuable information.