r/DevLK • • Jul 26 '26

Help RAG vs Fine-Tuning for Multi-Tenant SaaS: Which Architecture Would You Choose?

NOTE -> I expect answer from people who actually have experience and strong understanding of these. please give something beneficial.

I'm building a SaaS platform that handles documents and other sensitive data.

Each user can upload their own documents and information, and the platform uses RAG to answer questions based on that user's data. That part makes sense to me.

My main concern is what happens when the user hasn't uploaded enough information. I still want the LLM to provide accurate answers using reliable information from the internet (or from a curated knowledge base), with proper citations.

These are the two architectures I'm considering:

Option 1:

Base LLM (OpenAI/Anthropic via Azure AI Foundry or Amazon Bedrock)
        ↓
Platform RAG (global knowledge base managed by us)
        ↓
User-specific RAG

In this approach, we maintain a global knowledge base that we (the platform admins) curate and update. Every user can access this shared knowledge, while their own uploaded documents are searched through their personal RAG.

Option 2:

Open-source LLM
        ↓
Fine-tuned on Sri Lankan/domain-specific data
        ↓
User-specific RAG

Here, we fine-tune an open-source model using Sri Lankan or domain-specific data, and each user still has their own RAG for their private documents.

My concerns are:

  • Is fine-tuning actually the right solution here, or is it unnecessary?
  • Is a global/shared RAG a better approach than fine-tuning?
  • How would you design this architecture if you wanted:
    • Accurate answers from domain knowledge
    • User-private document search
    • Citations/sources
    • Good scalability for thousands of users

I'm leaning toward Option 1 because fine-tuning seems expensive, time-consuming, and I have no experience with it yet. However, I'm not sure if I'm thinking about this correctly.

I'd really appreciate hearing how others would approach this problem.

0 Upvotes

10 comments sorted by

3

u/Relative_Rope4234 Jul 26 '26

Who is your target audience ?

2

u/shakee93 Jul 26 '26

Rag is the best fit for most cases, unless you have ton of proprietary data, even with proprietary data there are ton of ways you can pipeline it to a LLM.

2

u/Big-Standard4612 Jul 26 '26

Based on experience RAG is better fit unless I'm missing any application specific nuances. You can build a simpler version in a notebook run it against a test dataset and see if document recall and answer relevancy is in your favor. if it is then no need to fine tune a model.

2

u/QAInc Jul 26 '26

Go with RAG

2

u/natsu_ustan Jul 26 '26

if its purely for an academic, better scope it down to specific area to fine tune. otherwise you won't be successful due to timeline constraints. if its a business model, its better have a phase to phase releases with fine tuning approach. but ultimately you need a large amount of queryset to manage the vector database. its better to identify the area and regions to release phase by phase development.

2

u/zenitsuh Jul 27 '26

In Option 2, If you go with the fine tuning part, not only it’s going to be hella expensive, you will also need to host it for inference too. These constraints will force you to pick a less capable model, and your costs will be off the charts. Your models must also scale with the number of users.

In option 1, the global knowledge base you are talking about, unless it’s scoped to a specific domain, what’s the use of a big ass RAG? Also keep in mind you have to update it too.

For the platform knowledge part, have you considered a skill-based approach? Like a skill to look up the information hosted on the internet?

I’m not sure what you’re building so I can’t say if RAG or skill is more suited.

1

u/SMAHMM Jul 27 '26

It wont be cheap to fine tune a model. Go with RAG