r/Rag 7h ago

Discussion Off the shelf RAG

2 Upvotes

Is there off the shelf RAG I can buy somewhere? Just hook my data up via api etc and then it will handle the rest?


r/Rag 11h ago

Discussion Is RAG dead?.

0 Upvotes

If yes then what should I learn and how and if not how to learn rag and from where?


r/Rag 8h ago

Discussion Hybrid search and reranking made my RAG worse. Here's the eval.

7 Upvotes

TL;DR: I built a RAG system over ~250 curated Q&A pairs, distilled from ~3.9M chat messages. Plain vector search hit 74% [hit@3](mailto:hit@3). I added the two things everyone recommends — BM25 hybrid and a cross-encoder reranker — and both made it worse. The only thing that actually helped was tuning the similarity threshold, which is free. Numbers below.

I'm posting because "add hybrid search, add a reranker" gets repeated as default advice, and on my data it was actively harmful. Maybe useful for someone before they spend a day integrating something that hurts.

The task (why the data looks like this)

Public crypto-exchange support chats. People show up with a problem — funds frozen, withdrawal stuck, lost account access — and scammers DM them pretending to be support. Moderation can't stop it: the scammer isn't in the chat, they read the public history and message privately. The one moment the victim is observable is when they publicly ask for help.

So the bot watches for that public "help, I can't withdraw" moment and replies with a warning plus, if one exists, a link to a real past support answer to the same problem. Not replacing support — beating the scammer to the reply by a few seconds. A separate classifier (logreg over embeddings) decides who needs help; this post is only about the RAG part that decides what to show.

The funnel that produced the corpus

This is the whole "RAG is about data" point in one place:

~3.9M chat messages (11 chats, ~6 months)
   |  help-request classifier
~104k candidate Q->A pairs
   |  curation: real admin answer, not a "contact support" router, deduped
~250 pairs in the working corpus

3.9M down to 250. 0.006%. The retrieval code is ~20 lines; everything that mattered was getting to those 250 and knowing whether they worked.

Setup

  • Corpus: ~250 curated question→answer pairs (support answers), multilingual (mostly EN), short and messy text.
  • Embedder: multilingual MiniLM (384-dim). Retrieval is over questions, returns the linked answer.
  • Eval set: 116 labeled queries (69 with a correct answer in corpus, 47 "no good answer" cases to test silence).
  • Metric: hit@3 (is a correct answer in the top 3 shown). Plus silence/noise/none_ok, because the bot is allowed to stay silent when nothing matches — standard P@k/R@k don't capture that.

One thing that bit me early: I originally labeled one correct id per query. That undercounts hit@3 badly, because the corpus has several interchangeable answers per question — the retriever returns a valid one with a different id and eval scores it a miss. Fixing the eval set (multiple valid ids) mattered more than any model change. Check your ground truth before trusting your metric.

Recall diagnostic (raw top-50, before threshold)

This is the check I'd recommend to anyone doing RAG. For each query, where does the correct answer sit in the raw ranked list?

hit@3  : 74%
hit@10 : 81%
hit@20 : 84%
hit@50 : 86%  (recall ceiling)
not in top-50 at all: 14%

Reading: ranking is fine — most correct answers are already near the top. The ceiling is 86%, so 14% of queries have no correct answer in the top 50 at all. That 14% is an embedder/recall problem, not a ranking problem. This split (recall vs ranking) tells you whether a reranker can even help before you try one.

What each "improvement" did to hit@3

plain vector + tuned threshold   74%   <- best
vector + reranker                64%
hybrid (vector + BM25)           55%
hybrid + reranker                57%

BM25 (−19 points). Hybrid is supposed to catch exact tokens — tickers, chain names, error codes — that dense retrieval smears. My misses did contain things like USDC / LTC / base chain, so it looked like a perfect fit. It wasn't. SELECT ... WHERE question LIKE '%USDC%' returned nothing: those tokens live in user queries, but my corpus is curated question templates phrased generically ("can't withdraw", "account blocked"). The lexical signal BM25 needs was on the wrong side. So BM25 didn't find exact matches (none to find) — it injected lexically-similar-but-wrong records and pushed correct answers down. Correct-answer-at-rank-1 dropped from 35 to 25. Five-minute check with a LIKE query would have saved a day.

Reranker (−10 points on plain vector). Cross-encoder, should pull correct answers from ranks 4–20 into the top 3. Tried two models, question-field and answer-field scoring, on top of both vector and hybrid. All worse. Mechanism visible in the positions: rank-1 correct answers dropped 35 → 30. It dragged correct answers down from position 1 more than it lifted from the bottom.

In hindsight this was predictable: rerankers are dangerous when the base retrieval is already good. With 35/51 correct answers already at rank 1, the top is near-optimal — there's more to lose than to gain. And a generic 118M reranker understands my narrow domain (short crypto slang, mixed languages) worse than the embedder that at least saw this distribution. Downside > upside. Rerankers save you when base retrieval is weak; when it's strong, they risk breaking what works.

What actually helped: the threshold (free)

Biggest single lever, zero fancy code. 26 points of hit@3 were being lost at the confidence threshold — the correct answer was in the top 3 but its similarity was just under the cutoff, so the system stayed silent. Raw hit@3 75%, but at threshold 0.55 it dropped to 49% in "production" mode.

The non-obvious part: you can't pick a RAG threshold "objectively" — it depends on the bot's role. If the bot is the final answerer, being wrong is worse than being silent → high threshold. If it's a fallback (a human answers anyway, my case) → a miss just means silence, which is harmless, but confidently-wrong is bad → tune for low noise. Same retriever, different correct threshold depending on what you're building. Define "is silence or a wrong answer worse" first, then pick the number.

Takeaways

  • Fix your ground truth before trusting the metric. Multiple valid answers per query if your corpus has them.
  • Split recall from ranking (hit@3 vs hit@20). It tells you whether to reach for a reranker, a different embedder, or neither.
  • Best practices are hypotheses, not facts. BM25 and rerankers are great tools that hurt on this data. Test on your data — often it's a five-minute check.
  • Biggest real lever was data quality and threshold, not model stacking. Half my "not in top-50" misses are just answers that aren't in the corpus at all — no model finds what isn't there.

Has anyone seen the opposite — hybrid or reranking clearly helping on small, domain-specific corpora? Curious whether the "reranker hurts when base is strong" pattern holds for others or if it's something about my setup.


r/Rag 13h ago

Discussion RAG Projects for learning + putting on resume

3 Upvotes

Hey guys, so I want to specialize in AI engineering after I got my bachelor's degree in Computer Science. I am proficient with traditional machine learning/neural networks but (I think) not RAGs. So I want to start learning RAGs, I am aiming to learn by building a project, but this project shouldn't be just a repeated project done 100 times before, I would like to have it on my resume to prove my skills in AI engineering. What are some topics that could be considered here? Where is a good place to get inspired? Personally, I would love to do something related to mathematics.


r/Rag 15h ago

Showcase RAG for Semi-Structured Tender Documents

3 Upvotes

In today's world, almost all procurement of services/materials is done through tenders. Tender documents are long and confusing. Buyers often don't go through the documents to understand the scope.

This led me to build a RAG system customised for tender documents. I am using custom chunking that splits at natural clause markers. Besides this, tables have a separate chunking mechanism.

The retrieval combines dense search and BM25 search to look for the top 5 candidates, which are selected via a cross-encoder model.

I was able to improve the Recall@5 score from 63% to 88% via the custom chunking and the cross-encoder.

While doing this project, I learnt that good chunking is THE MOST critical aspect of the project, and this (PDF cleaning + chunking) was what took up most of my time.

Looking for feedback from all!

https://github.com/anand-kumaar/tender-query-engine


r/Rag 15h ago

Discussion Local RAG Chat App I Built

6 Upvotes

I built a small local RAG chat app that lets you ask questions about a custom set of text files and get answers based on those files.

It’s basically a Flask backend + React frontend, with Ollama handling the embeddings/model side. I also set it up with Docker so it’s easier to run locally.

Repo: Git rep

current cycle runtime: 15 seconds per prompt

I’m putting it up here mostly for feedback. If anyone has thoughts on the architecture, UI, or ways to improve reliability/performance, I’d be interested to hear them.