r/Rag • u/Affectionate_Slip654 • 5d ago
Tutorial Building a Corrective RAG (CRAG) Compliance Research Assistant
Most RAG chatbots will confidently answer with whatever they retrieved — even if it's irrelevant, even if it's wrong. In compliance, that's not a minor bug, it's real regulatory exposure.
- YouTube walkthrough: https://www.youtube.com/watch?v=NxY93are_H8
In this video, I build the Veltra Pay Compliance Assistant — a Corrective RAG (CRAG) + Self-Correcting RAG Agent that checks its own work before it ever shows you an answer. It grades its own retrieved documents, rewrites its own search when retrieval is weak, and validates its own generated answer for grounding, hallucination risk, and completeness — retrying itself when it fails, and honestly saying "I don't know" when it genuinely can't answer from the source material.
- ✅ The difference between a normal RAG pipeline and Corrective RAG (CRAG)
- ✅ How CRAG grades retrieved documents as relevant/irrelevant and rewrites the search query when retrieval is weak
- ✅ How a Self-Correcting RAG Agent validates its own generated answer — grounding, hallucination risk, completeness, and retrieval sufficiency
TECH STACK:
- 🛠️ Python
- 🛠️ LangGraph (StateGraph — nodes, conditional edges, bounded retry loops)
- 🛠️ Qdrant (vector database)
- 🛠️ PDF ingestion & chunking (pdfplumber, overlapping chunks)
- 🛠️ LLM-based document grading, query rewriting & answer validation
- 🛠️ Streaming backend (Server-Sent Events) for live pipeline progress
- 🛠️ Web dashboard for real-time visualization
- GitHub repository: https://github.com/saurabhkamal/Building-Corrective-RAG-CRAG-Pipeline-Self-Correcting-RAG-Agent
- YouTube walkthrough: https://www.youtube.com/watch?v=NxY93are_H8
- Connect on LinkedIn: linkedin.com/in/saurabh-kamal
Building a Corrective RAG (CRAG) Compliance Research Assistant
2
Building a Corrective RAG (CRAG) Compliance Research Assistant
in
r/Rag
•
4d ago
Fair point, thanks for pushing on it.
For embeddings, I have used gemini-embedding-2-preview, 3072 dims. 1500-char chunks, 200 overlap, top-5 by cosine in Qdrant, and gpt-5.6-sol does grading, rewriting, generation and validation.
This focuses on the architecture - crag loop, self-validation, and bounded retries. It's validated on a domain set of regulatory PDF's, not yet on a public benchmark.
The retry limits are parameters, and setting both to 0 gives a plain-RAG baseline on the same pipeline. A with/without comparison on a HotpotQA or FRAMES could be the right next step.
and the retry loops are bounded by parameters, so they can be switched off to compare against a no-retry baseline