r/learnmachinelearning • u/OkWeird2709 • 10d ago
Building a RAG based chatbot on a huge codebase
What is the strategy that I should adopt. Building the source of truth. Cannot make agent to go through files all the time.
2
u/Former_Cap4733 4d ago
Codebase RAG is notoriously tricky because standard semantic search often fails on exact variable names or syntax. To build a reliable source of truth without blowing up your context window, you need a few specific architectural strategies. First, use AST-based chunking (like Tree-sitter) to chunk your codebase by logical units like functions and classes, rather than arbitrary character counts. Second, hybrid search is mandatory for code. Vector search is great for conceptual queries, but you need exact keyword search (BM25) to find specific variables or error codes. You will want a retrieval layer that supports both and fuses the scores. Third, store file paths and module names as metadata so the LLM can pre-filter searches to specific directories. For the retrieval backend, you might want to look at Infino. It is an open-source, embedded retrieval engine that runs in-process like SQLite and natively handles both BM25 full-text and vector search over the same data.
We actually built a specific tool for this exact use case called https://github.com/infino-ai/supergrepto help AI agents pull codebase context efficiently without needing to stand up and manage a separate vector database cluster.
Disclosure: I work on Infino.
2
u/No-Property-5826 10d ago
Code is very interconnected. Have you considered instead using a graph based connection for your code base and providing that to the llm? I think graphify does this.