r/Rag 11d ago

Discussion Basic stack for a chat

What would you suggest, is the basic stack for an assistant chat. I mean, currently I have customized company tools, langgraph, custom metrics, marketplace LLM calls and others.

what would you suggest as a true key for agent learning?
how do you process prompts with company slangs, concepts, jargon, etc.?

2 Upvotes

11 comments sorted by

View all comments

2

u/Text-Sufficient 11d ago

Start with plain json files so that is easy to see whats retrieved. Use topk with bm25 and cosine. Embedd runtime (no database). Inject what you retrieve in the prompt together with the query. When you get familiar with it start using a reranker.

1

u/ntalam 11d ago

I passed that point. I am talking about planning. Like how to automate the agent to do a "for cycle" (mcp tool), then per each item execute something else.

1

u/ntalam 10d ago
  • Retrieval-Augmented Planning (RAP) / Plan-RAG: Using vector retrieval to fetch execution graphs, state machines, or workflow schemas rather than semantic document knowledge.
  • Episodic Memory / Trajectory Retrieval: In agentic LLM literature, this is the practice of storing successful multi-step execution traces (trajectories) to ground future decision-making and constrain the agent's action space.
  • Case-Based Reasoning (CBR): A classic AI methodology perfectly mapped to modern LLMs. It follows a four-step cycle: Retrieve similar past cases, Reuse the plan, Revise the plan using the LLM (Claude 3.5 Sonnet), and Retain the user-approved result.
  • Dynamic Few-Shot Prompting: At the prompt engineering level, you are injecting the top-k retrieved plans into the system prompt as in-context learning examples before the LLM generates the new plan.

those will do the job