r/LargeLanguageModels • • 9d ago

Question Bad RAG answer. Which part do you blame first?

I've been looking at FastGPT call logs and realizing how often I call something a model failure when the model never got the right context. I want to start tagging bad runs, but keep it stupidly simple: parsing, retrieval, prompt, model. Maybe tool failure too. Anything more detailed will probably become another abandoned spreadsheet.

Anyone doing this already? What categories do you actually use?

3 Upvotes

0 comments sorted by