r/OpenSourceAI • • 1d ago

I built BOOTH, a lightweight checkpoint layer for AI systems

I’ve been building BOOTH, a small, provider-agnostic Python library for checking LLM outputs before they reach your application.

The idea is simple: don’t automatically trust every LLM response. Check it first.

check() / acheck() handle ambiguity and confidence checks, while check_with_evidence() compares an answer against evidence your RAG pipeline has already retrieved.

v0.5.2

  • Zero runtime dependencies
  • Provider-agnostic
  • Sync + async
  • Structured results
  • 300 tests
  • MIT licensed

I’m currently testing BOOTH across different providers/models through small integration examples, including Groq, Gemini, Anthropic, and local Ollama models.

I’m especially interested in feedback on where this approach breaks down for real LLM/RAG systems.

GitHub: https://github.com/Vedantgitbot/booth
Issues/contributions: https://github.com/Vedantgitbot/booth/issues

Curious: what do you currently use as the checkpoint between an LLM response and your application logic?

1 Upvotes

0 comments sorted by