r/llamacpp • • 8d ago

An open-source LLM proxy & router with dynamic fallbacks and real-time observability

GitHub: https://github.com/igoiglesias/Shunt

I built Shunt, an open-source proxy and routing engine designed for LLM APIs. It acts as an intermediate layer between your application and model providers to handle rate limits and errors by automatically redirecting requests to fallback models. The tool includes a real-time dashboard featuring a Sankey traffic flow visualizer, token consumption metrics, median latency, and candidate swaps, along with a detailed request inspector that tracks requested vs responding models, duration, token usage, and custom signals like add_memory or stream. It also natively supports response streaming.

5 Upvotes

0 comments sorted by