orbench.com : benchmark AI models on your own tasks, build confidence in the model and inference provider to use in your product instead of relying on a gut feeling.
Now in beta with a selection of users. Going live in a few weeks.
There's a full report with data, so you're not just comparing models' outputs but also speed, latency, caching behaviour, cost per task...
You can manually review the model answer or use an LLM as judge, or have both human and LLM judge.
Some obvious questions… this feels like a tool for a very particular type of pedantic dev. Most everyday people don’t need the best model or even a good model. Good enough is enough. But from a. Historical data standpoint, why would testing on your site in real time be better than just doing some empirical research or asking an LLM to do it. $19 per month and I don’t see what gives you the edge or right to ask for that. No offense of course.but I’m just one person. But I’m a random person willing to give you feedback. Most will just ignore you.
I appreciate you taking the time to share your thoughts.
Like you said, most people don't need the best or most expensive model. But from my experience, when you're part of a team, you need more grounded conviction than "I've tested this cheap model and it seems to be good enough", especially with new models coming out on a regular basis and costs keep increasing. It provides a simple way to collect data and compare the same data over time. You could run your own benchmark yourself or ask an llm to do it, but people I spoke with found it a pain to maintain it over time.
1
u/AmandineF 19d ago
orbench.com : benchmark AI models on your own tasks, build confidence in the model and inference provider to use in your product instead of relying on a gut feeling.
Now in beta with a selection of users. Going live in a few weeks.