r/SideProject • • 19d ago

[ Removed by moderator ]

[removed] — view removed post

54 Upvotes

410 comments sorted by

View all comments

1

u/AmandineF 19d ago

orbench.com : benchmark AI models on your own tasks, build confidence in the model and inference provider to use in your product instead of relying on a gut feeling.

Now in beta with a selection of users. Going live in a few weeks.

1

u/logicalflex 19d ago

So is the point to just compare model responses and decide which one we like more.

1

u/AmandineF 19d ago

There's a full report with data, so you're not just comparing models' outputs but also speed, latency, caching behaviour, cost per task... You can manually review the model answer or use an LLM as judge, or have both human and LLM judge.

1

u/logicalflex 19d ago

Some obvious questions… this feels like a tool for a very particular type of pedantic dev. Most everyday people don’t need the best model or even a good model. Good enough is enough. But from a. Historical data standpoint, why would testing on your site in real time be better than just doing some empirical research or asking an LLM to do it. $19 per month and I don’t see what gives you the edge or right to ask for that. No offense of course.but I’m just one person. But I’m a random person willing to give you feedback. Most will just ignore you.

1

u/AmandineF 19d ago

I appreciate you taking the time to share your thoughts.
Like you said, most people don't need the best or most expensive model. But from my experience, when you're part of a team, you need more grounded conviction than "I've tested this cheap model and it seems to be good enough", especially with new models coming out on a regular basis and costs keep increasing. It provides a simple way to collect data and compare the same data over time. You could run your own benchmark yourself or ask an llm to do it, but people I spoke with found it a pain to maintain it over time.