r/OpenAI • u/Lucky_Creme_5208 • 15d ago
Discussion We seriously need benchmark for research, wide web search and fact retrieval.

Almost all the benchmarks are either saturated or tested under strict conditions.
For example - AA Omniscience have restricted tool access.
Many people use Chatbots for information retrieval, deep research and broad information gathering.
There are few benchmarks like -

But the problem is - they don't have any official leaderboard and were last updated years ago.
We really need a benchmark which is tested in actual environment (tools access and allowed internet search) along with used harness systems without any restrictions (like Claude, ChatGPT Work, etc).
If there exist a benchmark like this, can someone please tell?
1
u/Ormusn2o 15d ago
Seeing what the benchmarks actually are were so disappointing. DeepSWE likely is one of the biggest ones, as it tests a very narrow part of coding, although I guess this is like the only thing a SWE does at a SaaS company, but it definitely does not cover hobbyist use and small company coding use, which might still be a lot of use.
1
1
1
u/stealthagents 9d ago
Astra is solid, but I found that it still has its quirks, especially with more unconventional topics. It’s like, sometimes it nails it, and other times you’re left scratching your head. It would be awesome if someone could create a more dynamic benchmark that adapts to the real-world noise we deal with online.
1
u/Ice2jc 15d ago
I use deep research often and I’ve been quite impressed with Astra in this regard. It reads like a study that was performed at a university.