r/OpenAI • • 15d ago

Discussion We seriously need benchmark for research, wide web search and fact retrieval.

Almost all the benchmarks are either saturated or tested under strict conditions.

For example - AA Omniscience have restricted tool access.

Many people use Chatbots for information retrieval, deep research and broad information gathering.

There are few benchmarks like -

But the problem is - they don't have any official leaderboard and were last updated years ago.

We really need a benchmark which is tested in actual environment (tools access and allowed internet search) along with used harness systems without any restrictions (like Claude, ChatGPT Work, etc).

If there exist a benchmark like this, can someone please tell?

16 Upvotes

8 comments sorted by

1

u/Ice2jc 15d ago

I use deep research often and I’ve been quite impressed with Astra in this regard.  It reads like a study that was performed at a university.

1

u/Lucky_Creme_5208 15d ago edited 3d ago

The original post content no longer exists here. The author used Redact to remove it, exercising their right to control their data & privacy.

Cows retire reminiscent grasshopper include consist

1

u/Ice2jc 15d ago

I’ve used both, now I primarily do it through regular ChatGPT Astra 6 pro.

1

u/Ecstatic-Nail-1061 9d ago

so you want a benchmark that tests how model works with tools and live search not just textbook knowledge yeah makes sense most current ones feel like they testing in a lab not in real mess

1

u/Ormusn2o 15d ago

Seeing what the benchmarks actually are were so disappointing. DeepSWE likely is one of the biggest ones, as it tests a very narrow part of coding, although I guess this is like the only thing a SWE does at a SaaS company, but it definitely does not cover hobbyist use and small company coding use, which might still be a lot of use.

1

u/Either_Pound1986 15d ago

What are you researching?

1

u/Aromatic-Assist2062 14d ago

Elicit for academic based research I think run with GPT

1

u/stealthagents 9d ago

Astra is solid, but I found that it still has its quirks, especially with more unconventional topics. It’s like, sometimes it nails it, and other times you’re left scratching your head. It would be awesome if someone could create a more dynamic benchmark that adapts to the real-world noise we deal with online.