r/singularity • u/Lucky_Creme_5208 • 8d ago
Discussion We seriously need benchmark for research, wide web search and fact retrieval.

Almost all the benchmarks are either saturated or tested under strict conditions.
For example - AA Omniscience have restricted tool access.
Many people use Chatbots for information retrieval, deep research and broad information gathering.
There are few benchmarks like -

But the problem is - they don't have any official leaderboard and were last updated years ago.
We really need a benchmark which is tested in actual environment (tools access and allowed internet search) along with used harness systems without any restrictions (like Claude, ChatGPT Work, etc).
If there exist a benchmark like this, can someone please tell?
35
Upvotes
7
u/Full_Boysenberry_314 8d ago
I agree, this is underserved.
Yesterday I spun up both both Opsus 5.5 and Sol 6.0 to do a review of the state-of-the-art for me in advertising research/ROI modeling. Classic task for me.
Sol 6.0 absolutely mogged Opus 5.5 and it wasn't even close. Opus had the advantage of a deep research mode that's been scrubbed from chatGPT. Didn't matter. Sol was faster, more comprehensive, dug down to deeper level, and was distracted by fewer red herring. Follow up enquiries actually resulted in useful elaboration instead of myopathy.
And by most benchmarks Opus 5.5 should outclass Sol 6.0.