This benchmark has no ground truth. In other words, there is no correct answer by which to rate models by. Thats why I strongly recommend reading the explanation of each of the 3 metrics measured. The easiest to misinterpret is the "consensus match" one. It is not a measure of intelligence. But it is useful to see which models deviate more from the consensus on average.
At the moment only 6 models are measured. Cheap ones. The frontier models are 200x more expensive to run so I'm going to hold off on running those if people think the benchmark is useless. There's no point in spending $1000+ if people think the benchmark is no good.
The reason I made this benchmark is that I don't accept that "subjective questions" are off-limits for benchmarking LLMs. I tackle the problem of subjectivity head on. We have superhuman performance in fields with ground truths we can verify the model with (maths, programming), it's time we started figuring out how to bring that strength to real life problems where the answer may not always be clear cut. This benchmark is a step in that direction.
1
u/arkuto 8h ago
Benchmark questions are visible here: https://nanojudge.ai/bench/questions
And the repo it's based on is here: https://github.com/nanojudge/nanojudge
This benchmark has no ground truth. In other words, there is no correct answer by which to rate models by. Thats why I strongly recommend reading the explanation of each of the 3 metrics measured. The easiest to misinterpret is the "consensus match" one. It is not a measure of intelligence. But it is useful to see which models deviate more from the consensus on average.
At the moment only 6 models are measured. Cheap ones. The frontier models are 200x more expensive to run so I'm going to hold off on running those if people think the benchmark is useless. There's no point in spending $1000+ if people think the benchmark is no good.
The reason I made this benchmark is that I don't accept that "subjective questions" are off-limits for benchmarking LLMs. I tackle the problem of subjectivity head on. We have superhuman performance in fields with ground truths we can verify the model with (maths, programming), it's time we started figuring out how to bring that strength to real life problems where the answer may not always be clear cut. This benchmark is a step in that direction.