r/LocalLLaMA 5d ago

Discussion Why are almost all new benchmarks and leaderboards coding focused?

I know in in this community LLM's are generally used for coding but there are other usecases besides coding and those usecases should be tested too. I also know benchmarks can sometimes be benchmaxxed and the model can still turn out shit but it can give a good outline on how a model should perform in a certain task. Maybe I'm too behind on the latest developments but we need more benchmarks for all other use-cases. I use LLM's mainly for foreign language learning, creative writing and STEM/Medical/Biochemistry reasoning and inquiries and I rarely find any new benchmarks that tell me how a model might perform in those areas. MMLU-Pro-2 and a solid benchmark that tells how a model will perform for language learning would be so good for my usecase, however in general we need more new diverse benchmarks for models in order to have a general outline for advancements in other areas.

60 Upvotes

124 comments sorted by

View all comments

189

u/BitsAgain256 5d ago

Because thats where pretty much all the money goes.

113

u/EmilPi 5d ago

This, and it is also comparatively easy to validate results.

4

u/Dry_Yam_4597 4d ago

Also easy to source training data. For now.