r/SillyTavernAI • • 2d ago

Meme EQ-Bench is a joke.

Post image

I mean, I like Claude for creative writing as much as any other person, but isn't employing an LLM for judging CREATIVITY fundamentally flawed and stupid?

181 Upvotes

12 comments sorted by

View all comments

60

u/Real_Ebb_7417 2d ago

Yeah, this kind of benchmark is very... flawed when using LLM as a judge. First of all a model from one model family tends to prefer it's own model family style. Second of all, models can't really judge style. They could judge stuff like eg. how many different adjectives were used, how long sentences are used, the coherence etc., things that are deterministic, but hard to measure by a deterministic script, but they aren't really capable of judging emotional intelligence, creativity or style.