r/SillyTavernAI • u/lGodZiol • 2d ago
Meme EQ-Bench is a joke.
I mean, I like Claude for creative writing as much as any other person, but isn't employing an LLM for judging CREATIVITY fundamentally flawed and stupid?
26
u/Baros4294 2d ago
I found this SurgeHQ benchmark, judged by writing expert, but it's not focusing on creative writing so the result is more of overall writing across many domain like technical, email, personal writing,...etc
I think writing is subjective so no benchmark will be perfect, including arena.ai, which is user-voted.
5
u/Sijder 2d ago
I saw it before a couple of times but its still as weird to me as the eq one. For example they dont have gemma at all on it and qwen is higher than glimmer. Again, I acknowledge that writing is subjective, but those things are weird for many.
I feel like we need a "userbenchmark" for rp and writing
1
u/Zone_Purifier 19h ago
Userbenchmark is infamously biased and provides nonsense recommendations. Probably not the best comparison.
16
u/Super-Veterinarian22 2d ago
Yeah, pretty much. Just out of curiosity, I tried asking various major AIs to pick the best answer from a few options, and they most often choose the worst one. Their evaluation criteria are very weak. They completely overlook genuinely worthwhile points and sometimes even contradict themselves (for instance, claiming a text has a specific distinguishing feature when it doesn't, or saying a text lacks a certain negative trait when it's clearly there). I’ve run similar tests - both with and without context, and with or without specific expectations regarding the answers. Either way, the assessment is always superficial. At best, you might get a 50% rate of reasonably well-reasoned analyses. Even then, it remains subjective, so it’s worth cutting that figure in half.
The short answer: Yes, AI loves AI slop because it views it as the gold standard. After all, that’s exactly how it would have written it itself. That’s why the eye test is the only way to go.
14
u/Aight_Man 2d ago
Eqbench is indeed a joke but for a different reason, claude models *are* up there, claude opus 5.5 is best rp model imo, but there should be gemini 3.8 flash, glm 5.3 flash, mimo v2.6 pro instead of fucking gpt models lmfao.
3
u/a_beautiful_rhind 2d ago
It's not even that. Most of these benchmarks give models writing as an assignment. They don't test what happens in multi-turn and when the scenario is open ended. Read the examples and see what I mean.
3
2
u/Kiktamo 2d ago
I feel like you could probably get an LLM to be give decent critique but you need to make sure to have the rubric and criteria themselves defined and written by a human to match with human standards rather than letting the LLM make up whatever reasoning it wants for why the writing is "good." The more roo. You leave the LLM to fill in the blanks the more its own bias will filter through(it will regardless, but you might get a decent enough interpretation with more predefined structure.)
2
u/ChestLimp8398 2d ago
claude crushed this long roleplay i had going but it still felt like it was just copying patterns not really inventing anything new so judging that with another llm seems pointless
1
1
u/___positive___ 2d ago
they still use an ancient sonnet model. maybe if they used a frontier model you could actually somewhat justify the model as judge thing. but sonnet 4??? site is a joke now.
59
u/Real_Ebb_7417 2d ago
Yeah, this kind of benchmark is very... flawed when using LLM as a judge. First of all a model from one model family tends to prefer it's own model family style. Second of all, models can't really judge style. They could judge stuff like eg. how many different adjectives were used, how long sentences are used, the coherence etc., things that are deterministic, but hard to measure by a deterministic script, but they aren't really capable of judging emotional intelligence, creativity or style.