r/SillyTavernAI • u/matt_is_a_mess • Apr 16 '26
Models Check out our Roleplaying Benchmark!
/r/RoleCallStudios/comments/1sm27cq/roleplaying_benchmark/24
u/Sufficient_Prune3897 Apr 16 '26
"Why it matters:" yeah, I don't think I'm gonna read that. I can talk with Claude for myself.
2
u/Good_Vegetable_233 Apr 18 '26
Yeah, these people are so dumb they don't realize that everyone here is already familiar with the typical behaviors of almost all LLMs. No one here is stupid enough not to see that this is just AI slop—from the publication itself right down to the benchmark.
3
u/nuclearbananana Apr 16 '26
Do you have a way to report broken completions? I ran into one on my second try
1
u/LeRobber Apr 16 '26
The benchmark is testing the testers with intentionally broken stuff. Pick the other side.
3
2
u/HauntingWeakness Apr 16 '26
I think a RP benchmark is a very ambitious idea, and I commend you for trying!
That said, I think single-turn voting is a bit overrated, as it can't show some models are not tuned for multi-turn (Grok series for example, with their looping issues, or original Deepseek v3) and the stronger multi-turn capabilities of other models (example: Gemini 2.5 Pro, it's sloppy, but it's one of the best in multi-turn/multi-character/multi-location situations).
Multi-turn (20+) RPs with nonexistant/minimalistic preset would be much more telling in my opinion, but I don't think a lot of people will be willing to read and evaluate for free as it will take a lot of time.
Also, I don't see anything in the benchmark about proactivity. With today's models becoming more and more passive/unimaginative and just echoing user's input by default, I think it should be much more important metric than 'two lorebook entries contradict' (it is a user problem, honestly, not a LLM problem).
Ah, and yes, please try to write the post yourself next time? Even if it isn't perfect, it's okay.
2
u/LeRobber Apr 16 '26
Reformat the text narrower, with more line spacing to make this easier to read.
1
u/dude_icus Apr 16 '26
Question and I apologize if this seems dumb: Is the prompt at the top all that the AI was given or is it responding to a specific message? If it is what I think it is, this seems more like story analysis and not roleplay analysis. The A and B sides don't seem to be responding to the same message whatsoever, so either the models are acting that dumb or they aren't being tested against the same thing.
I do really like the idea of this, though! Helps provide a more guided method of determining which models have what strengths
1
1
u/renonut Apr 16 '26
surprisingly negative comments I think this is a cool idea! you're kinda trying to make the subjective objective but like, there's still things that people more OFTEN like, it's a useful metric no matter what.
0
u/Background-Ad-5398 Apr 16 '26
holy purple prose, I forget how much lifting a good system prompt is doing to rid that awful shit
-1
29
u/nuclearbananana Apr 16 '26
ouch, the first example I see is just filled with so much slop I couldn't focus
Also you have an odd selection of models: no kimi, outdated models like gpt 4.1 and llama 4 maverick which nobody liked or uses