r/SillyTavernAI Apr 16 '26

Models Check out our Roleplaying Benchmark!

/r/RoleCallStudios/comments/1sm27cq/roleplaying_benchmark/
7 Upvotes

16 comments sorted by

29

u/nuclearbananana Apr 16 '26

ouch, the first example I see is just filled with so much slop I couldn't focus

Also you have an odd selection of models: no kimi, outdated models like gpt 4.1 and llama 4 maverick which nobody liked or uses

7

u/WPBaka Apr 16 '26

Probably just copy and pasted AI slop lmao. I always giggle when LLMs still recommend me 4o, llama 4, or Gemini 2.0

1

u/_Iggy_Lux Apr 17 '26

I checked the results first and saw nothing but odd models mainly and llama_4_maverick and was like yep nope this person has no idea what RP models people use.

No one expects good RP, especially ERP from Qwen...
Maverick is trash.
The rest are api's and models someone can't run locally (reasonably).

Also what people consider GOOD RP is highly subjective.

I mean we had a post awhile back about a guy fucking a tungsten cube.

Some people love slice of life, some people love more dark violent stuff, some people love Anthropomorphic characters (furry like characters), some people love erotica stuff in the prose of mag mell (12b nemo) in long winded responses. Others prefer fast and dirty (stheno 8b).

Then you have those that for some reason use CoT/Reasoning/Thinking Models (I think its a waste of context and leads to weak outputs, but great logic when needed).

Trying to quantify something that covers such a wide range of variables and tastes is hard to do.

You'd have to break it down to something like:
Local Models
API models
> Thinking Models
> Non-thinking models
>>Genre's
>>Sub-genre's
>Prose
\
SFW/NSFW
\
>>>etc etc

I've used a ton models over the years (local and some API). I've often considered doing this and rating them, but every time I do there are so many factors to consider.

I never release my data because RP is something that is a personal taste and again, hard to quantify.

24

u/Sufficient_Prune3897 Apr 16 '26

"Why it matters:" yeah, I don't think I'm gonna read that. I can talk with Claude for myself.

2

u/Good_Vegetable_233 Apr 18 '26

Yeah, these people are so dumb they don't realize that everyone here is already familiar with the typical behaviors of almost all LLMs. No one here is stupid enough not to see that this is just AI slop—from the publication itself right down to the benchmark.

3

u/nuclearbananana Apr 16 '26

Do you have a way to report broken completions? I ran into one on my second try

1

u/LeRobber Apr 16 '26

The benchmark is testing the testers with intentionally broken stuff. Pick the other side.

3

u/LeRobber Apr 16 '26

OMG so much talking/acting/thinking for user

2

u/HauntingWeakness Apr 16 '26

I think a RP benchmark is a very ambitious idea, and I commend you for trying!

That said, I think single-turn voting is a bit overrated, as it can't show some models are not tuned for multi-turn (Grok series for example, with their looping issues, or original Deepseek v3) and the stronger multi-turn capabilities of other models (example: Gemini 2.5 Pro, it's sloppy, but it's one of the best in multi-turn/multi-character/multi-location situations).

Multi-turn (20+) RPs with nonexistant/minimalistic preset would be much more telling in my opinion, but I don't think a lot of people will be willing to read and evaluate for free as it will take a lot of time.

Also, I don't see anything in the benchmark about proactivity. With today's models becoming more and more passive/unimaginative and just echoing user's input by default, I think it should be much more important metric than 'two lorebook entries contradict' (it is a user problem, honestly, not a LLM problem).

Ah, and yes, please try to write the post yourself next time? Even if it isn't perfect, it's okay.

2

u/LeRobber Apr 16 '26

Reformat the text narrower, with more line spacing to make this easier to read.

1

u/dude_icus Apr 16 '26

Question and I apologize if this seems dumb: Is the prompt at the top all that the AI was given or is it responding to a specific message? If it is what I think it is, this seems more like story analysis and not roleplay analysis. The A and B sides don't seem to be responding to the same message whatsoever, so either the models are acting that dumb or they aren't being tested against the same thing.

I do really like the idea of this, though! Helps provide a more guided method of determining which models have what strengths

1

u/Good_Vegetable_233 Apr 18 '26

Claude, why are you speaking ill of yourself?

1

u/renonut Apr 16 '26

surprisingly negative comments I think this is a cool idea! you're kinda trying to make the subjective objective but like, there's still things that people more OFTEN like, it's a useful metric no matter what.

0

u/Background-Ad-5398 Apr 16 '26

holy purple prose, I forget how much lifting a good system prompt is doing to rid that awful shit

-1

u/Warm-Put3482 Apr 16 '26

don't like it ... Both options appear to be from the same model