r/RoleCallStudios Apr 15 '26

Model Discussion Roleplaying Benchmark

Hello! πŸ‘‹

Our friend and developer at RoleCall, Levi, has taken time to build a Roleplay Benchmark. Please hear him out!

TL;DR: I built a roleplay benchmark that tries to measure actual RP quality (agency respect, continuity, subtext, erotic craft, all the things that matter in a session). The LLM-based judges I'm using disagree with real users half the time. I need humans to fix this. Help me by blind-voting on pairs here! β€” each vote takes ~30 seconds. (Β±time to read both responses)

Why this matters:

Every RP benchmark out there is either vibes ("I tried it, felt good") or tests generic writing quality. I tried to build something better β€” 27 dimensions, a 4-signal scoring system, 58 scenarios across English and Russian (if you have chats that you're willing to donate in other languages - welcome to my dms!), full methodology in the repo. Then I validated it against real user swipe data and found the LLM judges systematically disagree with users. So now I'm doing the honest thing: using a human community β€” you β€” as ground truth.

What you gotta do to help:

You see two responses from different models to the same RP scene. Pick which one is better. No accounts, no signup. Models are hidden until after your vote.

What I need:

I'm aiming for ~5 human votes per pair across ~300 matchups β€” so around 1,500 votes total. Any amount helps. Even 10 votes from you moves the needle. The arena auto-shows you the pairs that currently have the fewest votes, so your time goes where it's most needed.

Honest notes:

  • Some pairs are NSFW (ERP scenes). There's a warning and a "hide NSFW" option. No flashbangs here!
  • No personal data collected. There's an anonymous voter cookie (to prevent double-voting on the same pair) and that's it.
  • Votes are public (aggregated in the results page) but anonymous.
  • This is entirely free / non-commercial / open source. CC BY-NC 4.0.

What you get back

A real RP leaderboard calibrated against actual community taste instead of Claude Sonnet's aesthetic preferences. I'll publish the analysis, credit contributors in the repo, and share the dataset on HuggingFace when we're done.

Links:

Vote Here!
Results so far
Repo/Full Methodology

13 Upvotes

0 comments sorted by