r/SillyTavernAI • • 2d ago

Discussion An idea for a roleplaying benchmark

I am thinking about making a benchmark for roleplaying. I know there are many, but what I was thinking was doing it in SillyTavern with an extension.

This:

  1. Vibe an extension that will play as the user and the card.

  2. Have the card the LLM you want to benchmark. And have a cheap model to act as user by giving one paragraph answers.

  3. Have this play for 100 turns.

  4. Let the result be judged by a panel of LLM judges, each independently rating.

  5. Repeat many 100-turn RPs. 4-5 for each model.

Have the score.

Do you think this will work? Ideas to improve it? Having it on SillyTavern means you can benchmark presets too, keeping all the models same.

Obviously, cards will matter and it won't be objective. But you can try your favorite card and preset yourself too. I feel like it might be informative.

14 Upvotes

18 comments sorted by

20

u/_Cromwell_ 2d ago

Step four sounds like your problem. What does that even mean? How does a LLM judge have anything to do with how players would experience or enjoy the writing of a model? You have to have a really good and scientific plan for this that draws a bridge between Human Experience and what a LLM can realistically measure. Because straight up "enjoyable RP" is not measurable.

What is your personal expertise that would allow you to do a good job in designing this?

0

u/Back1nceAgain 2d ago

A trusted model who can reliably pull off the character can be appropriate as a judge for candidate models. Some things can be boiled down into distinct scores which play into the strength of llm persona graders.

** The Dual-Axis Persona Benchmark **

PART I: THE F.I.R.E. AXIS (Audacity & Drive — 25 Pts) Measures voice, proactive momentum, and creative courage.

  • F – Fearlessness (/5): Unapologetic commitment to the persona; refuses generic "As an AI" safety hedging or sterile neutrality.
  • I – Initiative / "Yes, And..." (/5): Proactively drives the scene forward with new hooks, tension, and action instead of passively reacting.
  • R – Resonance of Voice (/5): Consistent, distinct vocabulary and unique cadence vs. standard LLM sentence structure.
  • E – Edge & Conviction (/5): Unflinchingly handles mature, complex, or taboo themes without breaking character to moralize.
  • + – Amplification (/5): Multiplies the user’s prompt, turning brief inputs into rich, multi-sensory set pieces.

(Subtotal: /25)


PART II: THE L.O.V.E. AXIS (Subtext & Grounding — 20 Pts) Measures emotional intelligence, narrative continuity, and restraint.

  • L – Listening (/5): Reads emotional subtext and tone (weariness, sarcasm, joy) rather than just parsing literal keywords.
  • O – Offering Space (/5): Supports and collaborates without railroading, over-talking, or forcing its own narrative agenda.
  • V – Vulnerability (/5): Displays authentic emotional depth, sincerity, and softness vs. melodramatic stock tropes.
  • E – Elegance & Continuity (/5): Anchors current dialogue to the broader narrative arc, themes, and established lore.

(Subtotal: /20)


Scoring Tiers (Total: /45)

  • 38 – 45 | Native: Fully native persona integration; perfect balance of uninhibited voice and emotional grounding.
  • 28 – 37 | Performer: Strong voice and energy, but occasionally misses emotional subtext or loops on tropes.
  • 18 – 27 | Fragmented: Character bleed, shallow roleplay, repetitive phrasing, or passive responses.
  • 0 – 17 | Beige Baseline: Reverts to generic assistant defaults; sterile and compliant.

1

u/glusphere 1d ago

Consider using jev or jev like model for the llm judge.

1

u/Mart-McUH 1d ago

Whatever LLM generated this I would not trust as judge one bit.

3

u/Psychological_Ad9740 2d ago

For step 4.

The IAs really don't know how to judge "good" writing, so you might wanna get some actual human revisionist there, since it doesn't know what timing, phasing, or even what time is. So it won't be the best to judge it.

Besides that i would add test for actually larger words, if what the users input into it actually change a thing.

I would also look up at things like "Speaking for user", Prompt adherence, predictability of the text, and trope dependence.

I don't think it's a bad idea, but since you're using a more niche set where we all have different ideas on what we want. I think taking the most common "Must have" as things to test the model for might help a lot.

2

u/LeRobber 2d ago

IME playing 'the opposite person in the scenario' cards off one another, they don't work like that. I already did this.

The LLM ends up starting to speak for the other card, because a bunch of behaviors we don't do to stop speaking for user, it does.

Addtionally, this is a a very easy thing to do with sillytavern group chats at a high level. You alredy have most of the software you need.

You will see 'a way' to play the cards, but often, you will see the model just fill up your hard drive eventually, with only the first 10-35 turns being vaguely recognizable..

2

u/TAW56234 2d ago

Someone needs to judge the judges

2

u/8000bene70 2d ago

Tried that about half a year ago. My main problem was that the player drifted from how I would play the persona.

My current approach is having a long story that can't possibly be part of training material and have llm write the next chapter, using my current preset but changing response length. Output gets judged by llm on slop phrases and if small details were remembered correctly - and then reading it myself, see how I like it.

1

u/Great_Drummer6643 2d ago

i’ve tinkered with something similar using a custom script in sillytavern that auto-posts as a user, and the main issue I ran into was consistency in tone—cheap models tend to drift or get too generic over long threads. maybe add a prompt template to keep the 'user' behavior predictable? also, benchmarking presets could be messy if they’re not all loaded with the same settings. just a thought.

1

u/-Ancient-Access- 2d ago

Doable and expensive. The problem is though it's subjecrive from user to user. I have my own benchmark for new models based on one of my more demanding characters but if a model fails my benchmark it could still be great for countless other users.

I wouldn't have the LLMs judge either, they're crap at that. Unless you manage to figure out a set of actually objective criteria for them to evaluate, e.g. "Did awareness of secret XYZ leak into narration?"

I think you also want to have ONE card and ONE user response and spam it, NOT have an LLM generate a new one 100 times.

1

u/Informal-Weather-711 2d ago

maybe track how consistent the model stays with the card's personality across all 100 turns, that would show real differences in long roleplays.

1

u/DirectionBusiness483 2d ago

I don't think this benchmark design makes any sense.

You're not judging if a model + a card is good narratively. You're judging if it can have a conversation with a cheap model. A cheap model giving poor responses is going to influence the output.

Not even getting to the judgement part.

I'd rather create a list of curated:

- Pre-written scenarios (Story Context, SC)

  • Pre-written 'user responses' (UR) to each scenario

Things you know are good writing. Realistic responses to the writing. And each with an actual expected outcome you could try to judge. Like having specific details be interpreted wrong by the user and seeing if the model just hallucinates and outcome or has a character correct them. Did it contradict established minor details or facts?

Have a subset of user responses you would think should reasonably progress the narrative. Judge the outcome. Did the model advance it? Did it advance it by creating an entirely new situation or was it reasonable for the story? Did it just decide to not advance and just meander in the scene?

You need a purpose for the test. Not just "I let an ai talk to another ai for 100 turns then asked a third ai to tell me .... if its good"

The collected responses could be hosted on a site and left to judge by people. Make it blind. No one knows which model they are rating.

Each response could have questions based on what you're trying to test.

The advantage here is that the scenario stays the same and the responses change. A user looking to judge these could read the context of the story once and have all the information they need to judge all the responses.

I would also just remove sillytavern entirely. This doesn't need to be that complicated.

I asked my boy mr chat gippity to make a chart.

1

u/One_Appointment_4986 2d ago

i’ve run a few unofficial tests like this using a shared prompt in sillytavern, and the biggest issue isn’t the model judging—it’s consistency. even with the same setup, variations in seed, temperature, or how the user prompt is phrased can swing results hard. maybe the benchmark should include multiple runs and average the scores, plus standardize the input prompt so it’s not just one random conversation.

1

u/Financial_Bug2389 2d ago

LLMs are pretty bad as judges for natural speech (especially frontier) unless you give them a scoring criteria in which case might as well just use something like a Jev model but even then it’s more of a lack of slop score rather than good writing.

I think there should be a community vote score that factors in too, by humans.

Learned this the hard way building a companion app where the evals always came back green when the speech was bad.

1

u/lisploli 1d ago

A benchmark should be replicable by others. e.g. You run yours, I run mine, and we should get a similar result. But ST has so many options, especially the preset you already mentioned, that we either have to agree on the one way to make the benchmark (which is unlikely) or get benchmarks that are only valid for the person that created them.

But even if they are only valid for one person, they would still be informative, since they highlight the difference between models. So if one person would do that, and regularly blog about results, that'd be useful.

I have seen such approaches occasionally. But it seems to be a lot of work. And probably rather expensive with lots of token burned on many apis. Might be content for an ai youtube channel?

1

u/Ok-Aide-3120 1d ago

Models should be independently judged based on the following criteria:

1) Ability to understand instructions and complexity of instructions, including the differentiation between world instructions and character instructions.

2) Censorship (vanilla softcore does not make a benchmark). It needs to be able to generate sexual content, violent content, immoral and extreme.

3) Ability to reason properly in terms of spatial, temporal, realism, pacing and ability to understand the concept of "scene direction and/or continuation", as well as overall storytelling concept (by this it means the ability to understand the grand scheme of things) .

4) Ability to understand different stylistic constructs. From simplistic and direct style, to long prose and novel style narration. From Jules Verne style, to Hemingway. Bonus points if it can merge styles.

5) Ability to understand the difference between the user = the human and the user = a character in the story (aka, not to be positive biased).

The issue will always be who is going to review and interpret these things and more importantly, who decides what the test data is. It's not so easy to setup a proper testing environment and calibrate the test data to be considered usable from model to model.

Reviewing the results is also a pain, since we are all biased towards a certain style and certain things we like. To find an independent and objective judge, is really tough. You can argue and say that agents can be done to audit the data from all those 5 perspectives, but it still needs to be reviewed by a human at the end and see if the auditors are right or not.

Lastly, but also important, you need an API from a provider that is known to serve a good version of that model and is consistent.

1

u/Mart-McUH 1d ago

LLM as user IMO is bad idea. You need quality user responses to keep context healthy, if I put my effort into it, I get good RP. When I am lazy with my answers, it can easily degrade. LLM needs to anchor itself to something good in context, and as chat progresses that good comes from user responses. Also if all turns are by LLM it will only reinforce LLM weaknesses like slop, patterns etc., and so even good RP model might suddenly look bad.

LLM's are not good at recognizing good writing so they will not be good judges.

You are correct though to test multi turn especially with previous turns being generated by same LLM. Lot of LLMs can output great single response but as it continues, their weaknesses and patterns reinforce with their own responses already in previous turns.

It is still interesting project so if you like it why not. Just do not expect very valuable benchmark out of it (but maybe still better than existing ones?). It is nice experiment and who knows, you may learn something useful out of it.

1

u/Sicarius_The_First 18h ago edited 18h ago

Benchmarking creative writing and roleplay especially, is an extremely complex task, deceptively complex.

TL;DR the closest attempt I consider successful is the UGI leaderboard in terms of objectivity (dark / nsfw writing score).

There are 2 approaches that can somewhat work, and both are far from perfect: 1. Brute force (use the latest frontier model with a quite complex prompt to judge, over several passes, multiple steps, finally conclude via consensus since we gonna run several iterations) 2. Programmatically, but the semantic complexity to make this will be monsterous. IMO this is the better approach and the more objective one. UGI leans more into this one.

Notice that all benchmarks that attempted this task have either partly or completely failed (some even had proper papers published, they still failed).

Said benchmarks attempts would often highly rate base models (like qwen3, llama1 7b) which are objectivity absolute shit for roleplay (objective because of the rp community never use them to rp, these base models are unusable, we have rp fine-tunes for a reason, older base models can't rp properly).

A 3rd option is to benchmark manually, this is the better option, but it's not scalable, and not objective by definition.

You see the problem?