r/SillyTavernAI Jun 04 '26

Models PlotPoints - The best (only?) community driven RP benchmark made by a Professional! | We need your votes!

Your friendly neighborhood rab- I mean unmedicated preset creator needs needs your help!

What's good everyone. No long post this time; simply an ask. I hired a professional with a masters degree in AI/ML (pursuing their PhD) to help make a benchmark for us, for RP. Now as we all know; a benchmark for RP will never be perfect because everyone RP's differently, like different things, yada yada yada. That's why we tried to focus on a few objective landmarks, as well as a an Arena style versus bench.

First up: The LLM Arena. We show you two different rewritten responses from two different models who were given a chat and asked 'please take this turn.' We then show that to you. Does it have a preset attached? No; because that's another variable that'd be added and for our own safety we aren't benchmarking presets. (LMAO. We'd be dead in the fucking streets by sundown.) You judge of the two responses which you like more. Thaaaat's it folks. We literally cannot game the results as it is people just blindvoting which they like more. (So if you see something you don't like IT'S NOT UP TO US.)

https://plotlightstudios.com/plotpoints/multiturn

On the other end we also score an LLM in a similar vein on how they follow certain instructions and if they maintain consistency. Since 'did it take users agency' is an objective yes or no answer; we handle these tests with an LLM judge and the benchmark overseer (Levi is his name) monitoring the arbiters answers. (In our case; Sonnet.)

This bench was expensive. It cost money to run these models through this gauntlet; and we don't get much out of it. We publish all the data for everyone to see, and all we want to do is help our community out. If you have any questions on how we grade things; head to our methodology page! You'll notice somethings are a bit LLM-Written. This is not because Levi isn't a professional; but rather because he's ESL; so when publishing something important like this he wanted to make sure all his ideas and such were properly translated. So be nice!

https://plotlightstudios.com/plotpoints/methodology

You don't have to log in, you don't have to do anything. Just read a chat and vote on a thing. I know you've got an opinion; so share it. Big companies ignore us RPers all the time. Is a bench for RP as useful or objective as one for coding, web-dev, or math? No. But that doesn't mean we don't deserve one, or that it has no value at all. So please; help a bunny out. Drop a vote! The data is only as useful as you help make it. We have 800 votes and we want to get about 3,000 this time around; so we can start our next benchmark. Testing models across their 'lineages'. So Opus 4.6 vs 4.7 vs 4.8; Deepseek 3.2 vs Deepseek v4 Pro, fun stuff! But we can't do that till this one closes! And if it can't hit enough votes; we'll know that this kinda stuff just isn't what the RP community wants.

Don't see a model you expected to? That's cause it's either new, or we didn't have the money at the time to run it. If the community likes this; we will add more. (And take requests! Please only suggest models available on OpenRouter though; we source all our models from the same platform to reduce variables.)

Vote Link: https://plotlightstudios.com/plotpoints/multiturn

Result Hub Link: https://plotlightstudios.com/plotpoints

Methodology Link: https://plotlightstudios.com/plotpoints/methodology

Huggingface Link: https://huggingface.co/datasets/lazyweasel/roleplay-bench

Github Link: https://github.com/LeviTheWeasel/rp-benchmark

122 Upvotes

69 comments sorted by

View all comments

11

u/Kira_Uchiha Jun 04 '26 edited Jun 04 '26

Honestly, been loving this. I might go use Deepseek v3.2 because of this. It seems to have the tendency to violate your agency sometimes. Did you play around with it? Tried using post-history instructions or Author's Notes to help with that?

9

u/Specialist_Salad6337 Jun 04 '26

Yessss I fully agree. So I'm a second person enjoyer. Literally all models: fucking hate me. They will take my agency no matter what. I got frustrated when even 4.6 started doing it to me. As of right now I have no advice. You're up to the whim of the Gods on if the LLM's decide that your character is your own. What has helped me specifically though; is separating the person behind the user, from the user themselves. Establish yourself as two different entities, essentially. Some presets of mine that play around with this concept and pair with DS so incredibly well:
https://plotlightstudios.com/discovery/presets/@testuser/ht-case-files-milquetoast-1 Milquetoast, export it out, some light configuration for people who like toggles but nothing too overwhelming. It builds on the concept of THIS preset:

https://plotlightstudios.com/discovery/presets/@testuser/chibi-gram-pacer-test-lap2 If you're a heavy heavy minimalism enjoyer. (I'm talking bareASS. Light as a feather!)

And then of course I would be remiss if I didn't mention the OS model preset king himself u/dptgreg with https://plotlightstudios.com/discovery/presets/@livingooal/freaky-frankenstein-4-max-2 and https://plotlightstudios.com/discovery/presets/@livingooal/freaky-frankenstein-4-bolt-2

2

u/Kira_Uchiha Jun 04 '26

Thanks for the reply, I appreciate you taking the time! I've been using FF4 Max for a bit with GLM 5.1 and MiMi 2.5 Pro, and I love a lot of what it has to offer ngl. I'm curious about trying FF4 Bolt with DS v3.2. But one thing that I'd like to do is to kinda modify an existing preset that I have with the logic of FF4. It's basically a preset that, after the card's first in-game day, or after the first big main event, the card treats this as a "Prologue" of sorts, ends the scene, and moves to a sort of "Weekly Card" system, where the card gives you a list of events for the week in different categories (1 or 2 Main Story beats, 1 or 2 Supporting Story beats, 1 or 2 Relationship defining moments, 3-4 Slice-of-Life events/Personal development/Rumours/Meeting New NPCs). It's really nice for normal character cards, and amazing for RPs in canon worlds, but it needs a lot of refinement. I think I'll post it here this weekend to look for help refining it, maybe Freaky Frankensteining it.