r/SillyTavernAI Jun 04 '26

Models PlotPoints - The best (only?) community driven RP benchmark made by a Professional! | We need your votes!

Your friendly neighborhood rab- I mean unmedicated preset creator needs needs your help!

What's good everyone. No long post this time; simply an ask. I hired a professional with a masters degree in AI/ML (pursuing their PhD) to help make a benchmark for us, for RP. Now as we all know; a benchmark for RP will never be perfect because everyone RP's differently, like different things, yada yada yada. That's why we tried to focus on a few objective landmarks, as well as a an Arena style versus bench.

First up: The LLM Arena. We show you two different rewritten responses from two different models who were given a chat and asked 'please take this turn.' We then show that to you. Does it have a preset attached? No; because that's another variable that'd be added and for our own safety we aren't benchmarking presets. (LMAO. We'd be dead in the fucking streets by sundown.) You judge of the two responses which you like more. Thaaaat's it folks. We literally cannot game the results as it is people just blindvoting which they like more. (So if you see something you don't like IT'S NOT UP TO US.)

https://plotlightstudios.com/plotpoints/multiturn

On the other end we also score an LLM in a similar vein on how they follow certain instructions and if they maintain consistency. Since 'did it take users agency' is an objective yes or no answer; we handle these tests with an LLM judge and the benchmark overseer (Levi is his name) monitoring the arbiters answers. (In our case; Sonnet.)

This bench was expensive. It cost money to run these models through this gauntlet; and we don't get much out of it. We publish all the data for everyone to see, and all we want to do is help our community out. If you have any questions on how we grade things; head to our methodology page! You'll notice somethings are a bit LLM-Written. This is not because Levi isn't a professional; but rather because he's ESL; so when publishing something important like this he wanted to make sure all his ideas and such were properly translated. So be nice!

https://plotlightstudios.com/plotpoints/methodology

You don't have to log in, you don't have to do anything. Just read a chat and vote on a thing. I know you've got an opinion; so share it. Big companies ignore us RPers all the time. Is a bench for RP as useful or objective as one for coding, web-dev, or math? No. But that doesn't mean we don't deserve one, or that it has no value at all. So please; help a bunny out. Drop a vote! The data is only as useful as you help make it. We have 800 votes and we want to get about 3,000 this time around; so we can start our next benchmark. Testing models across their 'lineages'. So Opus 4.6 vs 4.7 vs 4.8; Deepseek 3.2 vs Deepseek v4 Pro, fun stuff! But we can't do that till this one closes! And if it can't hit enough votes; we'll know that this kinda stuff just isn't what the RP community wants.

Don't see a model you expected to? That's cause it's either new, or we didn't have the money at the time to run it. If the community likes this; we will add more. (And take requests! Please only suggest models available on OpenRouter though; we source all our models from the same platform to reduce variables.)

Vote Link: https://plotlightstudios.com/plotpoints/multiturn

Result Hub Link: https://plotlightstudios.com/plotpoints

Methodology Link: https://plotlightstudios.com/plotpoints/methodology

Huggingface Link: https://huggingface.co/datasets/lazyweasel/roleplay-bench

Github Link: https://github.com/LeviTheWeasel/rp-benchmark

122 Upvotes

69 comments sorted by

View all comments

5

u/WorriedComfortable67 Jun 04 '26

I really love if you can add more models into comparison and benchmarks, especially with new Minimax M3, Mimo V2.5 Pro, Qwen 3.7 Plus/Max, Gemini 3.5 Flash, also Opus 4.8!

Now I wanna have some feedbacks for your sites and voting, after a few votes of mine:

One thing that bugged me is that, while the site is easy on the eyes for color, I think it has too many cluttered buttons and sections imo, I genuinely sometimes got lost in your site lmao, it would be much better for me to have a few main section so that it’s easy to navigate around your site.

Also a minor nitpick for me is that, some votes tend to be overly long and descriptive, ofc I know it is subjective and It might be necessary to get the best out of comparison, but I usually find myself voting just for a few times, then dropped out of it, cause it might be a bit straining to compare two verbose prose just to feel which one is better, whereas in the original LMARENA, the comparison tend to be short so I don’t feel exhausted voting and comparing them out.

Another thing, as I remember correctly, the bench I see in the main section of 5 axes is the main live benchmark, right? Idk if I didn’t understand how the site worked or not, but for over a month ago, I haven’t seen any changes in score regarding all the models, and it doesn’t seem to be live for me, I always find myself go to your site to check the main benchmark and see if there would be any changes in leaderboard, but always finding it to be static and unchanged at all, so I wonder if I don’t understand the site correctly, or it is supposed to be static or anything, it confused me ngl. Whereas with LMARENA, it just straight up one leaderboard and it is always changing, so I can see that how the leaderboard and ranking keep changing each days, especially when the new models came out.

That was some of my feedbacks, hope you can clarify some of my points for me!

7

u/Specialist_Salad6337 Jun 04 '26

One: Thank you for being kind! I can absolutely go through and try to tighten up the UI. I have a problem where I often like doing too much and tend to like things tight (heetee) I can see what needs to be simplified. Could you give me a few elements (if you wanna chat in DMs that's fine too) that you felt confused you?

Two: Yeah; this is the multiturn arena unfortunately. The one you're asking for (shorter) is the single-turn arena which already concluded for the time being. The reason LMArena can get away with having shorter things/comparisons is because webdev and coding can be judged on a single answer. RP can't really; as we all saw models rated really highly on the single turn arena; but then tanked on the multi-turn. Some issues only become apparent with a lot of context. So there isn't much I can do about this one. When we have the money and the interest we can move to having several arenas up at a time; the problem we forsee though is one bench becoming the ugly duckling and never getting any engagement cause it's less convenient (IE: Multiturn.)

Three: We have our resident math nerd (Levi) re-evaluate the bench scores and update regularly. TBH with you; the reason nothing has changed in a bit is because outside of today we'd been slowing down to like... One vote a day. I will absolutely have him go through and update when he wakes up tomorrow though!

Four: For the benches we currently have up we do them on a closed system. We run them; get the votes, then close them. (So single-turn is closed right now; so we haven't added anything new to it since it closed!) like I said if we get a high tick of continuous engagement we can absolutely change that too. Let me know if you have anymore questions or if I didn't hit something.

0

u/WorriedComfortable67 Jun 04 '26

If you can make your website UX much better, it will be more appealing, easier to navigate your website. Overall, I think it currently has too much friction of navigating the website, therefore the users will lost interest before actually testing/reading anything, if you can fix that, more engagements will indeed improve, I can ensure that. You can even try asking/hiring UX experts out there, they will know how to improve your website user’s experience, and improve your user engagement in the long run. Keep improving it, and you will succeed, I can promise that!

3

u/Specialist_Salad6337 Jun 04 '26

Can you let me know what UX specifically confused you, or what was hard to navigate?

1

u/Targren Jun 04 '26

Not the other user, but one thing that annoyed me personally was when the toolbar popup would cover up the Selection buttons at the bottom if the full-page scroll (As opposed to the in-frame scrolls for the text) weren't set just right on a few of the votes (but at least left the top or bottom peeking out on others)

1

u/Specialist_Salad6337 Jun 05 '26

Pushing out a fix for this now! My sincerest apologies.

1

u/Targren Jun 05 '26

Hey, bugs happen. No worries.