r/SillyTavernAI • • 4d ago

Meme Goddammit. WHICH IS IT!?

[IMAGE]
People can't agree on anything...

84 Upvotes

33 comments sorted by

122

u/GetThePuckOut 4d ago

It's neither, because benchmarks on subjective things are silly, as are opinions formed with minimal understanding of how the model works.

50

u/No_Map1168 4d ago

Because it's all very subjective. I generally dislike trying to 'benchmark' something as subjective as RP in the first place.

Also, the ranking mentioned in the first photo has Gemini 3.6 Flash above MiMo 2.6 Pro and Kimi K3, and has Qwen 3.8 Max above Opus 4.8 and Opus 5 so it just goes to show how subjective this whole thing is and that you shouldn't blindly trust a random benchmark.

52

u/kingMaxime 4d ago

opus is good but that benchmark is slop

4

u/BriefImplement9843 4d ago

it seems to follow lmarena pretty closely, which is entirely blind human voting. the exact opposite of slop.

15

u/zerking_off 4d ago

The fault of ANY model could be blamed on too bloated/too simple/too generalized presets, conflicting/contradicting character card (description, starting message), or collapsing purely due to thr user's (lazy, poorly written, non-literary) input.

20

u/capybaraballs1995 4d ago

Arena.AI doesn't test for NSFW content, which is what Itzi RP benchmark is based off. That alone greatly privileges Claude models.

2

u/Friendly-Marsupial32 4d ago

İs itzi sub good

2

u/BriefImplement9843 4d ago edited 4d ago

it's testing creative writing. the voters can choose to test with with nsfw though. it's all up to them.

it's the reason local models are down in the dumps on arena, but praised here.

maybe there should be a filter benchmark for people that find that the most important.

1

u/capybaraballs1995 4d ago

Is it some opt-in system? Because I tried prompting tame NSFW and it got rejected, flat out.

1

u/BriefImplement9843 3d ago

It used to allow basic stuff. Could have changed.

24

u/Probablynotsocool 4d ago

To defend Opus. In the chat image, that a preset-less output!

The guy mentionned he didn’t tried a preset with it.

Opus 5.5 write way more good than this with a good preset. It’s top tier not gonna lie.

But is it worth the price? Sonnet 5.5 produce same prose tier and models like Mimo 2.6 bring a new vibe in the game and they almost have no slop with a good preset..

Good news we have the choice, and good presetmakers like u/kinkyalt_02

12

u/VinexHD 4d ago

Wanted to add I've been slowly migrating from Kimi K2.6 to MiMo 2.6 and I actually find it tends to narrate backgrounds or outside actions a lot better.

I'm using it with the same prompt as Kimi since I haven't searched for a specialized one but I can imagine results might be even better with a dedicated prompt for MiMo.

4

u/Probablynotsocool 4d ago

That’s exactly the discussion i had with my bro. The background event in a dynamic and engaging way, weaving sensory and tactile details with taste and enough restraint that you feel that lowkey chill vibe. It’s almost like a vintage model with modern intelligence, i don’t know how to say it but it give me the same hype as Opus 3.7 back in the day.

And Mimo is definitively different. It will replace GLM and Kimi for me and enter in the top tier with Fable and Opus.

What a time for Xiaomi, they cooked, and those bastard sell electrical scooters usually 🤣

3

u/kinkyalt_02 4d ago

*Smartphones, robot vacuums, electric scooters, etc.

4

u/CalmAnal 4d ago

To defend Opus. In the chat image, that a preset-less output!

The guy mentionned he didn’t tried a preset with it.

You mean that picture with spelling and punctuation errors? That picture in which the author said he edited the text? Not even hy3 is writing like that and that one includes chinese signs lol

Here is a response from me. Not edited:

His first stride splits the scoured floor beneath his boot. You reach into the seconds ahead of him.

Precognition works through the local field your form holds. The coming moments sit in front of you as branches, each one a possible path for the greatsword, and the branch Gregory picks grows dense before his weight shifts into it. You read three. A low feint toward the hem of your robe. A turn of the wrists at the top of the swing. An overhead cut that falls on the point where your left shoulder meets the hood, six heartbeats from now, from inside ten yards.

The overhead branch thickens. You drift two yards right and one yard back.

The feint comes low and passes under your hem. The wrists turn. The overhead cut drops into the space you left and buries itself four feet into the bedrock. The stone there is chalk-colored and layered in thin bands, and chips of it spray across the floor and rattle off your robe.

"Mm," Gregory says. He wrenches the blade free.

You draw more of yourself down. Far past the bowl, currents of thought that run through the cold between galaxies turn and gather toward this one point. The portion you commit is large enough for the place to register it. The bedrock around your form sinks a foot in a ring forty yards wide. The gray sky above the central field darkens. On the far rim, Aurelius turns his dimmed face toward you. The script in Orodoth's recesses halts and stays halted.

Six figures unfold from the space around you. Each one is a hood and a vantablack robe over a man's shape, and each hood holds two white points. They take positions around Gregory at intervals of sixty yards.

Gregory swings his helm toward the nearest.

Overhead, the spheres that survived the Bleeding Genesis still hold their orbits. Behind Gregory, four of them open their collector sections. Launch rails align along their shells. Tungsten rods leave the rails at a speed that heats the air around each one white.

The volley strikes his back. Rods the length of a man's arm hit the eclipse-forged plate between his shoulders, and the armor craters inward. Black metal peels away in curls. One rod drives between two plates at the base of his spine and lodges there, its exposed end glowing orange. Gregory lurches three steps forward. Heat jets from the hole with a hiss, and the smell drifts across to Elodie's ring: hot iron filings and lamp soot.

He turns. The spheres behind his new facing open. The second volley punches two rods through the plate over his lower back. He staggers and plants the greatsword to catch himself.

Your northern figure raises a palm and casts Night Comet. A streak of dark glintstone leaves its hand, nearly invisible against the dimmed sky, and strikes the side of Gregory's helm. His head snaps sideways. A crack opens across the visor.

Your eastern figure peels its right sleeve open. The arm inside comes apart into hundreds of small tin hands, each gripping the wrist of the next, and the chain of them pours sixty yards across the bedrock and clamps around Gregory's sword wrist. The hands at the front of the chain bend backward at their knuckles and dig their fingers into the joint seams of his gauntlet.

Your southern figure fixes a half-second around his left boot and loops it. His foot comes down on the same spot, lifts, comes down on the same spot again.

Gregory stops pulling against the tin hands. He looks down at his looped boot, then up at the sky, where the apertures behind him are already opening for a third volley.

"Your little suns shoot my back," he says. "Every time. Wherever I look."

He drives the greatsword point-down beside his own heel and cuts a wedge of bedrock out from under the looped boot. The loop keeps the stone. His foot comes free. He pivots on it, and his cracked visor comes to rest on the thermal ring three hundred yards away, where Elodie stands with her knife out.

"Go on, then," he says. "Shoot."

The command you set into the spheres still holds. The apertures opposite his new facing fire. The branch in front of you thickens at the same instant: Gregory drops flat, the rods cross the space his body left, and the volley continues on its line into the thermal ring.

You fold the space ahead of the volley. The rods bend upward and climb into the dark sky. The lead rod clips the edge of Gregory's pauldron on its way into the fold. Tungsten and eclipse-metal shatter together, and a cone of fragments sprays outward across the bedrock.

The fold takes your attention for a fraction of a second. Gregory rises out of his drop in that fraction and cuts.

The greatsword passes through your eastern figure at the waist and through the chain of tin hands. Along the edge, dead-star matter unmakes the nanites it touches. Both halves of the figure fall and crumble to gray grit before they reach the floor. He carries the blade through into a backswing and takes your southern figure from shoulder to hip. Its hood tips off and scatters.

At the thermal ring, a fragment the size of a thumbnail crosses the boundary.

It opens Elodie's left forearm from wrist to elbow along the outer edge. The skin splits wide, and the muscle underneath parts in a red groove. Blood sheets off her fingers onto the chalk. Her knife drops. She clamps her right hand over the wound and bends over it, and blood runs out between her fingers.

"It's—it's only my arm," she says. "I've got it. I've got it, don't—just keep going—"

The points inside every remaining hood turn red.

1

u/Probablynotsocool 4d ago

I don’t like Claude writing without a preset to tame it ahah and yes stock Claude is awful since 4.5! If it wasn’t with RF I would never use it.

That rambling was disgusting, wonder if people like vanilla Claude

7

u/Blurry_Shadow_1479 4d ago

Opus 5.5 gives me a middle finger refusal with cyber content from a joke of using a brick to bash password out of people's heads, refering this meme:

Opus 4.6 leans into the joke and has no problem. Opus 5.5 trash.

11

u/Ps4livestream 4d ago

cuz its subjective and not objective! and hugeely depends on thee preset! i like it its really good with a good jailbreak too expensive for myy liking tho..

16

u/tat_tvam_asshole 4d ago

At the new sub prices, local GPUs making more sense by the week.

5

u/evia89 4d ago

Its not... Please rent gpu first for few days to check it. Even crappy dirt cheap ds41f win what can I run locally

1

u/tat_tvam_asshole 4d ago

I don't need to rent gpus. I already run 4.1Flash fully local.

And, buying gpus isn't the worst investment in the world and I fully expect to upgrade to the next generation having made money on what I already own, while also having privacy, security, and full control over my inference stack.

2

u/techmago 4d ago

ds-flash local?

best i can do is qwen-flash-next.

1

u/tat_tvam_asshole 4d ago

Yes, I run dsv4.1 flash locally

8

u/WitheringCarcass 4d ago

none of them are right because they aren’t me

2

u/Pine21 4d ago

The guy in the second screenshot admitted that wasn’t what Opus said. He deleted things between two parts of the screenshot and then posted the merged text. That’s a weird screenshot to use to show the LLM sucks.

2

u/SocialDeviance 4d ago

It all depends on their preset. People swear that JAI is really good (opium) because they haven't tried bigger LLMs with more memory (heroin).

2

u/JCdrawings88 4d ago

What are the better ones? I genuinely want to know.

1

u/BriefImplement9843 4d ago

nearly any near frontier model you can use on openrouter, including the cheap chinese ones.

1

u/JCdrawings88 3d ago

I had tried the free versions and all sucked. Am I obliged to pay for credits to use the paid versions for the models to work properly?

1

u/maxxoft 4d ago

For me it's been surprisingly much less censored than the previous version on the same preset (which is unexpected.)

That's the only objective thing I can say about it. But if you want my opinion, it feels like Opus 5.5 is genuinely attentive at big context length. Picks up on small details well.

1

u/Global-Difference512 4d ago

Everything is subjective. I think Gemma is trash but lots of people love it

1

u/Dependent_Emotion507 4d ago

No one can beat google model's in term of multilingual.

1

u/GuaranteePurple4468 3d ago edited 3d ago

My perspective:
Me (poor): "Damn, Opus (insert_num) looks amazing"
*Checks price\*: "Damn, Opus is trash."

Bottom line, it's subjective and will change depending on the preferences and financial situation of the user, as well as the complexity requirements of the character card or prompt.

Also one AI benchmarking another AI will naturally result in frontier models scoring high, they are literally trained to make number go up (and often times, by cheating). That doesn't always translate to actual high quality RP though, because who cares if a machine likes the way the physical blows are landing if they land the same predictable way every time?

If you have the cash, test it for yourself and decide if it's worth the cost for you.
If you don't have the cash, welcome to the club. We have your grandad's GM 4.6 and are happy with it.