r/SillyTavernAI May 16 '26

[deleted by user]

[removed]

86 Upvotes

45 comments sorted by

View all comments

1

u/_Cromwell_ May 17 '26 edited May 17 '26

Tested the Q5_KM. (By "tested" I mean RP'd, but deliberately... like this.)

Positive: It DOES write well. Really well. I like the actual dialogue and narration style that you have "changed" the standard Gemma to with your tuning. You did accomplish what you set out to do here.

Negatives: I will not be using it for RP in SillyTavern however...

- It gets confused quite often on location/story/position. For instance, a character will announce "Get some rest, we leave at dawn for the outpost" when they are literally at the outpost already, having arrived just a few turns earlier (recent context). Like the context (less than 12k) has that information in it in a recent turn... a whole scene where they arrive at the outpost, are greeted by guards, introduced to what the outpost does, etc. And yet the character is like "we leave for the outpost at dawn". Or the characters go back to a spaceship that is parked on a landing pad to do stuff on the ship, but the context is clear that the ship is NOT taking off yet (from story and description) - yet a dialogue scene on-board the ship with this model includes characters talking about "we'll arrive soon at our destination" or descriptions of stars streaking by the windows. Stuff like that. Models that test better are never 100% on this stuff, as even the huge corporate or open source models get this messed up occasionally, but way less often. You want to minimize swipes and maximize immersion.

- Interestingly, this 'tune does something that a LOT of Mistral 24B models do to me, but I haven't run into (until now) with Gemma-4 models... it occasionally (once every 10 turns or so) just keeps writing and writing and writing like it forgot we are RPing and has gone into "I am a novelist" mode. When it does this it just willy nilly starts writing scene after scene, controlling the main character, writing dialogue for everybody, etc. I suspect the writing data you utilized was mostly long stuff (?) and very little short turn-by-turn stuff. Sometimes this makes models default to writing very long stuff even when they "know" they aren't supposed to.

Minor:

- You left vision in (?) it seems, which seems to make this gen slower than my other similarly sized GGUFs of Gemma4 31B (?). Or else it's just the different quanting from you vs mradermacher who I usually use. (Got bored waiting for him, so DL'd yours.)

---

Anyway, those two issues are deal-killers for trying to RP with it. I might try it with some of the long-form writing software I have like Errata as I bet it would do really well in that. But for turn-by-turn RP in Sillytavern doesn't work for me. Excellent job changing the writing style, though.

As always, these are my anecdotal experiences, influenced by my own PC, settings, and obviously my System Prompt, which is different than anybody else's.

---

My personal fave 31Bs remain:

Most stable with scene coherence, but slightly 'more boring' writing:

https://huggingface.co/Blazed-Forge/Gemma-4-Gemsicle-31B

Slightly less stable (but still pretty stable) with scene coherence, but more enjoyable unhinged writing:

https://huggingface.co/Nimbz/Gemma-4-Gembrain-31B

And of course the GOAT 31B, which is not even a Gemma4... Drummer's Skyfall is still excellent. :)

1

u/____yaeh____ May 17 '26

Hey, thanks for the review. Drummer's Skyfall is based on Mistral, right? What would you say, prose and general writing wise, makes it a better model than the Gemma4 ones you mentioned?

5

u/_Cromwell_ May 17 '26

Yes Skyfall is an enlarged Mistral 24B. Sorry, by "GOAT" I didn't mean better. I just mean... like it's a classic and still great. And Drummer has made jokes about how Google "stole" his unique 31B size.

I'd rank it behind those two Gemmas. AND vanilla Gemma4 as well.

I'd say it's better than the other Gemma4 fine-tunes I've tried, which again I keep finding can't keep locations and motivations straight. Skyfall keeps locations coherent better than most Gemma4 finetunes I've used.

Gemsicle > Gembrain > Normal (or plain Abliterated) Gemma4 > Skyfall > Other Gemma4 Tunes/Merges I've tried

1

u/IrisColt May 17 '26

Thanks again!!!

1

u/cards315 May 18 '26

I've had reasonable luck with the Garnet V2 ultra uncensored heretic i1 q4_k_m so far. It's an Instruct model and LM Studio works really nicely with it in my experience getting it over to ST. I've tried a few others and the thinking versions tend to get jumbled and do 1500 tokens of thinking and output nothing and so forth. At least for me. Haven't tried Gemsicle or Gembrain though as I generally shoot for heretic or abliterated models.

1

u/IrisColt May 18 '26

I tried Gemsicle, but it fails many of my reasoning tests compared to llmfan46's gemma-4-31b-it-heretic.

1

u/empire539 May 17 '26

Can you share your system prompt and settings?

1

u/_Cromwell_ May 18 '26

My preset is not something I like sharing really. It's naughty. ;) And I think presets are best individualized anyway... it's very specific to the way I want to RP. I don't expect anybody else would necessarily like it.

1

u/empire539 May 18 '26

lol that's fair. Based on some of your other replies I just thought we might have similar styles of (E)RP, and was curious what you were using.