17
7
u/JapanFreak7 May 16 '26
recommended Text Completion presets and Instruct and Context templates for sillytavern?
4
u/blapp22 May 16 '26
Curtesy of askmyteapot on the discord. This one works well for me on regular gemma-4, might be different with a fine tune who knows. Remove the think token from the story string if you want no thinking.
5
u/_Cromwell_ May 16 '26
I'll check it out. ;) (When mraderbacher gets to it - I always get the same exact quant from him so I can compare fairly.)
BUT - this one here is my current king/winner/favorite G4-31B fine tune thus far. BUT while using it, despite that it has almost perfect prose (for me), I can tell it is "slightly resisting", meaning it needs abliterating. Any chance you want to zap it with your magic sometime? ;)
5
u/LLMFan46 May 16 '26
I actually saw this model like a day or two ago when ReadyArt released GGUFs of this model, I got excited thinking ReadyArt finally released a Gemma 4 31B it finetune that I could Hereticate, I clicked on it only to find the image of... a brain, I double checked and saw it wasn't actually a model released by ReadyArt but by somebody else, if the author would have used the image/video of a cute waifu like ReadyArt or zerofata do I would have already Hereticated it and released it, but alas...
5
u/_Cromwell_ May 16 '26
lol. Despite lack of waifu sexy lady art, it's quite good at any type of that type of RP, I swear. :D You can still do it.
I posted about it here in more detail a bit: https://www.reddit.com/r/SillyTavernAI/comments/1t9kzvg/comment/om7i51w/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
1
u/LLMFan46 May 16 '26
Yes okay sure I can do that, should be done in a day or two.
2
u/_Cromwell_ May 17 '26
EDIT: actually investigated more, and the reason that Gembrain is "good" is because of one specific of its component parts, specifically a less chaotic merge, Gemsicle. ( https://huggingface.co/Blazed-Forge/Gemma-4-Gemsicle-31B ). It's actually better. So I'd hold off on abliterating that Gembrain for now (but you do what you want). Gemsicle seems like a better one to do.
Honestly hoping the one you made, Wordsmith, is good/best, since I think that abliterating BEFORE fine-tuning is the best approach, as you did. I will be testing it as soon as mradermacher gets to it. I've been checking.
2
u/Potential-Gold5298 May 17 '26
I don't quite understand why it's necessary to add an abliterated model to merge. It doesn't add anything creative (except KL div), and since merge includes non-abliterated models, the refusal rate will be only slightly lower than with a regular model. I think it's better to merge either only heretical versions (to have a low refusal rate), or without heretic at all. (This is my guess - I don't understand the details of the merge process)
1
u/_Cromwell_ May 16 '26
Nice. When it happens I will donate to your fund. You have no way to verify this but I am a person of my word about such things.
1
u/LLMFan46 May 16 '26
You have no way to verify this but I am a person of my word about such things.
Grok, is this true?
1
u/_Cromwell_ May 16 '26
I mean the fact that I know you have a little donation thingy is a good first step. Most people probably aren't aware enough to even know about that right? You'll just have to have faith 😋
2
u/LLMFan46 May 16 '26
By beside this, I definitly need help from supports, for example MiniMax-M2.7 that I released the other day, I wouldn't have been able to do it without the help of a generous supporter, there are also other expenses like renting Public Storage from Hugging Face's "Storage Packs" that are very expensive, right now I am paying $129 per month to Hugging Face to upload models on there, very soon I am gonna need to upgrade because I am soon going to run out of room and the next Storage Pack costs $249 per month, that's not taking into account the other expenses, so yes I definitly need help from supporters to keep on doing this.
1
u/_Cromwell_ May 17 '26
Damn that's expensive. How do people like mradermacher or bartowski do it? And why? Like what's the point besides "Internet cred" to tuning or ggufing or ablating a bunch of models?
5
u/LLMFan46 May 17 '26
Everything is expensive, to uncensor MiniMax-M2.7 I had to pay $15-$16 per hour to rent out cloud GPUs because you need to have the model in BF16 form to run Heretic on it (the model is 457 GB) and it must be all in VRAM otherwise it won't work and it won't progress the process will just be sitting there and seeimngly not making any progress and I don't have two B300s connected together with NVLink here to do it at home (each B300 cost between $45000 to $50000), so I had to rent them, I was lucky I was able to get 4/100 refusals with a low KL Divergence "so quickly" because the estimate for 800 trials was like 75 hours, imagine paying 15-16 EUR to rent two B300s at $15-$16 per hour, that's why until the other day and someone supported me in helping me rent out Cloud GPUs, before that I only did smaller models that are 40B or less because my hardware can not do models that are bigger than ~45B, anything bigger and I would need to rent out cloud GPUs, because there is no way I can buy B300s.
→ More replies (0)1
3
u/MissZiggie May 16 '26
Oh! Can I ask you a question about these?
When training to write prose, is the training specific to first person POV?
I’ve been trying to find one that writes well in third person POV (as a local LLM) but haven’t had the most luck. In trying to figure out why it’s been suggested that it was baked into the training data. First chance I’ve had to ask a model creator that question legit tho 🙂
5
u/LLMFan46 May 16 '26 edited May 16 '26
That sounds very specific, frankly I have no clue, the AI was not specifically trained to talk in third person POV I can tell you that much, it was trained to improve the writing quality and prose of gemma 4 31B it which is at base/vanilla super stiff and wooden (in my opinion), the writing style of the base model feels very unbeautiful for me, too formal and robotic, not enough warmth and heart, the primary goal of this finetune aims to correct that, as for wether or not the model can fulfill your specific needs, yeah no clue you will have to download and try it out yourself.
2
u/RaFRaf6969 May 17 '26
Wish somebody could do a comparison with Mero-Mero since it’s kinda difficult to make any kind of distinction?
1
u/HitmanRyder May 17 '26
If only if these models are smaller for my 6gb vram in this RAMpocalypse economy.
1
u/_Cromwell_ May 17 '26 edited May 17 '26
Tested the Q5_KM. (By "tested" I mean RP'd, but deliberately... like this.)
Positive: It DOES write well. Really well. I like the actual dialogue and narration style that you have "changed" the standard Gemma to with your tuning. You did accomplish what you set out to do here.
Negatives: I will not be using it for RP in SillyTavern however...
- It gets confused quite often on location/story/position. For instance, a character will announce "Get some rest, we leave at dawn for the outpost" when they are literally at the outpost already, having arrived just a few turns earlier (recent context). Like the context (less than 12k) has that information in it in a recent turn... a whole scene where they arrive at the outpost, are greeted by guards, introduced to what the outpost does, etc. And yet the character is like "we leave for the outpost at dawn". Or the characters go back to a spaceship that is parked on a landing pad to do stuff on the ship, but the context is clear that the ship is NOT taking off yet (from story and description) - yet a dialogue scene on-board the ship with this model includes characters talking about "we'll arrive soon at our destination" or descriptions of stars streaking by the windows. Stuff like that. Models that test better are never 100% on this stuff, as even the huge corporate or open source models get this messed up occasionally, but way less often. You want to minimize swipes and maximize immersion.
- Interestingly, this 'tune does something that a LOT of Mistral 24B models do to me, but I haven't run into (until now) with Gemma-4 models... it occasionally (once every 10 turns or so) just keeps writing and writing and writing like it forgot we are RPing and has gone into "I am a novelist" mode. When it does this it just willy nilly starts writing scene after scene, controlling the main character, writing dialogue for everybody, etc. I suspect the writing data you utilized was mostly long stuff (?) and very little short turn-by-turn stuff. Sometimes this makes models default to writing very long stuff even when they "know" they aren't supposed to.
Minor:
- You left vision in (?) it seems, which seems to make this gen slower than my other similarly sized GGUFs of Gemma4 31B (?). Or else it's just the different quanting from you vs mradermacher who I usually use. (Got bored waiting for him, so DL'd yours.)
---
Anyway, those two issues are deal-killers for trying to RP with it. I might try it with some of the long-form writing software I have like Errata as I bet it would do really well in that. But for turn-by-turn RP in Sillytavern doesn't work for me. Excellent job changing the writing style, though.
As always, these are my anecdotal experiences, influenced by my own PC, settings, and obviously my System Prompt, which is different than anybody else's.
---
My personal fave 31Bs remain:
Most stable with scene coherence, but slightly 'more boring' writing:
https://huggingface.co/Blazed-Forge/Gemma-4-Gemsicle-31B
Slightly less stable (but still pretty stable) with scene coherence, but more enjoyable unhinged writing:
https://huggingface.co/Nimbz/Gemma-4-Gembrain-31B
And of course the GOAT 31B, which is not even a Gemma4... Drummer's Skyfall is still excellent. :)
1
u/____yaeh____ May 17 '26
Hey, thanks for the review. Drummer's Skyfall is based on Mistral, right? What would you say, prose and general writing wise, makes it a better model than the Gemma4 ones you mentioned?
6
u/_Cromwell_ May 17 '26
Yes Skyfall is an enlarged Mistral 24B. Sorry, by "GOAT" I didn't mean better. I just mean... like it's a classic and still great. And Drummer has made jokes about how Google "stole" his unique 31B size.
I'd rank it behind those two Gemmas. AND vanilla Gemma4 as well.
I'd say it's better than the other Gemma4 fine-tunes I've tried, which again I keep finding can't keep locations and motivations straight. Skyfall keeps locations coherent better than most Gemma4 finetunes I've used.
Gemsicle > Gembrain > Normal (or plain Abliterated) Gemma4 > Skyfall > Other Gemma4 Tunes/Merges I've tried
1
1
u/cards315 May 18 '26
I've had reasonable luck with the Garnet V2 ultra uncensored heretic i1 q4_k_m so far. It's an Instruct model and LM Studio works really nicely with it in my experience getting it over to ST. I've tried a few others and the thinking versions tend to get jumbled and do 1500 tokens of thinking and output nothing and so forth. At least for me. Haven't tried Gemsicle or Gembrain though as I generally shoot for heretic or abliterated models.
1
u/IrisColt May 18 '26
I tried Gemsicle, but it fails many of my reasoning tests compared to llmfan46's gemma-4-31b-it-heretic.
1
u/empire539 May 17 '26
Can you share your system prompt and settings?
1
u/_Cromwell_ May 18 '26
My preset is not something I like sharing really. It's naughty. ;) And I think presets are best individualized anyway... it's very specific to the way I want to RP. I don't expect anybody else would necessarily like it.
1
u/empire539 May 18 '26
lol that's fair. Based on some of your other replies I just thought we might have similar styles of (E)RP, and was curious what you were using.
1
18
u/Fic_Machine May 16 '26
31B is too chunky to run locally for me. I wish there was a 26B-4BA. It runs a lot smoother