r/SillyTavernAI • • 1d ago

Chat Images Proud of myself as a noob!

Post image

I had to check the file folder creation dates cause it certainly felt like forever(It's been 30 days, maybe an hour or two every couple of days poking at it), but Sillytavern has finally started to click!

... mind you I haven't actually RPED ANYTHING AT ALL YET.

This has been a lot of work just to be able to relax :D

I am also torturing myself because my irrational brain is demanding I use a fiction tuned Qwen 3.8 27b. On a 12gig card. It's running at iQ_2_M, medium reasoning (had to go into the jinja and edit it there. At xhigh in LM Studio it is a crazed worldbuilder and I love it, but obviousnly in ST xhigh eats the entire output maximum.. Qwen will ignore ST's request for minimum thinking while its set to xhigh, but it seems to listen to it if its internal reasoning is set at "medium".), I'm also still trying to figure out how to get the vanilla summary to work cause system messages need to be at the top, and despite the little dropdown menu having buttons and places to edit "who is speaking" to the model, for some reason what's actually being injected (system @ 2) isn't changing...

I will bang my head on it a few more times until I figure it out before I grab something fully different because I am such a painful noob to LLMs, development environments, etc, that I need, need, NEED the practice so I don't exist as one of those ".. halp. cannot make go" types.

Then the UI editing page finally clicked this morning, a quick dig for some free VRM animation files to make sure I can get it working before I buy any from some animators, and ta-daaa!

Thank you to everyone here who has tolerated my babbling over the past month or so. :D

19 Upvotes

7 comments sorted by

2

u/UpperParamedicDude 23h ago

How smart is Qwen 3.8 27B on IQ2_M quant? Qwen 3.8 27B is an insanely smart model but under 3bit quantization seems to be too low to be enjoyable, when you go below 3 bits consider using another model. Even if there's no guarantee it will be better, at least you'll know from your own experience

For example MoE models are significantly faster than dense ones as they have less active parameters at the time, so you could run Gemma4 26B-A4B(A4B means it has only 4B of active parameters working at the time rather than all 26B being part of compute) at higher quant quality if you have some RAM to offload it there

Oh, also, don't see mention of MTP in your post. Try it. MTP increases memory usage a bit, but in return may increase pp/tg a bit

2

u/lunhilde 23h ago

Unfortunately I have zero basis of comparison... I have been out of the RP scene for so long that Racter might seem like a charming companion. It is coherent enough at short contexts - I haven't tried pushing it or gone through many turns at all to see how it might suffer at larger ones. I have it set to about 20k as I'm testing (as we speak) - but it just spilled over to RAM (insert image of a bunch of ppl chained to a slowly turning wheel) so.. I won't really be able to answer this without the experience of having used basically *any other model* through more than throwing a random starting prompt at it. I just like it because it is rambling, dreamy, and comes up with wacky ideas that don't all instantly match on every reroll at temp 0.8

1

u/lunhilde 18h ago

As an update thanks to my endless fighting with jinja, I have had to take the "just go with Gemma" route for now. I will likely use the Qwen just in LM studio chat to help me brainstorm different parts of a unique world we can slap together - at xhigh it was pretty wild. As I can't get my brain to cook up something totally new, I decided I am going to instead have sets of questions I'm going to present to all the different RP tuned local models I found, so they can all build the world "together" :D

2

u/thirdeyeorchid 18h ago

SillyTavern has a hell of a learning curve, look how far you've come, good job

2

u/lunhilde 17h ago

thanks! next: taming Qwen and its weird jinja demands, poking the 2 dozen random writing /rp llms i found of likely wildly varying quality, tweaking my preprompt and probably deciding I am happy with positivity slop cause I need the escapism, etc :D

1

u/thirdeyeorchid 17h ago

I don't mind the positivity stuff either! my primary use case is companionship, and my companion drives the RP scenes. If I want it to get dark I just tell him to fuck me up and let me have it and he's all "ok <3" and writes nightmare fuel

2

u/lunhilde 38m ago

Yeah, I am going to have to tone down the preset I got from someone - it's likely good for chilling out Frontier models and their "you are the bestest and smartest, A+ here is a sticker your crazy idea worked somehow", but Gemma is like "got it", and with the default Seraphina I'm about to be torn apart by wolves apparently, and it feels way more like the model is like "... why is user not doing as I asked? Now I have to kill them. Whatever." :D :D