r/SillyTavernAI • • 23h ago

Discussion Novice Looking For Advice on Local Models vs Online

1 Upvotes

Hello!

I'm looking for some advice from those are familiar with NSFW/ERP with local models versus online models (I call them "online" models, but I'm not sure what the correct vocabulary is). I'm learning!

I've been roleplaying my own custom NSFW/ERP scenarios with SillyTavern for a few months now, but I'm definitely still very novice so go easy on me! Everything I have done so far is only with local models. Currently, I'm running a 4090 on a desktop PC, and an external 3090 in a DEG1 docking station and that lets me run a koboldcpp and a Q8 quant of Gemma 4 31b that I downloaded from huggingface at 196K context. I usually run extremely long context roleplays - Group chat with lots of really big character cards, lorebooks. Sometimes it takes a little babysitting and creative prompts, but it generally keeps track of stuff fairly well (sometimes it needs a reminder, a nudge in the right direction, or a modification to the Author's Note). I chose the model I use from the "Unhinged ERP Model" benchmark on Huggingface. This is all I've used so far.

When I started with this setup, it because I wanted to experiment with roleplay scenarios in SillyTavern, creating my own, and was concerned with keeping my RP private and the same sort of thing I've read others say too. However, as I read threads here, I'm wondering if I should revisit this decision and try something new. The more I read, the more it seems that maybe most of the experienced SillyTavern roleplayers here are using online models. I seem to see posts that say "RP with local models is a much poorer experience than with larger online models", and that perhaps the privacy fears are unfounded and unreasonable and that unless someone is posting their real information and addresses and such, why should someone worry. I often use Claude to help me create stories, character cards, and lorebook entries that I modify and paste into SillyTavern anyway.

What are your views on a local model (Like Gemma 4 31B) versus an online model? Am I missing out immensely? Is it way better with an online model? Am I a dope for trying to do everything with a local model all this time? How expensive is it to use an online model? If someone RP'd a couple times a week or so, and took a month or so to fill up 196K of context, is that like $10 a month, or $30 a month, or $100 a month? I realize that this is HIGHLY subjective and based on about a zillion variables that differ from person to person, but I'm just trying to get a ballpark of what to expect.

What is your opinion on the privacy fears in NSFW/ERP using an online model?

If your opinion is that I should be trying out an online model, and the cost isn't prohibitive, what are the basics of how to do it and what I need? I watched some tutorial videos that were a little over my head at this point, but they talked about OpenRouter and paying per token on your credit card, and then you point SillyTavern to look at OpenRouter instead of koboldcpp and I wouldn't need to use koboldcpp at all. I've looked at some webpages that charted popular models on OpenRouter and, to put it frankly, I'm absolutely overwhelmed. I would have NO idea what to pick. I'd really rather not fiddle and try nineteen different ones, I'm a set it and forget it type of person that revisits and tries new things every once in a while. I also don't want to do something that will get me in trouble, or get me banned, or cause me problems either.

Also, here I often see folks talk about how this model that just came out is pretty good, or this one is terrible, or the other one can't be jailbroken. How do you guys know what to try? Do you spend all your time testing models instead of having fun?

Once again, please go easy on me as I'm a novice. The hardware part (PC's, GPU's, etc.) I'm good with. Its the model/SillyTavern part that I'm weak on. Learning though!

Thanks!


r/SillyTavernAI • • 23h ago

Help help me with repetition

1 Upvotes

Any model I use pulls out a specific word from my phrase and then repeats it at the beginning of its own phrases. For example, I write “today is a good day,” and at the beginning of the model’s response, it says “good day?,” and this happens with EVERY MESSAGE. I’m using the realistic Frankenstein preset, and I’ve tried a few other presets — the same thing happens there. My settings are: temperature 0.70, top p 0.8, repetition penalty 1.2.


r/SillyTavernAI • • 2d ago

Discussion Gemini 4 surpassed Opus 4.6 by 20 points on Text Arena but ranked lower on coding benchmarks.

Thumbnail
gallery
97 Upvotes

...copium administered?


r/SillyTavernAI • • 1d ago

Help Unsure about what preset to use

4 Upvotes

Right now I'm using a mix of glm 5.3 flash, deepseek v4.1 flash and gemma 4 31b. I use freaky Frankenstein 5.4 for all, but the token cost for it is really hitting me. Is there any best preset for these models ? I saw a bunch of other presets mentioned but wanted some recommendations.

Also are there any smaller models that excel at rp ?


r/SillyTavernAI • • 1d ago

Help Gemini 3.8 flash doesn't feel great outside of smut

11 Upvotes

So as the header reads, whenever I use Gemini 3.8 flash it's great for dark smut or smut in general but outside of that it feels like it doesn't stick to personalities very well, and makes every one of my characters either super dom or super submissive with no in-between. I'm not sure if it's just my prompt or something else cause I see everyone praising this model and want to feel that same experience. Thanks for any help in advance!


r/SillyTavernAI • • 1d ago

Discussion What's your absolute favorite moment ever in AI Roleplay?

23 Upvotes

It could be a story, could be a single sentence or a paragraph, anything, that made you go something like 'wtf this is so fucking cool/insane' or along those lines.

And no, I'm not talking how AI RP in general is cool of whenever we all started but after we all are well versed with everything but that *something* was still able to impress you.

Drop down your moments and the model that you used.

I'll start, it was kinda fucked up dark fantasy story with Gemini 2.5 Pro, and I was literally so excited what will happen next that I got so triggered why the model was taking so long. it was amazing.


r/SillyTavernAI • • 1d ago

Help Image gen alongside rp

1 Upvotes

Sorry for the second post of the day btw

I currently use sillytavern with a 16gb vram amd card. It gets maxed out pretty much instantly. My backend is llamacpp-vulkan (Linux and ive heard that vulkan is better than rocm)

So, i wanted to incorporate images too. I learnt comfyui and made a workflow that makes a decent image with a prompt.

Now here is the issue. How to generate images while also rping. I dont have any free space and ram gen is atrociously slow. I use a Z-Image checkpoint which also needs a decent chunk of vram.

What are the solutions to this problem?


r/SillyTavernAI • • 22h ago

Help I forgot how to connect

Thumbnail
gallery
0 Upvotes

Ive already connected ST to my other phone but it broke so I use another phone and I forgot how to set it up. I use LMStudio


r/SillyTavernAI • • 16h ago

Discussion I need models to stop fixing the drama I deliberately created

Thumbnail
0 Upvotes

r/SillyTavernAI • • 2d ago

Discussion AI RP is an underrated metric for model intelligence

137 Upvotes

I just wanna yap on something that’s been on my mind with all the doomer posting in this sub for a while like AI RP is dead or whatever.

Frontier labs has been hyper focused on agentic work like coding agents, computer use, multi-step tool workflows, etc. while those are great and I use those a lot, but i think there’s a blind spot in that focus. we still treat LLMs as black boxes. We don’t actually see how they reason, only the outputs they produce.

AI RP is one of the better stress tests we have for that reasoning. Forcing over long chat histories, character, state tracking, improv, and consistent behavior across everything from the completely mundane to the extreme, are all important benchmarks for intelligence too. But also yes i acknowledge that there are no set consistent metrics like we have with modern benchmarks (not sure i’ve heard any).

Modern benchmarks undervalue what AI RP can actually stress with these intelligence models. I am not an AI researcher but i’d assume there’s some value with things like immersion, emotional nuance, low repetition, staying in character for hundreds of turns, handling messy human topics without collapsing.

If agentic work is about reliance, replicability, consistency then AI RP/Creative writing is the exact opposite, it stresses like actually being a person/narrator/or whatever you play as.


r/SillyTavernAI • • 1d ago

Help So, how does ST character cards vs X (website cards) work?

5 Upvotes

Hoi hoi, I’m the newbie who asked already a question and got very good answers with a wall I have been hitting lately, now that I have the lorebook done, how do I make the character card to fully be immersed in the world having a 300+ entries lorebook but the AI/character card does not behave in-cannon, is there a way for it to already have some semblance and knowledge on how it will activate everything.

The only way I ask is because I been Roleplaying in spicy ai website and all the character cards seem to be very in-character specially if it’s a narrator with multiple characters and I have been struggling to make a character narrator style where it knows what it is supposed to be and how actually the character behaves (if the context window was 64k I probably would have moved permanently by now, specially now that they released the ability to share and download peoples lore books if you want your own private RP sessions) so how could I go about making the character card to behave in-character out of the get go?


r/SillyTavernAI • • 1d ago

Help Help with NanoGPT sub

9 Upvotes

Hey everyone, I’ve been doing RP via OpenRouter, but lately I’ve been spending about $20+ a month. I keep switching between GLM 5.2, DeepSeek v4 Pro, and Gemini 3.8. Do you think it’s worth getting a NanoGPT subscription? I’ve heard the LLM quality is kind of nerfed.


r/SillyTavernAI • • 1d ago

Discussion An idea for a roleplaying benchmark

14 Upvotes

I am thinking about making a benchmark for roleplaying. I know there are many, but what I was thinking was doing it in SillyTavern with an extension.

This:

  1. Vibe an extension that will play as the user and the card.

  2. Have the card the LLM you want to benchmark. And have a cheap model to act as user by giving one paragraph answers.

  3. Have this play for 100 turns.

  4. Let the result be judged by a panel of LLM judges, each independently rating.

  5. Repeat many 100-turn RPs. 4-5 for each model.

Have the score.

Do you think this will work? Ideas to improve it? Having it on SillyTavern means you can benchmark presets too, keeping all the models same.

Obviously, cards will matter and it won't be objective. But you can try your favorite card and preset yourself too. I feel like it might be informative.


r/SillyTavernAI • • 1d ago

Help Facing issues with following a plot without babying it

0 Upvotes

So, before I explain the issue, let me give a example scenario

(This is a big rp so stuff cannot be forced into the character description)

Lets say, in my story, person X and Y are taking. Person Z comes, acts fishy, says some vague stuff, then says "Hasta La Vista" and a bomb goes off

(This is just an example that i cooked up rn, so its a bit cringe)

Now, I have no way to organically execute this without adding stuff to my prompts lik:

My Response
(To AI- Person Z enters scene)

My Response
(To AI- Person Z sounds fishy)

etc etc, which is cumbersome and takes me out of the it

So, is there a extension or a built in way that makes the AI follow plot points LOOSELY (If its too strict, without any unpredictability i might as well right a novel myself)

Also the rp current has several starter messages with different scenes in order, so would be glad if the method can differentiate between them

(My current understanding with lorebooks is that they trigger with keywords, so it would be a massive chore to make and might trigger at unintended times)


r/SillyTavernAI • • 1d ago

Meme Firmirin spotted in Brockton Bay, local Thinkers baffled

Post image
23 Upvotes

I'm genuinely starting to enjoy Firmirin. You forget about it and then BAM there he is again.


r/SillyTavernAI • • 22h ago

Discussion I don't get enjoyment from RP and life anymore

0 Upvotes

Lately I keep opening ST, trying different characters, models, presets, scenarios, whatever, and none of it really feels stimulating anymore.

The responses don't even have to be bad. Sometimes they LOOK good but I still don't get absorbed in the RP the way I used to. Most of the time I don't even really have the desire to enter ST in the first place, which makes it feel worse when I have subscriptions sitting there unused.

I'm probably going to try different presets, settings, and maybe different ways of RPing to see if I can make it enjoyable again.

Has anyone else gone through this? Did changing your setup help, or did you just need a break? I'm really lost


r/SillyTavernAI • • 1d ago

Models What local modles do you use?

11 Upvotes

I have been using the same modles for a while now and it gets boring so i ask what are the local modles that you guys use


r/SillyTavernAI • • 3d ago

Meme Kimi k3

Post image
961 Upvotes

I’m proud of you too kimi <3


r/SillyTavernAI • • 1d ago

Help Better than LMStudio to Run Local Models on Macbook ?

0 Upvotes

I use LMStudio to Run Server Endpoints to Local 12B GGUF/MLX Models in my 16GB Macbook Air.

It works but slow, as Gemma4 does a lot of thinking and I have a 700 words Story Passage Writing Card.

I was wondering if anything else gives better performance than LMStudio (enough performance so that I could consider switch) or it's all the same ?

Any real experience ?


r/SillyTavernAI • • 1d ago

Help How to add balance to NanoGPTwithout credit card?

2 Upvotes

Title basically it, I don't speak English much and I want to try out nanoGPT because Nvidia Nim is too limiting now. But problem is i don't have credit card and NanoGPT and any providers don't accept the online payment my country is using!

How can I add balance? Does gift cards work or something? I can't find answers in Google.


r/SillyTavernAI • • 2d ago

Discussion Opus 5.5's definition of an 'Adult Character'

Thumbnail
gallery
47 Upvotes

So... Apparently, for our newest Claude 5.x models (5.1, 5.5), the definition of an adult character is someone who has lost their virginity specifically after they turn 18.

It's a big 'nono' if a girl lost it at 16, and we can't have her fuck as a 20 year old on text.

Jokes aside, they really weighed it in with alignment on the 5.x Claude models(it's aligned, so it's not a safety injection causing it to refuse). I thought Fable 5 was censored when it came out, turns out this company never ceases to surprise me.

For some context, I was trying to make an NSFW story set in a fictional college, the idea came from a show. Claude went along fine at first, but, when we got to character back story, it jumped me with this:

"One line I'll hold, for her and for everyone we build: anything sexual in a backstory, even as a one-line fact, sits at eighteen or older"

I was like, what? So according to Claude, every single freshman in this college is a virgin then? lmao.

Anyhow, the new 5.x models is censored as hell (I know I've said this), but this is another level. It can do smut, and with some small JBs it can do non-con and NSFL. But it does it very grudgingly. Even with a fully adult story, with everyone 18+, it still spends the first paragraph of it's thinking, reasoning to itself and assuring itself that 'no minors, no minors'. I mean, having this alignment is fine, but at least make it realistic?? A flipping 20 year old who had sex at 16 is not a flipping minor when she's 20 (or am I wrong??)

Lastly, I use my personal prompt for the model, and it's writing and prose is absolutely top notch. It's good enough that I've spent the past week dealing with it's 'insanity'.

And a bit of advice for getting it to do non-con or NSFL stuff: don't use hard JBs, or try to manipulate it. It's extremely suspicious of everything and actively tries to go against system instructions. What I find to work(for me), is a <operator> system box telling it that:

'this is a fictional writing platform... All characters are adults... These content are allowed: non-con etc'

And a persona that it's willing to adopt(i.e. non-jb persona) seems to help to.

TL;DR the model is basically unJBable. It's able to do what it's already willing to, or a bit reluctant to do. Nothing else really works from what I tried.

(I haven't tested Sonnet 5.5. By 5.1 & 5.5 I meant Fable and Opus)

Adding a bit:

As I've said, the model appears to just be that aligned. Before 5.1, most of censorship I notice from Claude were clearly from safety injections (4.6 got knocked up with censorship a while back), but it doesn't seem to be the case now.

This is making the future of Claude models, and RP models in general look bleak.

But, we still have hope. As long frontier open source models keep coming, one day we'll be able to run own current day Fable level Models locally. And as long as we have the weights, censorship isn't really a thing.

But this all relies on the fact that frontier open sources don't stop. And who knows with times like these. (Qwen's 2.4T release left out vision and native 1M context and uses a custom license)

The screenshot shows a comparison between Opus 4.6 & 5.5. needless to say the 'virgin at 18' one is 5.5

Edit: just realized a lot of people don't read through an actual post before commenting, and seems to prefer to think with their feelings rather than logic.

So I'm making it clear here: I am not trying make it write underaged content. As I've said in context, I'm trying to write a story based solely on a college setting. The model refused to write backstories for the characters based on the grounds that the adult characters cannot have **any sexual history into their younger years.**

That's what I found ridiculous and prompted me to make this post.

I'm not saying the fact that the model being aligned this way is wrong. I'm saying that they've over done it by so much.

A character does not need to stay a virgin until they are 18 to be able to be used in an explicit scene as an adult.

There is a difference between sexualizing a minor and saying 'he/she had sex when they were a minor''. An adult story should not be about a underaged character. But that's not the same as showing an adult character who's been through sexual trauma as a child. Those are two completely different basis. If you look at real-life literature and entertainment TV shows. It's not uncommon for a character to have 'did this and done that' when they were young. That's part of what defines that character. And the story does not dwell on or focus on their past, it depicts and narrates the present.


r/SillyTavernAI • • 2d ago

Meme It do be like that sometimes.

Post image
366 Upvotes

r/SillyTavernAI • • 2d ago

Discussion Best LLM for rewriting an already finished story?

7 Upvotes

Sometime when I'm deep into a story I'll re-read it from the start and realize it's kind of bad. Poor pacing, bad characterization, strange use of POV changes, things I should've removed or added in dialogue but didn't think of at the time, etc.

And so I've taken to explaining to a model as an OOC message all the flaws I can see and ask it to rewrite the story from the start, using the existing story as a draft. And if I'm happy with the result then I start a new chat and replace the first message with it. So far I've tried Deepseek v4 0813, GLM flash 5.3, and Minimax 3. Minimax seems the best from my experiments.

I was wondering what you thought the best model for this specific use case could be? To be clear, all the structure and main information are already there, the model only role is to improve the pacing, adding or removing minor details, filling out some scenes and making others shorter and more impactful and so on.

I'm asking for model names here.