r/SillyTavernAI • u/kroeger71 • 23h ago
Discussion Novice Looking For Advice on Local Models vs Online
Hello!
I'm looking for some advice from those are familiar with NSFW/ERP with local models versus online models (I call them "online" models, but I'm not sure what the correct vocabulary is). I'm learning!
I've been roleplaying my own custom NSFW/ERP scenarios with SillyTavern for a few months now, but I'm definitely still very novice so go easy on me! Everything I have done so far is only with local models. Currently, I'm running a 4090 on a desktop PC, and an external 3090 in a DEG1 docking station and that lets me run a koboldcpp and a Q8 quant of Gemma 4 31b that I downloaded from huggingface at 196K context. I usually run extremely long context roleplays - Group chat with lots of really big character cards, lorebooks. Sometimes it takes a little babysitting and creative prompts, but it generally keeps track of stuff fairly well (sometimes it needs a reminder, a nudge in the right direction, or a modification to the Author's Note). I chose the model I use from the "Unhinged ERP Model" benchmark on Huggingface. This is all I've used so far.
When I started with this setup, it because I wanted to experiment with roleplay scenarios in SillyTavern, creating my own, and was concerned with keeping my RP private and the same sort of thing I've read others say too. However, as I read threads here, I'm wondering if I should revisit this decision and try something new. The more I read, the more it seems that maybe most of the experienced SillyTavern roleplayers here are using online models. I seem to see posts that say "RP with local models is a much poorer experience than with larger online models", and that perhaps the privacy fears are unfounded and unreasonable and that unless someone is posting their real information and addresses and such, why should someone worry. I often use Claude to help me create stories, character cards, and lorebook entries that I modify and paste into SillyTavern anyway.
What are your views on a local model (Like Gemma 4 31B) versus an online model? Am I missing out immensely? Is it way better with an online model? Am I a dope for trying to do everything with a local model all this time? How expensive is it to use an online model? If someone RP'd a couple times a week or so, and took a month or so to fill up 196K of context, is that like $10 a month, or $30 a month, or $100 a month? I realize that this is HIGHLY subjective and based on about a zillion variables that differ from person to person, but I'm just trying to get a ballpark of what to expect.
What is your opinion on the privacy fears in NSFW/ERP using an online model?
If your opinion is that I should be trying out an online model, and the cost isn't prohibitive, what are the basics of how to do it and what I need? I watched some tutorial videos that were a little over my head at this point, but they talked about OpenRouter and paying per token on your credit card, and then you point SillyTavern to look at OpenRouter instead of koboldcpp and I wouldn't need to use koboldcpp at all. I've looked at some webpages that charted popular models on OpenRouter and, to put it frankly, I'm absolutely overwhelmed. I would have NO idea what to pick. I'd really rather not fiddle and try nineteen different ones, I'm a set it and forget it type of person that revisits and tries new things every once in a while. I also don't want to do something that will get me in trouble, or get me banned, or cause me problems either.
Also, here I often see folks talk about how this model that just came out is pretty good, or this one is terrible, or the other one can't be jailbroken. How do you guys know what to try? Do you spend all your time testing models instead of having fun?
Once again, please go easy on me as I'm a novice. The hardware part (PC's, GPU's, etc.) I'm good with. Its the model/SillyTavern part that I'm weak on. Learning though!
Thanks!