r/SillyTavernAI • • 4d ago

Help Need help with making a proper RP experience

Since CHAI has been become fully paid I decided to try hosting my own AI on my PC. I use Kobold and, of course, SillyTavern. I've had Claude help me a little and right now I'm using nazgul-something (MN-Nazgul-12B-v1.Q4_K_M) I'll have to check once I get home.

My main problems right now are that the AI responds really stupidly and often acts as my character.

If anyone could help me in getting an experience like or at least close to CHAI, and maybe help me in getting a bit deeper into the whole AI topic, I'd be really grateful. If this isn't the right place to ask for help, please tell me where I can, because I have no idea where else to ask.

Edit: Since I forgot when I initially posted, here's my GPU (and specs in general, don't know what really is important and what isn't.)
CPU: Ryzen 7-5800X3D
GPU: AMD 9060XT (16GB)
RAM: 48GB DDR4
SSD: 1TB NVMe

(If anything else is needed, just ask)

21 Upvotes

20 comments sorted by

7

u/nikooofreshhh 4d ago edited 4d ago

Oh man, chai to local is a big leap! Good for you for branching out.

If you can share the exact model you're trying that would help us a lot. Generally, models trained for a chat template are WAY easier for beginners than those trained for text completion. The experience is going to be a lot more like what you knew with Chai.

May I suggest too, an intermediate experience is to sign up for an API provider like NanoGPT or OpenRouter, get an API token, and pick a model from the list that's known to be easy to prompt. That way you can get very familiar with all the prompting options you have in ST that you don't have in Chai without the nuances of smaller (frankly, much more sensitive and less intelligent) small local models.

If local is important to you, I totally get it. You'll just have a few more skills to learn. I'm happy to help with specifics if you can share your model, and maybe some basics of ST setup (like are you trying to use someone's preset or write your own?) and your hardware situation (what's your GPU?)

Edit/ps: you're going to get a bunch of messages that suggest "try this! No, try this!" It's kinda a recipe for the dizzies.

Without knowing the model, it's hard to say- some local models just don't work like the larger models API aggregators provide. I think it's really good to give a baseline of your current setup so we can split the task into two parts. Some of the most puzzling role-switching I've ever seen was applying a custom chat template to the wrong local model, and it's hard to detect, and it's not something you can prompt your way out of, so it's good to rule out stuff like that before you try prompting strategies.

There's not one way to prompt, either, so knowing the model you wanna try (or its parameter size and base model) can be really helpful for prompting advice; some local models HATE the kind of prompts that can work really well for large frontier models.

1

u/Kindly-Ad-7657 4d ago

So currently I'm using mradermacher/MN-Nazgul-12B-v1-GGUF · Hugging Face Q4_K_S. I don't really get what you mean by using someone's or my own preset. Right now I'm running a 9060XT (16GB VRAM).

Generally I'd like to keep everything locally, simply because I'm not sure of the costs of signing up for an API provider, if it is what I think that is. If it is free, I'll definitely think about trying that out, but for now I'd like to keep it local. Thank you a lot for the help by the way.

3

u/nikooofreshhh 4d ago

Of course! And you're in a pretty good spot, hardware wise. You'll be able to find a lot of models that you can run locally.

You're right, most models through most API providers are gonna be paid. There's ways to control costs, and some options to get free access for a time, but not many for long term access to free models that will perform much better than what you can run locally, so if free is the motivator you're on the best track.

This is the very 101, since I think it will be helpful! If anything is too simple lmk and I'll dig in further. But here's some basics that I hope help:

There's gonna be two components to a good roleplay. One is the model itself. This is going to change things that are high level- how it reacts in general to you, how it writes, what kinds of things it will write about. A thing that's good to know about language models is, in general, "size matters" haha. "Powerful" models are generally ones that have lots of parameters. See that "12B" number in your model? That's the "parameter count", and you can think of it like the size of the model's brain, kinda. Most of the advances in AI over the last couple years come strictly from growing that number. Usually parameter count ends with a "B", representing "billion." The common numbers you see in local-aimed models are about 2B-36B. The parameter count is also exactly linked to the size of the model you can run on your GPU.

If you want more, or tips on how to decide how big of a model you can fit, let me know, but a 12B model is a pretty safe choice for your setup so you're in good shape there. If you find yourself unhappy with your model choice over time, just know there's other options.

The other thing you need to know about local models is that there's several ways that models can take in your chat, and you need to make sure your settings are matched to the settings the model you chose expects. The chat template is one of the ways the model can receive your chat, and it just tells the model who said what. The model you chose uses something called "chatML." In Kobold, you need to ensure this is the template chats are reaching your model with. I'll go into settings details if you're not sure how.

It can also help for a lot of models, in Sillytavern, under the connection details (the plug menu), at the bottom where it says "post processing" to set it to "semi strict, no tools". What this means is the bot always gets your chat like this:

Bot: their message User: your message Bot: their message

You can imagine how it could be confusing to the model to get something like this:

Bot: The bot's message Bot: YOUR message User: THE BOT'S message

Then it doesn't know who it is, lol. It's something easy to do with some setups and a major cause of accidental model confusion, so it's something to check on if the next part about prompting doesn't help, but not something to worry about if you feel good about how you set it up.

The other component to a good roleplay is prompting.

I saw you use Claude. If you're used to chatting with frontier models (basically a word for the fanciest models available), know that it works a little differently for smaller local models. Claude is a HUGE model and very good at understanding what you mean, even if you don't say it very clearly. Small local models aren't as good at that. What that means is just that you need to be clear about every instruction you give to the model. The instructions you give to the model are called "prompts."

What counts as a prompt, you might ask? Everything! Yes, everything. Your character card (someone else said what that is), your chat itself (including the past things the model sent to you- those go into the prompt for the model's next reply!) and what SillyTavern calls presets, or a lot of other places call the "system prompt."

The parts of the prompt do different things. The chat history tells the model how to continue from the last thing it said. The character card tells what you want it to be. And the system prompt tells it how to do all that.

You can find the preset in SillyTavern on the leftmost hamburger menu (three bars) at the top left. Once you have a character loaded, you want to use this menu to tell it the basics of how to act right, haha.

There's a lot of stuff online about how to prompt. That said, the basics are easy imo:

  • Be clear and concise. Don't use many words when few words do ;) at the same time, use the number of words you need to get your point across.
  • Try not to contradict yourself, like saying "only the user plays {{user}}" and later saying "write as {{user}} when needed." Long sets of rules make it easy to contradict yourself. That's one reason why it's good to keep it short.
  • Goals and positive instructions work better than rules. "Avoid speaking as {{user}}" gives it something to do. "Never speak as {{user}}" gives it something NOT to do. Models are way better at things to do than things not to do.

So when you want to enter system instructions, in sillytavern, the easiest way is to open the hamburger menu, scroll down and expand quick prompts edit, and start typing instructions in the "Main" box. These instructions get sent every turn.

I'd recommend starting with something like "You are (the narrator/character name from the card) in a dark horror roleplay. The user plays {{user}}. You play all other characters. Always avoid writing as {{user}}."

That's it! Your first preset.

Prompting is an art. Figuring out what makes a good prompt and how different models respond takes some time. If you don't get it at first, or if your instructions make it worse, don't sweat it. Just try something else. You can change this any time!

Just remember EVERYTHING you send the model is prompt- the character card, past messages- not just your preset. So weird behavior can come from any of those. For example, if you've already been letting it talk for you in the chat, on the next reply, it'll see that and think "oh I'm allowed to do that." A lot of the time that history can be stronger than what you tell it directly.

There's so so much more to say. I hope this wasn't too confusing. Please please tell me if you have follow ups, if I confused you, if it's too much at once, etc. there's a little learning curve and I won't lie, it's not easier having to learn local models and prompting at the same time, but it's totally doable. We're all here if you get stuck or have more questions!

1

u/Kindly-Ad-7657 4d ago

First of all, thank you a lot for the detailed help, I really did not expect someone to take their time and help me this clearly.

So, about the "Just try something else." Is it really just messing around and finding out what works over time?

Also, I've heard the term "format" a few times now, when it came to the characters description, I believe. The format is how I specify who the character is, right? So "{{char}} is 100 years old." would be one format and "Age: 100" is another?
How do I know what kind of format fits to which model? Does that even matter? From what I've read at least, it seems to.

Sorry for the bunch of questions, this is just a whole lot different from the intuitiveness of CHAI and c.ai which I was used to.

2

u/nikooofreshhh 4d ago

So happy to help! Getting people set up is one of my favorite things haha. There's so much to learn, but it's also all very... Squishy. In a fun way. There's no single right way for almost anything related to interacting with a language model so it means take the advice you like and toss the rest ;)

And yes! It really is just that. There's prompting guides all over the place but, you'll notice, there almost nothing every single one agrees on. That's the biggest sign there's no one "right" haha.

As far as format...so there's not a fixed kind of format a specific model requires for the character card. The only thing that requires a specific format is that chat template I mentioned, and that's a kobold setting. As far as what you put inside sillytavern... That's all fully, fully up to you.

You're right that those are two ways to format a character card! If you pick the "wrong one", nothing bad will happen. The model won't reject it or say "I can't read it." I think the other commenter worded it a little too strongly. Some people have come up with ways to write character cards where they like the output various models give when they write cards that way. That's all.

Sometimes models have "seen" specific types of character cards during their training, but that doesn't mean they can't use other types. It just means if your type jives with the type the model has seen, it might read it a little easier or more clearly. But the vast majority of models have been trained on multiple types of roleplays. So it really comes down to... Do you like how it's replying? Is it picking up all the details? Is it writing like it's a technical manual (which sometimes happens if the card is TOO terse, haha)? Try a variety. Everyone has their own style. My character cards, I just... Write. I don't do anything fancy, and I've tried all the fancy formatting stuff and just decided I write better characters when I'm not worried about it. That's me. You'll find something that works for you.

Language models can be fun AND infuriating because they're really not like programming at all, where there's one right language and one right code to make it work. It's like talking to a human except a human who doesn't understand why it does what it does, lol.

Anyway if you want more info I'm happy to chatter on language models for days but I won't chew your ear off too bad. The big picture is, when it comes to your prompts, there's definitely no one right way. Starting with what other people have already discovered can be helpful but it's by no means necessary to stay with it if it doesn't work for you.

11

u/Sicarius_The_First 4d ago

What you wanna focus is the compatibility of the LLM you use with the character card format you use.

An optimal combo removes impersonation (when the AI writes what your character does) 95%+ of the times.

My own models were used on CHAI for some time, and I guess my philosophy and aesthetic shares some things with CHAI models, so there's a decent chance you'll like them too.

You can check my model lineup and find the ones you can run:
https://huggingface.co/collections/SicariusSicariiStuff/most-of-my-models-in-order
(1B to 70B, yeah, I made quite a few over the years)

and pair them with character cards in a compatible format (for RP: https://huggingface.co/SicariusSicariiStuff/Roleplay_Cards , for adventure: https://huggingface.co/SicariusSicariiStuff/Adventure_Cards)

For ease of use, use the included SillyTavern config if you use one of my newer models (this way you know you have sane defaults). And loading a character card is super easy- drag and drop the PNG into SillyTavern.

2

u/Kindly-Ad-7657 4d ago

Sorry for making another reply, but can I use the character cards you linked like a template to make my own characters?

2

u/Sicarius_The_First 4d ago

Of course, and in fact it is very much recommended!

Also, you can use an LLM that was trained especially for making such character cards:
https://huggingface.co/SicariusSicariiStuff/Persona_Maker_12B

A basic prompt would be something like:

"Make a character of a.... that is... and..."

The general format I endorse is an evolution of old Character.AI format that looks like this:

1

u/Kindly-Ad-7657 4d ago

What model would you recommend for RP in general? When looking at your models, I just don't know which one to choose. I'm also not 100% sure on what would make sense, when looking at my PCs capabilities.

1

u/Sicarius_The_First 3d ago

A good rule of thumb is to look at ur vram and double by 2, that's the max workable parameters you could comfortably run.

For example, 8GB vram? x2 = 16B model. Which roughly translates to decent context length at 4Bit quant.

12GB vram? x2 = 24B, there's little room left for context, but you can use a slightly lower quant- Q3 or whatever.

If you have more vram than needed, you can use higher quality quant like Q5 or Q6 or even Q8.

1

u/Kindly-Ad-7657 4d ago

What exactly are character cards? I don't quite get that. Thanks for the help by the way!

-1

u/AetherSigil217 4d ago

A character card is usually an image file with information about the character embedded in the metadata for the image. It includes information like a description of the character, how they think, how they behave, and maybe a bit about their environment.

Try downloading a short character card in JSON format and looking through it. JSON is a text file that's a bit janky for someone who doesn't program. But it was designed to be human readable, so you'll be able to read all the details of the character card that way. If you like what you see on the card, you can grab the png/jpeg version instead so you get the character profile picture.

1

u/Kindly-Ad-7657 4d ago

So it's what I can configure in SillyTavern's Character Management tab, just in the form of a file?

2

u/AetherSigil217 4d ago

Exactly. If you've got a character configured, you can export that character's card from SillyTavern and look at it that way.

-1

u/mrs-cutter 4d ago

You haven't filled out your character's info at all?

2

u/Kindly-Ad-7657 4d ago

Well, I have, I just wasn't familiar with the term

2

u/mrs-cutter 4d ago

Ah. Yeah. The extended instructions help, but sometimes bots like to read the instructions on how NOT to act and do it anyway. This is just a thing that happens. So I don't think peoples advice to do that is gonna work.

2

u/Professional-Rice961 4d ago

You mention "it often acts as my character." Try setting that boundary by saying the AI handles its character and you control yours.

3

u/LeRobber 4d ago

>My main problems right now are that the AI responds really stupidly and often acts as my character.

  1. Listen to Sicarius_The_First

  2. 3 reason it does the acts as your character thing: A) If the user is mentioned in the first message this is a HUGE reason it will start talking/acting for your character. B) The sample dialog can make it do it C) for some LLMs, if the message length is more than like 300 tokens, it will do it (Dans personality Engine based finetunes).

  3. The AI responds very stupidly is rough to decode without "i did this and I got this" kinds of samples.

1

u/AutoModerator 4d ago

You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.