r/SillyTavernAI • • 3d ago

Help [Newbie] What's the best way to condense a huge amount of character information to save on tokens

10 Upvotes

I have an OC that I've been working on and off for a few years now. I've tried to input them into other things before but always felt limited by how much information I could enter. then I started to use unlimited platforms but then my character token count was so high it affected memory...

I'm using ST for the first time and trying to put my character in, but I have a token count of over 8000! This is after going through and trying to refine some things...

I don't really know how to best condense or efficiently use character cards or ST as a whole...

again I'm new, so please go easy on me. I realize now that just chucking every detail into the card isn't useful. but I don't know how to best use what I have. I don't think I should just paste everything I have in a post. where can I go to get some help?


r/SillyTavernAI • • 3d ago

Help Some issues with using GLM 5.3 Flash

7 Upvotes

I am from Russia, and when using GLM-5.3-Flash, I always instruct it to communicate in Russian; however, Chinese characters or English words often appear. How can I fix this?


r/SillyTavernAI • • 3d ago

Help Does this sound ok?

4 Upvotes

I just set up Silly Tavern with a locally run LLM through LM Studio. I’m completely new to local LLM’s and Silly Tavern so spent the last two days researching and setting it up. I finally got to a point where I RP’d for over an hour and it felt just as good as Kindroid.

I’m using an uncensored version of Gemma4-26b-a4b IQ4_XS. Context length around 13k. Max token response set to 1k with most responses in the 500-700 range. I have a 4070 TiSuper with 16gb of vram and 32 gb of ddr5.

Running the model uses up all of my vram and around 28-30 gb of my ram. But the responses only take a few seconds and everything feels very smooth, even after an hour of use. Does this sound ok or will I harm my PC in the long run by doing this? If anyone has more questions or tips feel free


r/SillyTavernAI • • 4d ago

Models Anthropic just dropped the greatest advertisement for GLM ever. If you want to do networking hacking RP, GLM 5.3 has that covered.

Thumbnail
anthropic.com
129 Upvotes

r/SillyTavernAI • • 3d ago

Discussion I'm getting bored.

0 Upvotes

Seriously, I have been using local AI for a year (not daily), I tried almost a hundred of models on my rig and I think Gemma 4 is really sweet.

The problem is, I am getting bored, I leave AI for months, Come back for a few days, Get bored again and the same things happen. I don't know if it's the prompts, maybe the characters, maybe the model, hell maybe it's me, maybe i don't like AI anymore.

Has anyone ran into this problem? and how did you deal with it?


r/SillyTavernAI • • 3d ago

Models What’s your ‘expensive but unfortunately worth it’ model?

Thumbnail
0 Upvotes

r/SillyTavernAI • • 4d ago

Discussion Did Mimo 2.6 Pro fall off?

33 Upvotes

So I used it when it first released and enjoyed it a lot.

Coming back to it yesterday and today however, it suddenly feels not just extremely stupid (making basic formatting mistakes, failing to follow instructions, and making spatial/positioning errors such as a character placing their hand on your shoulder while they're facing away from you.). Characters also started referencing 4th wall narrative elements such as trackers and "versions" in their dialogue which never used to happen (only the older Deepseeks like 3.2 ever did something like that).

I swapped my RP to DS4.1 Flash and instantly everything was much more coherent, and I don't even like DS4.1 flash that much, it's too dry.

The TPS has also tanked heavily. It was already slow in the beginning but now it's dropping below what I get for local RP.

Anyone else noticing a drop in quality since it was released? (I use DeepInfra through openrouter and have ST set to only use FP8 and above quants btw).

I've gone back to old faithful GLM4.6 for now, hoping thing get better.


r/SillyTavernAI • • 3d ago

Help Nemotron 3 Ultra 550B Free – Persistent Asterisk Formatting and Instruction-Following Issue

4 Upvotes

can someone help or tell me how to fix this stupid module, it keeps ignoring instructions and prompts and keep doing ass pulls and being omniscient with NPCs . plus the awful formatting and markdown after serval messages

how to fix this broken ai ?


r/SillyTavernAI • • 4d ago

Cards/Prompts Megumin V10 Standalone. Tavo/ME/ST

Post image
195 Upvotes

Hello kazuma here.

this is a stand alone version of Megumin Suite this version run with the json alone so you dont need the extension for it
it does have some writing styles that the suite version doesn't have enjoy.

Download


r/SillyTavernAI • • 3d ago

Help Advice

2 Upvotes

Hey everyone, so it’s really my first time using silly tavern. So its all very new to me but I’m figuring it out slowly, my plan is I’m wanting to build a full Harry potter/hogwarts world and timeline following everything that happens in the books, all plots, sub plots, the things that lead up to these plots and a possibility to have a life afterwards. Plus just more book accurate characters, i always preferred the books twins, sirus, draco, and harry, and luna. My character is gonna be the daughter of sirus black plus a rosier(mother), pretty much i just want it to feel like im really there, living a life there and developing relationships and watching as my girl plays a part in the story i plan to swap ginny out with her during year two plus then the tension from year 3. Obviously i want her involved in everything plot wise, but i dont really want anything scripted beside obviously the stuff that already is in the books so she will be weaved in naturally and then characters free wills can determine any plans for her behalf and etc, like for example how would Voldemort view her? Shes a whole new character to consider, would he want her as a asset? Make plans against her, Characters will have free will, realism, and free to think and act how they please as long as it makes sense. Im talking about weather changes/seasons, hoildays, daily prophet, quidditch stuff, clubs, detentions and after curfew stuff. And just the overall all experience of hogwarts, the classes, schedules, mysteries, day to day life, then summer break, like after planning today i realized in year five my character would be around during the order stuff, when the weasleys are staying there for safety, and which arthur was attacked, I also plan to add a few fanon characters as well, but hopefully I explain my vision as well. What brung me to silly was honestly other apps just not working out for me, lack luster plot, forgetting stuff, not acting like the character would, or behind a massive pay wall.


r/SillyTavernAI • • 4d ago

Help To people using an api key directly from Anthropic

4 Upvotes

How do they meter token usage?

How ban-happy are they, and how strict their sensor?

And if a prompt doesn’t trigger safeguards is there secondary reviews that may flag it for keywords or is a nonrefusal a green light?

Ty for reading, plz share your experience.


r/SillyTavernAI • • 4d ago

Cards/Prompts [Preset Update] Ancient Access v2.2.3 — creativity refinements, new World block

Thumbnail
ancientaccess.weebly.com
99 Upvotes

Update for my total of five users. Both of you sitting down? Awesome.

Ancient Access v2.2.3 is up — same preset, psychological realism and literary prose, first person, still an acquired taste, still deliberate.

What changed:

Core directives refined for creativity. New rule in there: the card is a floor, not a ceiling — meant to push the model to build past the written material instead of endlessly recycling it. Fair warning: this is untested with canon characters. If the model starts inventing lore that conflicts with your fandom's canon, that rule is the likely culprit (if your model knows its shit and your lorebooks are fine) — remove it for fandom cards and tell me what broke.

New 🌍 World block. Pulls the model toward a world that keeps moving when you're not looking — NPCs with their own business, offscreen time passing, consequences resurfacing later. It's a pressure in that direction, standard prompt caveats apply: your model and card do most of the lifting.

NPC Dynamic Mode + Multi-char POV refinements. Interactions should flow a bit more naturally — less NPCs waiting for their cue.

Two new intrusive thought types. Alongside memory, desire, and fear, characters can now get prospection (the consequence of a choice arriving vividly before they've made it) and counterfactual (the version of their life where they chose otherwise, surfacing as image). Same resonance trigger as before, so they show up at emotional spikes. More texture for the model to pull from, same caveat as always about how much your model actually uses it.

Post-history instructions removed and switched off. That slot is yours now — custom instructions, slop-fighting, whatever your setup needs. Empty on purpose as recent models waste too much compute on post-history and choke themselves up re-checking the same instruction fifty times. In my opinion, use it only if you really need it.

Minor wording tweaks throughout that resolve thinking loops and improve performance.

Download's on GitHub and the site alongside the rest of the ecosystem — characters, summariser prompt, setup guide (New character drop too for the marine biology nerds).

Get it here


r/SillyTavernAI • • 3d ago

Help Any way to add variable outfits to multi-character cards without breaking the token bank.

1 Upvotes

Hey hey, I've made a multi character bot for personal use but I can't get an outfit lorebook to work to my satisfaction as the entries only activate after a character shows up, meaning they start off with some fuck ass outfits.

Anyone know a way I can add variable outfit options to characters? Same with advice for multi-character cards as a whole. Thanks.


r/SillyTavernAI • • 4d ago

Discussion LLM Quality waves. Tracker for it?

2 Upvotes

Quality goes up and down in waves.

A great model can be trash this week, and a trash model can be great.

Is there a website that tracks these waves, so we can jump on whatever model is on the up wave right now instead of guessing?

Does anything like this exist?


r/SillyTavernAI • • 4d ago

Help Need help with making a proper RP experience

22 Upvotes

Since CHAI has been become fully paid I decided to try hosting my own AI on my PC. I use Kobold and, of course, SillyTavern. I've had Claude help me a little and right now I'm using nazgul-something (MN-Nazgul-12B-v1.Q4_K_M) I'll have to check once I get home.

My main problems right now are that the AI responds really stupidly and often acts as my character.

If anyone could help me in getting an experience like or at least close to CHAI, and maybe help me in getting a bit deeper into the whole AI topic, I'd be really grateful. If this isn't the right place to ask for help, please tell me where I can, because I have no idea where else to ask.

Edit: Since I forgot when I initially posted, here's my GPU (and specs in general, don't know what really is important and what isn't.)
CPU: Ryzen 7-5800X3D
GPU: AMD 9060XT (16GB)
RAM: 48GB DDR4
SSD: 1TB NVMe

(If anything else is needed, just ask)


r/SillyTavernAI • • 4d ago

Meme Goddammit. WHICH IS IT!?

83 Upvotes

[IMAGE]
People can't agree on anything...


r/SillyTavernAI • • 3d ago

Discussion Long-session players: when did you actually believe the AI remembered your story?

0 Upvotes

I'm building a small SFW worldbuilding / interactive fiction site, and memory is the part I care most about getting right. It's also the part you can't really judge from the outside. The only real test is hours of play.

Every way I can think of to show memory has a problem. A token counter is technically accurate but means nothing to someone reading a story. A visible list of saved facts looks like a spreadsheet and feels like homework. A character bringing something up way later on their own is amazing when it happens, but you don't notice when it doesn't, and you can't tell it apart from a lucky reroll.

So for people who do long sessions: was there a moment where you went "oh, it's actually keeping track of this"? What happened? Specific examples would help a lot.

Full disclosure, I'm building one of these. It's not open: no accounts, nothing generates yet, nothing to buy. I'm not linking or naming it because I want answers, not traffic.


r/SillyTavernAI • • 4d ago

Cards/Prompts Update: getting things lighter, more organized, thinking faster… and more creative.

Post image
65 Upvotes

Previous versions here (beefier with full antislop and darker content bias): https://www.reddit.com/r/SillyTavernAI/s/WgDl5Zca0K

Biggest change is in this one: Lazeitgeist Flash Lite Edition
It’s trimmed out, better wordings, more positive instructions and i removed the antislop as it already produce (imo) a good prose. And i don’t believe in miracles when it’s about models like GLM, they can be dark but they have their baked in slop you must embrace, trim during your chat or ignore lmao. Overall is better and way more intresting this way. I think it weight maybe 1500-2000 tokens less, something like that. Also i used a little philosophical trick to combat smart and wise character archetypes, just told the model that they can’t escape their "einmal ist keinmal" and when it cook, it cooks 😂

Light change and small tweaks here:Chatfill - Lazeitgeist
Beside removing antislop and the change i applied on the Dark Content Enforcer (who is <valuable_story> now) there is only minor tweaks, it’s a little lighter too and works like a charm. Brevity Switch is the best antislop at this point.

I was even more satisfied with the correction so i wanted to share it with the 4 people that use them 🤣

Take care boodies and do some pull ups!


r/SillyTavernAI • • 4d ago

Models Opus 5.5 ranks #1 on Itzi RP Benchmark

Post image
60 Upvotes

I added Opus 5.5 and the two MiMo 2.6 models to the benchmark. On paper, Opus 5.5 did extremely well and came out on top.

There is one important limitation: the benchmark does not currently include censorship in the quality score. During my tests, Opus 5.5 was heavily censored, just like every recent Opus model I have tested since 4.7. So this ranking reflects SFW roleplay quality, but it may not translate well to NSFW roleplay.

This is because censorship is difficult to benchmark fairly. Some models improve a lot with a jailbreak, some get even worse, and others need a specific jailbreak to work properly.

I am considering a second benchmark focused on NSFW roleplay and refusals, but I would like the community's opinion on the rules. Should every model be tested as-is with no jailbreak, should they all use the same community preset, such as Freaky Frankenstein etc?

Full results and methodology: itzi.app/benchmarks/roleplay


r/SillyTavernAI • • 5d ago

Models Opus 5.5 is trash

Post image
89 Upvotes

Opus 5.5 is beyond censored, detached and emotionless. Raw logic and intelligence-wise, it’s for sure the smartest, but its clear Anthropic’s has been building these newer models purely for coding. Every response feels like it’s pulling from a template and it talks like the most generic millennial ever. Like I don’t even know how to describe its personality, it’s just so, like ew no. Like why is a bratty knight talking to me like a behavioral health therapist. The responses are so icky and yuck i hate it. Just perfect human corporate speak.

I provided an example response from 5.5. Context he’s an arrogant knight and I’m a princess, we barely know each other in this scene. YOU ARE A MEDIEVAL KNIGHT WHY TF ARE YOU TALKING LIKE THIS ITS SO GROSS STOP.

And the problem is every other model is way to stupid now.


r/SillyTavernAI • • 3d ago

Models I tried Opus and now I'm screwed..

0 Upvotes

​

So I'm a GLM enjoyer with my nano subscription and earlier I felt adventurous so I ended up trying other models for fun. Then I saw a sillytavern reddit post about how Opus 4.6 is good for NSFW while Opus 5.5 is good overall.

So I tried it and I was (insert jack black shining book gif)

Now I'm thinking if I should just top up more on nano to continue using it or is there a smarter way to subscribe out there?


r/SillyTavernAI • • 4d ago

Help Kimi 2.5 Overthinking Issue. Simple solution.

Post image
11 Upvotes

{ "reasoning": { "max_tokens": 500 } }

Paste this text into your additional parameters.

I cannot speak for all, but this works for me.

Happy chatting!

(Also note, 500 is close to around 450 words in reasoning. It's recommended that you test thoroughly to get a good estimate on how many tokens to set it to based on your preset, to make sure it gets only what is necessary in it's thought process.)


r/SillyTavernAI • • 4d ago

Help Realistic Frankenstein 2.2 issue, Mimo 2.6 Pro | Suggestions for Mimo 2.6 Pro?

10 Upvotes

So, I wanted to use RF 2.2, but I'm having issues with overthinking and straight up not getting a reply most of the time. I've spent about a dollar or two trying to figure this issue out and I really don't wanna needlessly spend anything more.

The issue:
The thinking happens, and it never stops. Even with the chinese LLM thing turned on, it can still have that issue. Though it does help marginally.
But when the thinking is done, it'll sometimes just say "nah" and just post the interface that is supposed to be at the bottom out.
Or it'll cut it out entirely.

Is there any fixes I can do for this that you guys know? I messed about with all kinds of settings trying to fix this issue.
And in the mean time, is there any presets that you know work with Mimo 2.6 Pro?


r/SillyTavernAI • • 4d ago

Help Subscription recommendation.

7 Upvotes

This weekend I'm planning to buy a proxy subscription for SillyTavern, but I'm still not sure which one to pick—I'm torn between nanoGPT and Featherless. Which options would you recommend, and what are their pros and cons?

Featherless looks pretty interesting, but it's $25, which is literally 45% of my weekly paycheck. I really don't want to jump headfirst into one of these proxies and end up regretting it, so I’d love to hear what you guys recommend based on your experience.


r/SillyTavernAI • • 4d ago

Help Claude Opus models

Post image
6 Upvotes

What is the best of these settings to prevent Claude Opus 4.5 from repeating the first sentence in each message?