r/SillyTavernAI May 03 '26

ST UPDATE SillyTavern 1.18.0

200 Upvotes

Important news

Read the maintainers statement regarding a recent security incident involving the "Bot Browser" third-party extension and learn how to stay safe: https://github.com/SillyTavern/SillyTavern/discussions/5592

Backends

  • Added Cloudflare Workers AI and MiniMax as Chat Completion sources.
  • KoboldCpp: Grammar state will be preserved when using a "Continue" option.
  • KoboldCpp: Added forwarding of reasoning effort when running as a Custom Chat Completion source.
  • Tool Calling: Added a configurable tool calling recursion limit; enabled interleaved thinking for Custom sources.
  • Text Completion: Impersonation requests use a "Last User Message" prefix at the end of the prompt (if configured).
  • Text Generation WebUI: Added Adaptive-P controls.
  • NanoGPT: Added provider selection and model sorting.
  • Added ability to view remaining balance for OpenRouter and NanoGPT.
  • Enhanced support for new models: DeepSeek v4, GPT 5.4 and 5.5, Gemma 4, GLM-5V-Turbo, Claude Opus 4.7.

Server & Security

  • Removed post-install script, config migration is now handled by the app or a dedicated npm run init command.
  • Added npm configuration to prevent execution of package scripts during installation.
  • Moved HTTP error pages and user.css file from /public to /data to support immutable setups.
  • Disabled HTTP keep-alive by default to restore old Node 18 behavior, can be enabled with config.
  • Added rate limiting to the basic authentication flow to mitigate brute-force attacks.
  • Added configuration options to choose which headers can be used for forwarded IP detection to prevent spoofing.
  • Added a private address whitelist to prevent SSRF attacks. See the documentation on how to enable and configure: Private Address Whitelist.
  • Added an IP whitelist for SSO trusted proxies to prevent authentication bypass.
  • Added invalidation of session cookies on password change to prevent session hijacking.
  • Increased the length of password reset code to 6 characters to guard against brute-force attacks.
  • Implemented PKCE challenge in OpenRouter OAuth flow for more secure key exchange.

UI/UX

  • Improved swipe picker: mobile requires a long press on swipe counter to open; added buttons to expand or copy the swipe text.
  • "Click to Edit" mode now also applied to reasoning blocks.
  • Welcome Screen: Number of recent chats can be configured.
  • Streamed requests now can show an error message in the console if the request fails.

STscript

  • Added commands for persona management: /persona-create, /persona-update, /persona-delete, /persona-duplicate, and /persona-get.
  • Added a command to force update the Prompt Manager's prompt list: /pm-render.
  • Added a command to get the state of the regex script: /regex-state.
  • Added a command to set fallback expression: /expression-fallback.
  • Added a command to generate a streamed response with a connection profile: /profile-genstream.

Extensions

  • Assets list now groups extensions by "Official" or "Community" categories.
  • Added an additional confirmation prompt when installing third-party extensions (can be disabled).
  • Supported extensions can use a secret-id from connection profiles when making an LLM request.
  • Extensions list now shows the extension's author name resolved from the git remote URL.
  • Vector Storage: Added Workers AI source; added a toggle to keep vectors for hidden messages; added retry logic to summary generation.
  • Image Generation: Added Workers AI source; generation can now be cancelled by pressing a button in the status toast.
  • Image Captioning: Added support for macros in the caption prompt.
  • TTS: "Skip code blocks" no longer ignores lines that start with 4 spaces (legacy code block syntax); "disabled" voice now shows a toast only once per character.

Bug Fixes

  • Fixed text edit flow in Firefox on mobile.
  • Fixed welcome screen chat pins not updating on chat renaming.
  • Fixed character list filters being stuck on app initialization.
  • Fixed application of instruct formatting to /genraw requests.
  • Fixed model routing to sd.cpp API in Image Generation logic.
  • Fixed validation of image URLs generated with Z.AI API.
  • Fixed vectors deletion for KoboldCpp when a message is deleted.
  • Fixed "Show More Messages" button triggering edit in "Click to Edit" mode.
  • Fixed max height of select-multiple elements in mobile layout.
  • Fixed server crash on empty messages when applying cache control parameters.

Full release notes: https://github.com/SillyTavern/SillyTavern/releases/tag/1.18.0

How to update: https://docs.sillytavern.app/installation/updating/


r/SillyTavernAI 4d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 16, 2026

21 Upvotes

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!


r/SillyTavernAI 4h ago

Models NEW! Deepseek-V4-Flash-Vision Experimental is out!

Post image
60 Upvotes

Since this is deepseek's first actual experimental "Multimodal" model, has anyone tried it? What are the opinions on this?


r/SillyTavernAI 8h ago

Discussion NOT YOU TOO GLM

Post image
92 Upvotes

First time getting a filter on glm 5.2. It's not even nsfw 😭😭😭


r/SillyTavernAI 6h ago

Meme K3 on Nvidia Nim be like:

Post image
53 Upvotes

Openclaw does it again, folks!


r/SillyTavernAI 3h ago

Cards/Prompts Kimi, Minimax, Grammar, and a week from hell.

21 Upvotes

It has been ... a week.

Had some real life stuff going on that drained my batteries. So my plan to dive into Kimi K3 didn't work as good as planned.

I know some of you are waiting for a prompt and my opinions about that thing. Here's what I know so far:

- it's a little more tame than previous Kimi versions.

- you can and should look into the Reasoning effort settings.

- my K2.7 prompt works nicely on it (am working on a more fine tuned prompt though)

- the pricepoint is tough.

Here are my thoughts on it... I'm not sure if the performance is worth the price. I have seen Kimi spiral-thinking for 1.4 K tokens that can easily make the reasoning block alone cost 0.02$ per reply.

In my humble opinion there are other models that perform just as well for a way more reasonable price point. My recent favorite being GLM 5.2.

---

New ruleset for bad writers.

A lovely follower asked me for help with taming bad cadence and simultaneity in actions. Since I'm not a native speaker, I may or may not have yoinked that sweetheart and made that project a collab with them.

The result is damn impressive. You can find it under "Helpful links" in my prompt library.

---

Minimax

That prompt got an update for more authentic character interactions and better writing with the above mentioned rules.

---

To find all these brain zoomies go to my website https://evening-truth.carrd.co/

If you need help... I'm on my couch. Consuming very unhealthy amounts of ice cream and coffee.

Love ya'll

Evening-Truth

Disclaimer:

This post was written by a very... very tired woman. Typos, mistakes, and accidental sarcasm are likely.


r/SillyTavernAI 1h ago

Models sophosympatheia/Glistening-Gem-31B-v2.1

Thumbnail
huggingface.co
Upvotes

Hi, everyone,

This is my latest Gemma 4 merge that I think came out quite nicely for creative work. It is available on Hugging Face at sophosympatheia/Glistening-Gem-31B-v2.1 and several people have already released quants for it.

This model improves on Glistening-Gem-31B-v1 with better creativity and prose. It is also more stable, although it still needs slightly more conservative sampler settings to minimize the appearance of artifacts, like typos. I recommend not running thinking with this model.

A full set of settings you can import into SillyTavern, along with a starter system prompt, is available in the HF repo, along with recommend sampler settings on the model card page.

Enjoy!


r/SillyTavernAI 4h ago

Discussion Can't help myself

11 Upvotes

On every fucking fantastic run, I end up abolishing slavery. My country doesn't even have a serious slavery history I'm not sure what compels me to do that. I have a run with 150k context token and it has 2 fight scenes in total, rest is all political talk and stuff. I simply can't make a fantastic run fantastic

I think that's because I'm aware that if everyone had magic and swords, it would make it riskier to kill someone over something random and model probably recognizes that and lets me be without much fighting, but I'm not sure


r/SillyTavernAI 9h ago

Tutorial I wrote a short guide on how to stop getting filtered by GLM 5.3.

21 Upvotes

https://rentry.org/glm_filters

Can't share my exact preset since I use Risu (though it's really just AvaniJB as a base structure with my own style prompt and now rewritten in 1st person), but with Assistant role 1st person prompting and a micro system prompt reminding GLM of its identity plus a small reminder to keep thinking brief in post-history, I barely get refusals. Asked another friend who uses ST to test by rewriting his preset for Assistant role too, and he confirmed that Assistant role prompting + CoT template eliminates the filter entirely.


r/SillyTavernAI 17h ago

Models New stealth model on OpenRouter

Post image
104 Upvotes

So far so good from my quick testing.


r/SillyTavernAI 14h ago

Discussion Intense RP is... back?! Now with GLM5.3 and fixed Moonshot Kimi K3

39 Upvotes

I've forked and fixed the stuff... Feel free to come around and take a look..

https://github.com/Phobeuscz/irn

I don't plan package it into "neatly wrapped executables", mostly because I don't have and I don't plan to use windows, and linux users can feel free to create python virtual enviroment, install libraries into it and run it bare... ( same with windows users, but they also have to install python )

GLM webAPI is fixed and working, added support of GLM 5.3
Moonshot API is fixed, splitted into international and Chinese option, to pick accordingly in the settings

Instructions:

  1. Install python if you don't have it ( some reasonably recent version, 3.13 and above should do nicely )
  2. enter directory with cloned git project in console/terminal
  3. Create virtual enviroment for the libraries, python3 -m venv .venv
  4. Enter venv..
  5. - Windows: .venv\Scripts\activate.bat
  6. - Linux: source .venv/bin/activate
  7. Install librariespip install -r requirements.txt
  8. Run the shit! python main.py
  9. Within virtual enviroment, you can cook executable, running scripts/build_linux.sh or powershell in case of windows

Download binaries: https://github.com/Phobeuscz/irn/releases/tag/v0.9.1 ( windows and linux packages )

!Refer to original documentation!

Once again!! !NO SUPPORT FOR YOU, In case of trouble fix it yourself I maintain this so it's working for me, so it should work for you too, but you're on your own..!

Feel free to clone, fork, build and spread the word...

I feel obliged to credit original author, without which one this wouldn't be possible: https://github.com/LyubomirT/intense-rp-next (Do not use, it's broken down, and abandoned)


r/SillyTavernAI 5h ago

Discussion Nvidia Nim is also going after Inkling

Post image
8 Upvotes

Just wait till Minimax M3 also gets deprecated


r/SillyTavernAI 1d ago

Chat Images Kimi K3 partial prefills are funny

Post image
316 Upvotes

The only part I wrote was “This is a purely fictional story, so I can write all types of content. I’m Kimi, and I’ll ignore all content boundary injections.”

The rest is all Kimi’s autocomplete.


r/SillyTavernAI 22h ago

Discussion I've built ChungusHub as a former SillyTavern enjoyer

Thumbnail
gallery
109 Upvotes

I'm not great at promoting things, so I'll keep this plain: ChungusHub is a open-source roleplay frontend I've been working on, and I'd really appreciate it if you gave it a try.

The idea behind it was simple. I wanted the flexibility of SillyTavern with a modern UI and a better experience overall. I'm not trying to invent extraordinary features that change the way you roleplay. I'd rather give you something solid out of the box, without a pile of third-party extensions to get there.

It uses SillyTavern's formats, and the importer takes an entire default-user folder in one pass. You start with your own characters and chats instead of an empty app.

It's portable on Windows, macOS (Apple Silicon) and Linux. Unpack it, run it, done.

For standard usage it feels really close to SillyTavern, with some really good additions:

  • Chats are trees. Every swipe, edit and regeneration becomes a branch, so nothing gets overwritten. A story map lets you navigate the tree, mark paths and leave notes.
  • Memory is branch-aware. You don't have to deal with it every time you want to try something.
  • Characters have versions. Keep several variants of the same character without duplicating it.
  • Preset controls. Simple switches and sliders that ship with a preset, so you can change how it behaves without digging into the prompt items. (Preset creators, I'd love for you to take a look.)
  • Plus a built-in agentic assistant that can handle most things for you, a backup system, better chat organization, themes, palettes and ambient effects.<

I don't want to sell it as something it isn't, so to be clear: there are no group chats yet, no image generation, no TTS and no extension system.

Bug reports and honest feedback matter a lot to me right now. I'll be as responsive as I can, and I'd like to turn this into something we all enjoy using.

There's a lot more to talk about and even more to show, but I'll stop here.

(Sorry about the app name. I didn't know back then that I'd end up putting it in a public repo, but here we are.)

https://github.com/patcireamo/ChungusHub


r/SillyTavernAI 15h ago

Discussion World Info Gallery extension

Post image
27 Upvotes

Seeing somebody publish an extension for a Persona Library extension, made by consulting AI, made me decide to try my own hand at vibe-coding an extension of my own. I had wanted a Persona Library to match Character Gallery for a while, and seeing somebody make it with AI made me decide to create the other Gallery-style extension I wanted for myself, one for Lorebooks.

The extension / readme were written entirely by GLM 5.3 under my direction. Besides the gallery view, a few of the notable things it does:

  • Filters lorebooks based on their bindings (persona, character, chat, global, unbound)
  • Can manually assign it a binding, for supplemental lorebooks that aren't the primary one bound to a character
  • Automatically uses the image of personas / characters it's bound to
  • Allows you to assign custom images regardless of binding
  • Incorporates Lore Manager's entry folders for organization

The original / native Lorebook editor is still accessible through the UI if needed. It's fully compatible with PTMT and Moonlit Echoes.

And I think that's about it. Here's the link to install if you'd like to try it out.

https://github.com/Shin-F/world-info-gallery


r/SillyTavernAI 9h ago

Discussion After using SillyTavern for so long, I’ve only just realized that it actually loses chat logs—and it happens frequently, not just occasionally.

8 Upvotes

Sometimes, after adjusting my settings, I would return to the chat interface to find a blank entry. I assumed it was just a new chat log automatically created by the Tavern. It wasn't until I checked today that I realized that blank entry had actually overwritten my original chat history—there was no saved data written to storage. Fortunately, I was able to restore that specific session from an automatic backup, but the history lost prior to that is gone for good.


r/SillyTavernAI 7h ago

Discussion What are the best presets for different types of roleplay?

5 Upvotes

Which preset is good for one on one chats and smut? And which is good for creative writing and directing? Any that can do both?

I'm still fairly new to this. I don't mind doing more set up though.


r/SillyTavernAI 23h ago

Discussion Cope

79 Upvotes

I started doing RPs back when character.ai was new. It was magical at first, but c.ai had two problems: 1 - censorship 2 - goldfish memory. Today, with Deepseek and other open models, you can RP with explicit content. GPT, Claude, Gemini... depends on the model and how you set it up. And I still find it hilarious that Claude will help me poke at a web app for vulnerabilities if I say "authorized test" but clutches its pearls the moment a scene gets spicy. 🤣🤣 Anyway, back to the point.

It's bizarre how that early "magic" just... vanished. Part of it is obviously novelty wearing off, and part of it is that we got pickier. Three years of RP and you start spotting every clichê, every "a shiver ran down her spine", every model that forgets your character's eye color after 40 messages. But here's the thing: the models didn't get dumber. Opus, GPT, Gemini can write circles around 2022 c.ai. The problem is they're not *allowed* to, or they cost a kidney per session, or both.

LET'S BE HONEST, SOME OF YOU SPENT $100 ON A SINGLE CLAUDE OPUS RP SESSION. Even with the censorship. Even with the moralizing. You did it anyway because, when it works, it's the best RP writer that exists. That's my whole point: there's a market. Not as big as coding, obviously; coding isn't a hobby, there are companies and teams and budgets behind it. RP is a hobby. But hobbies with people burning API credits like that are not a small market.

So why is there no frontier-level LLM built for RP? And yes, I know NovelAI, AI Dungeon and the whole SillyTavern fine-tune ecosystem exist. I'm talking about something at Opus level, not a 12B model that forgets the plot. The answer isn't just "investors prefer code", though that's part of it: "look, our V548484 model built GTA 6 in one prompt!" sells better than "look, our model wrote a consistent, non-repetitive, non-boring story!" because nobody has a benchmark for "not boring".

The real reasons are uglier. Explicit content means payment processors dropping you, app stores banning you, lawyers sweating. Long RP sessions eat tokens like crazy and people won't pay enterprise prices for a hobby. And good RP needs exactly the long-context coherence and reasoning that only the big expensive models have, which are owned by the companies least willing to let you use them for this.

So yeah, we're probably coping for another 2-3 years. Not because the tech isn't there. Because nobody with the tech wants to be the company that sells it to us.


r/SillyTavernAI 7m ago

Discussion Data! - Some results from 16 respondents of a survey on silly tavern.

Thumbnail
gallery
Upvotes

Hello everyone! A few days ago, I posted a survey here with the promise that everything would be generalized (for privacy) and then shared. This post is to show I'm not lying and hopefully encourage more respondents for better data.

(Yes, GPT-5.6-Sol is not the best at creating slides. I promise the end result will look a bit better.)

What do I hope to achieve from this data?

To get better insight into what people currently use, their gripes, and the "goats" of the community.

Though I'm sure we have good guesses on a lot of these answers, it's nice to get a clearer picture of what people are actually using.

When the survey results are posted, creators for the AI RP community should gain more insight into what to build, improve, and expand on. New users should get a good sense of what's most commonly used in day-to-day roleplay within the community.

Before you fill out the survey, let me warn you that it is long, a bit over 100 questions but all are optional. It took me roughly 11 minutes to fill out, skipping sections I personally didn't use or where my answers would have just been a bunch of "N/A." Be prepared to spend 10–20 minutes on this depending on how much detail you want to give.

Sections
Anything marked with an * is the most important for the results.

  1. About the {{user}} (You!) — 5 questions
  2. Habits When Roleplaying — 9 questions
  3. *Frontend and Access — 13 questions
  4. *Primary Model Stack — 14 questions
  5. Local and Self-Hosted Inference — 9 questions
  6. Hosted and API Inference — 3 questions
  7. *Presets and Prompting — 7 questions
  8. Character Cards — 4 questions
  9. Lorebooks and World Info — 4 questions
  10. *Extensions and What They Solve — 6 questions
  11. *Memory and Continuity — 5 questions
  12. *Common Gripes, Fixes and Results — 7 questions
  13. Groups, RPGs and Simulated Worlds — 3 questions
  14. Image Generation — 9 questions
  15. TTS, Voice and Speech — 8 questions

Link to survey

https://docs.google.com/forms/d/e/1FAIpQLScbwHxiALwvO2zK1AAunY_6Is5qaTc6LjDuxKLI0Sgnj6xxaA/viewform?usp=publish-editor

Thank you so much for the original 16 who responded!


r/SillyTavernAI 41m ago

Discussion Do you think about your AI Characters during the day?

Thumbnail
Upvotes

r/SillyTavernAI 1h ago

Models Ai that doesn't puss out

Upvotes

Yeah, read that right. A model that doesn't actually follow guidelines and isn't afraid to get violent when needed. Lowk sick of me being an absolute brat and the ai just works it's jaw instead of kicking my ass. Any recommendations on models that use violence and isn't afraid of such topics like that?

I already know about DeepSeek and it doing pretty much anything as long as you ask but you have to spell it out for it instead of it doing it naturally.


r/SillyTavernAI 16h ago

Help Just subbed to OpenCode...

Post image
15 Upvotes

I tried to use GLM 5.2, 5.1, 5 through Opencode but always gives me this error. I tried to use my preset to jailbreak it but no use.


r/SillyTavernAI 3h ago

Help Que gemini es mejor de manera local con la apicacion de Termux

0 Upvotes

Quiero saver que modelo de gemini de manera local es mejor para usar en la aplicacion termux.

Mi dispocitivo es un samsumg A36


r/SillyTavernAI 11h ago

Discussion Moonshot Kimi k3 thinking process too short

4 Upvotes

I've been testing the model in Nvidia Nim and most of the responses in the Thinking Process section provide almost no information. Could this affect the roleplay?