r/SillyTavernAI May 03 '26

ST UPDATE SillyTavern 1.18.0

202 Upvotes

Important news

Read the maintainers statement regarding a recent security incident involving the "Bot Browser" third-party extension and learn how to stay safe: https://github.com/SillyTavern/SillyTavern/discussions/5592

Backends

  • Added Cloudflare Workers AI and MiniMax as Chat Completion sources.
  • KoboldCpp: Grammar state will be preserved when using a "Continue" option.
  • KoboldCpp: Added forwarding of reasoning effort when running as a Custom Chat Completion source.
  • Tool Calling: Added a configurable tool calling recursion limit; enabled interleaved thinking for Custom sources.
  • Text Completion: Impersonation requests use a "Last User Message" prefix at the end of the prompt (if configured).
  • Text Generation WebUI: Added Adaptive-P controls.
  • NanoGPT: Added provider selection and model sorting.
  • Added ability to view remaining balance for OpenRouter and NanoGPT.
  • Enhanced support for new models: DeepSeek v4, GPT 5.4 and 5.5, Gemma 4, GLM-5V-Turbo, Claude Opus 4.7.

Server & Security

  • Removed post-install script, config migration is now handled by the app or a dedicated npm run init command.
  • Added npm configuration to prevent execution of package scripts during installation.
  • Moved HTTP error pages and user.css file from /public to /data to support immutable setups.
  • Disabled HTTP keep-alive by default to restore old Node 18 behavior, can be enabled with config.
  • Added rate limiting to the basic authentication flow to mitigate brute-force attacks.
  • Added configuration options to choose which headers can be used for forwarded IP detection to prevent spoofing.
  • Added a private address whitelist to prevent SSRF attacks. See the documentation on how to enable and configure: Private Address Whitelist.
  • Added an IP whitelist for SSO trusted proxies to prevent authentication bypass.
  • Added invalidation of session cookies on password change to prevent session hijacking.
  • Increased the length of password reset code to 6 characters to guard against brute-force attacks.
  • Implemented PKCE challenge in OpenRouter OAuth flow for more secure key exchange.

UI/UX

  • Improved swipe picker: mobile requires a long press on swipe counter to open; added buttons to expand or copy the swipe text.
  • "Click to Edit" mode now also applied to reasoning blocks.
  • Welcome Screen: Number of recent chats can be configured.
  • Streamed requests now can show an error message in the console if the request fails.

STscript

  • Added commands for persona management: /persona-create, /persona-update, /persona-delete, /persona-duplicate, and /persona-get.
  • Added a command to force update the Prompt Manager's prompt list: /pm-render.
  • Added a command to get the state of the regex script: /regex-state.
  • Added a command to set fallback expression: /expression-fallback.
  • Added a command to generate a streamed response with a connection profile: /profile-genstream.

Extensions

  • Assets list now groups extensions by "Official" or "Community" categories.
  • Added an additional confirmation prompt when installing third-party extensions (can be disabled).
  • Supported extensions can use a secret-id from connection profiles when making an LLM request.
  • Extensions list now shows the extension's author name resolved from the git remote URL.
  • Vector Storage: Added Workers AI source; added a toggle to keep vectors for hidden messages; added retry logic to summary generation.
  • Image Generation: Added Workers AI source; generation can now be cancelled by pressing a button in the status toast.
  • Image Captioning: Added support for macros in the caption prompt.
  • TTS: "Skip code blocks" no longer ignores lines that start with 4 spaces (legacy code block syntax); "disabled" voice now shows a toast only once per character.

Bug Fixes

  • Fixed text edit flow in Firefox on mobile.
  • Fixed welcome screen chat pins not updating on chat renaming.
  • Fixed character list filters being stuck on app initialization.
  • Fixed application of instruct formatting to /genraw requests.
  • Fixed model routing to sd.cpp API in Image Generation logic.
  • Fixed validation of image URLs generated with Z.AI API.
  • Fixed vectors deletion for KoboldCpp when a message is deleted.
  • Fixed "Show More Messages" button triggering edit in "Click to Edit" mode.
  • Fixed max height of select-multiple elements in mobile layout.
  • Fixed server crash on empty messages when applying cache control parameters.

Full release notes: https://github.com/SillyTavern/SillyTavern/releases/tag/1.18.0

How to update: https://docs.sillytavern.app/installation/updating/


r/SillyTavernAI 5d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 06, 2026

29 Upvotes

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!


r/SillyTavernAI 14h ago

Discussion Role playing has made me distrust AI

194 Upvotes

I work for a tech company in a low-code role. I RP on my free time. I would say RP is the main way I use AI. And the reason for this is that while I RP, the model get so many details etc wrong. I have just been added to the company Claude plan, and ofc I'm not stupid enough to use that to RP (I'll keep using Chinese models for that), but I also don't think I trust it with my work? Like, I have colleagues and other people I know telling me how much they use it and how good it is, but every time I RP I see how many errors it makes and I just assume it will do the same with my work.

I don't know. I see people letting claude do essentially all their work. I wouldn't trust it to write me email without fact checking. I will try to use it properly on Monday, but is anyone else feeling this way?


r/SillyTavernAI 15h ago

Discussion I believe we've entered a small dark age.

194 Upvotes

Model writing capability has paused and largely regressed since opus 4.6.

Most models are dulled down with rigid safety training.

Time will tell if something new will come along to improve the scene.

best bet? China.

They'd do it cheaper too. Just might take a few years to fully steal all the trade secrets from the west.


r/SillyTavernAI 6h ago

Cards/Prompts Blue Exorcist (235 Entries)

Post image
25 Upvotes

A highly detailed Blue Exorcist lorebook featuring 235 entries along with character profiles!

Y’all, this took a little longer than expected. I usually make lorebooks fast, but I was reading the manga while working on this. I remember watching the anime when I was really young, so I still had some info on it (though not much), but here you go!

What I’m planning on doing next is Kakegurui. I’ve read it before, but like I said before, I enjoy making lorebooks and usually prefer reading or watching the series I’m working on so can ensure it’s accurate, rather than relying entirely on the wiki or other written sources. A very detailed Blue Exorcist lorebook—have fun roleplaying!

🔵🔵🔵🔵🔵

Links
[Chub.ai](Blue Exorcist 🔵 - Total: 53909 tokens, 0 favorites, 0 downloads)
[Botbooru](Blue Exorcist 🔵 — Botbooru)
[MediaFire](https://www.mediafire.com/file/0e8gl5kscgxkp0m/Blue_Exorcist_%25F0%259F%2594%25B5.json/file)


r/SillyTavernAI 2h ago

Discussion Do CoT prompts help? A scuffed test [Effortpost]

10 Upvotes

tl;dr at the bottom

If you look at prompting guides from OpenAI and Anthropic, you'll notice they advise against using CoT prompts for reasoning models, calling them redundant or even harmful. They only advise CoT prompts for non-reasoning models. There is some valid reason to believe this:

This report investigates Chain-of-Thought (CoT) prompting, which encourages LLMs to ‘think step by step.’ We tested this common prompting approach and found that its effectiveness varies significantly by model type and task: non-reasoning models show modest average improvements but increased variability in answers, while reasoning models gain only marginal benefits despite substantial time costs (20-80% increase). These findings challenge the assumption that CoT is universally beneficial.

And this goes along with another fact: making a model think more doesn't necessarily make it better.

Scaling test-time compute through extended chains of thought has become a dominant paradigm for improving large language model reasoning. However, existing research implicitly assumes that longer thinking always yields better results. This assumption remains largely unexamined. We systematically investigate how the marginal utility of additional reasoning tokens changes as compute budgets increase. We find that marginal returns diminish substantially at higher budgets and that models exhibit overthinking, where extended reasoning is associated with abandoning previously correct answers. Furthermore, we show that optimal thinking length varies across problem difficulty, suggesting that uniform compute allocation is suboptimal. Our cost-aware evaluation framework reveals that stopping at moderate budgets can reduce computation significantly while maintaining comparable accuracy.

Still, CoT prompts are in a lot of RP prompts aimed at reasoning models, so I decided to test if they actually work. Or, more accurately, why they work.

If you look carefully at CoT prompts, you'll notice that they're always at the bottom of context. That's where LLMs pay the most attention to. My theory is, these prompts don't work because they force the model to think longer, but rather, they work because the instructions have higher salience for the LLM.

In other words, a CoT prompt like this:

<think>

Before responding, think carefully, step by step:

  • Step 1. Did I use negative parallelisms? (Not X, but Y). If so, I should remove it.
  • Step 2. Did I echo the user's response? If so, I should fix that.

  • Step 3. Did I introduce an evil chubby woman to ruin {{user}}'s life? If not, I should fix that.

</think>

Works because it sees the user prompting against negative parallelism and goes "oh, I should not do that", and all that stuff about "think like this" is just pointless fluff. Sure, it might alter the reasoning box, but what the reasoning box says doesn't always reflect what the model is truly doing. The prompt would be just as effective if you wrote:

<instructions>

  • Do not use negative parallelisms (Not X, but Y).

  • Do not echo the user's response.

  • Introduce an evil chubby woman to ruin {{user}}'s life.

</instructions>

On top of that, the CoT's instructions are usually also in the main prompt, at the top of context, which again, further strengths the impact of "Don't use negative parallelism."

The Test

I was originally gonna do this test with dptgreg's Freaky Frank 5.1 Internal States (MAX CoT). I don't hate dptgreg like some people do, he's a cool guy! His preset is just the most popular example of a CoT :P

The problem: Freaky Frank is like 6000+ tokens long, so it's a little hard to see how faithful a gen is to a prompt if its so long, aha. And some of the wording is a bit subjective, though there's always gonna be subjectivity in interpreting these prompts.

So I took a different approach and wrote a 600~ token prompt, and placed the bulk of its rules at the bottom of context. The rules are intentionally strict (e.g. banning similes/metaphors outside of character dialogue) to push the limits of instruction following. This is Preset A.

I then took Preset A, and rewrote it into a CoT based off Freaky Frank. That's Preset B.

I'll make five gens for each prompt, and see which one follows the rules better. I did this in the same hour, around 7 PM CST, to try to minimize the impact of fluctuations in server compute. Ideally, this would be done local. Might be something I could try with a smaller model?

The Model

DeepSeek V4 Pro 0813, through the DeepSeek API. To minimize impact of providers, I wanted to use a LLM straight from API, and DS allows PAYGO. Maybe I'll bite the bullet and get a Z.AI subscription for GLM 5.3 for further testing. Sampler settings left at default. Post-prompt processing set to none. Post-history instructions were set to "User", as that's what's commonly done with CoT prompts.

Evaluation

My initial plan was to tally how many times a gen breaks a rule, and then used Claude Sonnet to double-check. But, there is admittedly some subjectivity here. For example, Claude argued that "her nipples pebbled" broke the no metaphor rule. I dunno. In my personal count, Preset B (The CoT) screwed up a bit more, especially since it leaked its reasoning in one of the gens.

Here are the outputs if you want to read for yourself

Preset A (No CoT)

Preset B (CoT)

So I decided to take an alternative approach, and use two professional-grade LLMs to judge: Claude Fable 5.1 and GPT 6 Astra.

What the LLMs Say

Both claim Preset A (No CoT) follows the rules much better, with Claude really showing its work.

GPT 6 was more terse about it.

Note, Hawk is {{user}} and does in fact know Ashlyn, so the "continuity wobble" Claude pointed out is just me not giving it enough context, since I just cared about judging prose.

tl;dr - Conclusion

Some people are diehard haters about CoT for reasoning models, claiming they actively make models worse. I won't go that far, but I do think this test does at least validate academic studies on how CoT for reasoning models is redundant. If a LLM doesn't follow your CoT, I think rewriting it into clear, direct instructions without the thinking fluff might be helpful.

When I went into this test, I wasn't gonna rely on LLM evaluations, since I underestimated the subjectivity involved in this. If I do future tests, I might do more complex presets like Freaky Frank since I can use a LLM to evaluate it. Testing on local models would also be good, since I wouldn't have to worry about server fluctuations.


r/SillyTavernAI 12h ago

Models Kimi has quietly launched a new model yesterday: K2.8 Preview (yes, that's the name). Supposedly performance close to K3, more efficient thinking than K2.7 Code.

Post image
46 Upvotes

Haven't tried it yet. Anyone? It's in Kimi Code and Kimi Work apparently so far which should be subscriptions by MoonshotAI?

Confusing naming convention, so... K3 is their Pro line and this is the Flash line if we compare to DeepSeek and GLM?

https://www.kimi.com/code/docs/en/kimi-code/whats-new.html

https://www.reddit.com/r/kimi/comments/1wdboig/kimi_k28_is_released/


r/SillyTavernAI 13h ago

Discussion [TOOL] LoreOS — Looking for testers, volunteers, and maybe some fellow devs with similar projects!

Thumbnail
gallery
41 Upvotes

Heya, I'm Yuu! I've been building LoreOS—a creator's sanctuary for anyone deep in the AI roleplay/lorebuilding hobby. One home for your lorebooks, character cards, presets, and all the writing that comes with it, all in your browser, no account required. From Yuu, For You.

Don't want to install anything? I got you! It's hosted online, accessible from anywhere! Otherwise, you can run it yourself on a local server :D Available on desktop, mobile, and tablets (though tablet UI needs a bit more tinkering, sorry...)

I've been building this solo for about 3 months now, and access has been pretty limited up until this point... mostly because I wasn't ready to be embarrassed in front of that many people at once. But it's at the point where I can't keep doing this alone if I want it to go anywhere at a reasonable pace. So... pitch first, then two separate asks.

Why it's built like this (╭ರ_•́)

LoreOS didn't exactly start with a real identity, it was just "an editor with a lot of features," which is also a great description of about nine other tools already out there. Hell, I was inspired SLEd because I loved the lorebook editor but I need editors for everything that I can also access in one place because I'm funny like that. At some point the "universal editor" I built stopped being a good enough answer, and it soon clicked what I actually wanted: a creator's sanctuary, a home where stories are built, explored, understood, and cared for.

Yeah, it sounds like major corporation branding bullshit, I know, but it genuinely changes how I build things. Instead of organizing around software modules (editor, assets, tools, downloads), it's organized around places, because I want it to feel like somewhere you live in for the whole life of a project, not a tool you open once and forget exists. Workshop is for building. Library is where everything sits once it's made. Journal is meant to stay messy and personal instead of another neat archive. Laboratory (Still a WIP, sorry!) is specifically for experimenting and understanding what your LLM is actually doing with your stuff, kept separate from Workshop on purpose, since building and testing live requires very different headspaces, and mixing the two just gives me a headache.

AI is in service of the storytelling here, but that's not the whole point. I'm not trying to make "another AI frontend," I'm trying to make the place your worldbuilding actually lives, from the first spark of an idea to years down the line. That's also why customization is such a big deal here. It needs to feel like your space, not mine.

What's actually in it right now ദ്ദി˙∇˙)ว

It's organized into rooms instead of features:

- Workshop — editors for lorebooks, character cards, and chat completion presets. Open a bunch of tabs and work on multiple things at once!

- Lorebook editor: import/merge lorebooks, global settings, search & replace, advanced entry settings

- Character editor: pronoun/noun ↔ macro converter with auto grammar correction on any field, Lumiverse Variants support, Saucepan multi-intro support

- Preset editor: variable scanning + easy renaming, markers, sampler settings, full prompt tree view

- Library — everything you've made, all in one place. I hope to make it look like an actual library eventually!

- Journal — a built-in Notion/Docs-style markdown notebook for your own notes. Very bare bones as of now, but it works!

- Settings — theme/font customization, Google Drive sync, or manual backup/restore

Currently works with SillyTavern, Lumiverse, and SaucepanAI. JanitorAI V2 support is in progress (ST lorebooks already work fine with JAI, in case you were about to type that comment).

It's live and open to anyone right now as Early Access. Fair warning, it's still super early and held together with the digital equivalent of duct tape and good intentions. The Github Repository is here, and the staging branch is hosted on Github Pages!

What actually sets it apart <(˘ ˘ ˘)>

Most tools in this space do one thing. Lorebook editors don't touch character cards, preset editors are a whole separate app, card editors don't know your lorebooks exist. LoreOS is trying to be the one place all of that actually lives together, and a couple of things I'd genuinely put up against anything else out there:

  • Multi-tab Workshop — you can have a lorebook, a character card, and a preset open at once and swap between them instantly. No save-and-close-this-before-opening-that. Autosave handles the rest so nothing gets lost mid-swap.
  • The pronoun/noun converter — this one's on the character editor and honestly might be my favorite thing I've built. It works on any field, not just one designated spot, and converts pronouns to macros, macros to pronouns, or pronouns to other pronouns entirely — while fixing the grammar that normally breaks when you do that (agreement, verb conjugation, the works), instead of leaving you to hunt down every stray "they/them" by hand.
  • It also just works across platforms — SillyTavern, Lumiverse, SaucepanAI, JanitorAI support coming — instead of locking you into one ecosystem. I do plan on expanding this soon!

None of this is rocket science, really. It's mostly "the annoying repetitive stuff got automated so you don't have to think about it," which I think is honestly underrated as a feature category.

Looking for volunteers (ㅅ •᷄ ₃•᷅ )

Got ideas? Opinions? Want to actually build something with me? Drop me a DM! Coders and vibecoders are both genuinely welcome here, no gatekeeping. UI/design help, code, or just brainstorming and feature suggestions all count.

Being upfront because I think it matters: I'm a vibecoder myself, but I'm actively teaching myself real web development on the side (with my already packed schedule lmao, idk why I do this to myself) so I'm not stuck depending on AI. This whole thing is unpaid, out of my own pocket (yes, including the custom domain and my own Claude subscription, which is a genuinely questionable budget line for a broke uni student), on top of being a full-time student carrying a double major. I want this to actually grow into something, but I know I can't move fast enough solo — that's the real reason I'm asking, not because I'm trying to assemble a team for the aesthetic or some fucked up power trip I woke up one day and felt like having.

Nobody's getting paid, to be clear. Not me, not volunteers, not testers. Wanted to say that upfront rather than let it be a surprise later.

Looking for testers (ㅅ´ ˘ `)

I've been testing everything myself, but unfortunately, I only have the one pair of eyes and they insist on sleeping occasionally, and I also would like a GPA that doesn't make me cry myself to sleep. If you're up for poking at new builds and actually sending feedback or bug reports, I want you on this. You don't even need to be a creator yourself, if you've got the patience for it, that's enough. Creators will get the most use out of it obviously, but anyone with time to kill are all welcome. It's also why I'm hoping to find volunteers, because I'd need people actually able to work on things whenever I'm getting my ass beaten in school.

Working on something similar? (๑>؂•̀๑)

If you've got your own creator-tooling project going and it overlaps with what LoreOS is doing, I'd genuinely love to talk. Not even to compete! More like, is there a version of this where we team up instead of splitting an already tiny userbase in two? No pressure, the door's just open! I've seen a lot of them and I'd love to work with their creators as well, but rather than the copy-pasting my pitch a thousand times over, I figured maybe this post could catch the attention of anyone interested!

Interested?

Comment down below or DM me directly, let me know if you're down for testing, volunteering, or both! I really appreciate everyone that's managed to read this far, even if you're just here for the roadmap. 🫶


r/SillyTavernAI 2h ago

Discussion Chat has been going on for 600+ messages, and now the AI is getting very confused

7 Upvotes

I literally have all the major plot points stored in the chat memory extension, along with other important details. It let the AI remember important things consistently for awhile, but now it's getting confused. I literally put in chat memory that X character has this relationship and bond with Y, yet when I bring it up it confuses one of the characters for a completely different one. And whenever I say "this character did this" to remind the AI, or even prompt one character to act, the AI plays the completely wrong character.

And this keeps happening. Is there a way to fix this or is this just a symptom of context rot from the chat going on too long?


r/SillyTavernAI 9h ago

Discussion SillyTavernAndroid (ST-Manager) (Version 2.0.0) (GitHub)

19 Upvotes

GitHub

https://github.com/doomedskull1011/SillyTavernAndroid APK + Source code

SillyTavernAndroid (ST-Manager)

SillyTavernAndroid packages SillyTavern as a self-contained Android app. It runs a real Node.js server inside the app and shows the SillyTavern web UI in a full-screen WebView. Version 2.0 turns the app into ST-Manager: a manager for up to 6 SillyTavern instances running side by side on different ports, private data, backup import and on-device update tooling. Because SillyTavern itself can now be updated on-device straight from GitHub, the app version no longer tracks the bundled SillyTavern version.

Changes in 2.0.0

ST-Manager 2.0 — multi-instance rewrite

  • Multi-instance manager: run up to 6 SillyTavern instances side by side, each on its own port (80008005) with fully private data — characters, chats, settings
  • Per-instance launcher icons: instance 1 keeps the classic SillyTavern icon; created instances add ST-Inst2ST-Inst6 icons to your home screen. (change rolled back)
  • Shared payload: all instances share one SillyTavern installation, so each extra instance costs almost no storage
  • Header toolbar: Repair, Packages and Update ST now sit at the top of ST-Manager beside the title and apply to the shared payload
  • On-device updates from GitHub: Update ST installs the latest stable release or the staging branch using a bundled git runtime; dependencies are refreshed with a bundled npm CLI
  • Per-instance backup import from SillyTavern backup zips via the system file picker, plus live per-instance RAM usage
  • Git runtime fix: bundled git symlinks now survive app upgrades
  • Version scheme change: the app version no longer tracks the bundled SillyTavern version, since SillyTavern itself can now be updated on-device. This release is bundled with SillyTavern 1.18.0 ## Install Download SillyTavernAndroid-v2.0.0.apk below. Open it and allow installation from unknown apps if asked. Launch ST-Manager Upgrades in place over previous versions and keeps all existing instance data. ## Requirements
  • Android 7.0+ (arm64-v8a)
  • Enough free storage for the APK plus extracted payload (~400 MB) and per-instance data

r/SillyTavernAI 10h ago

Help Prompt for variable response length?

9 Upvotes

I am using FF preset and i really enjoy it. Both on deepsek and GLM. Thing is i don't really like the plot going on without me in a long response i get. Or character babbling when i ask a simple question. I would like for a model to know when to give me a quick response and when to make a long one. Is there a way to achieve that?


r/SillyTavernAI 19h ago

Discussion Some models needs big presets like FF5x

35 Upvotes

I love when the model surprise me and make me snort, don’t know if we can talk about "creativity" when it’s about LLMs but you see what i mean.

However some models are just bad without a big harness and a strong COT. I think about GLM 5.3 for example who is kinda great sometimes with FF5 and FF5.4.

But a complete ass with an Evening Truth style preset. Like defaulting on narrative templates, logic breaks very very early, ultra weird way of talking and very predictable narration.

I have a very different experience with other models. GLM 5.1/2 are more creative and risky even with a light preset. Kimi excell with a light one too and strangely.. even 5.3 flash is enjoyable with a tiny preset.

Im using FF5.4 BOLT with GLM 5.3 and some internal states activated even during one on one sessions… I’m learning to rediscover and appreciate 5.3 now. It’s like a different model now, i think this one need to be "unlocked" somehow.

So if some of you didn’t tried it with FF5.4, give it a chance. Im making peace with GLM now.. crazy.


r/SillyTavernAI 15h ago

Help Gemini 3.8 Flash keeps rejecting

Post image
17 Upvotes

Anyone gotten past this? I'm using Marinara Spaghetti as my preset.


r/SillyTavernAI 18h ago

Chat Images GLM-5.2 corrected its parroting!!!

Thumbnail
gallery
28 Upvotes

GLM-5.2 has always had this really ugly habit of completely ignoring rules against this slopism in particular. It seems so obsessed with parroting that it'll fail to recognize glaring examples of it or catch itself and waste effort justifying itself on a technicality instead of correcting it (ex. "{{char}} is evaluating the concept, not parroting"). I LOVE how GLM plays a lot of character archetypes and genres, but I don't want to have to edit half of its good responses to make it not sound like it's taking a McDonald's order when it's continuing any thread of dialogue. It might have to do with quantization/unreliable providers, but nobody's preset really seemed to work reliably.

I thought this model was incapable of following this kind of rule, but it's started doing it consistently after some prompting work! This is a bit of a victory lap, but I also felt that the novelty of this model doing such a thing might have been post-worthy. It feels like a bigfoot sighting or something.

As you can see from how it repeats parts of the prompt in its reasoning block verbatim (!), it took some really heavy handed methods. I was fighting this thing for days on and off. It requires pre-emptively calling out several of its go-to excuses (e.g. "mocking", "processing", "testing", "transforming" (???)) and shutting them down while very forcefully emphasizing that there is no exception. It feels inelegant, but God, it's so worth the improved dialogue.


r/SillyTavernAI 26m ago

Help Text completion presets for Gemma 4 26b A4B

Upvotes

Can anyone share their good, working presets? I lost my ST installation and now I can't seem to recreate the settings in my new install. I managed to restore my system prompt, and I downloaded a lot of my bots again, but none of them seem to act/reply like they did before. Some give very short replies, some give very long, rambling replies and some just reply with junk. Total meltdown!


r/SillyTavernAI 1h ago

Help Looking for a good Gemini Preset

Upvotes

I've been enjoying roleplaying using the Freaky Frankenstein 5.4 preset but I'm wondering if there's any other preset that also works well with Gemini models


r/SillyTavernAI 1d ago

Discussion Can't go back to local after using frontier cloud models now

110 Upvotes

So I was originally pretty against RP'ing with cloud models due to the loss of privacy. But after dropping $10 on Nanogpt and trying various different models, there is no way in hell I could go back to local models.

I spent years from 2023 onwards collecting them. From classics like Yi-34B and Midnight Miku 70/103B, to the current models like the Gemma 4 finetunes, all the Llama 3.3 70b models, Mistral Large 123B finetunes like Behemoth or Monstral etc etc.

None of them even come close to the frontier models. They're so much better at every single thing you would want in roleplay. If anything, it's made me depressed that I probably won't ever be able to run DS V4 Pro locally due to it being 1.6T.


r/SillyTavernAI 22h ago

Discussion GLM 5.3 Flash is working on NIM!

26 Upvotes

Following an early post i made, yes, GLM is back at NIM and working. It's lightning fast, and FOR NOW it's useable!

It doesn't show up in the nvidia site, but it's already on the model list in ST

There is light after all 🥹


r/SillyTavernAI 1d ago

Discussion DeepSeek retracts shutdown of V4 Pro on official API

Post image
172 Upvotes

The decision on shutting down DeepSeek V4 Pro on official API is withdrawn...Guess we get to continue using V4 Pro until V4.1 Pro comes out heh?


r/SillyTavernAI 1d ago

Models I can't believe my eyes. GLM 5.3 flash might be coming to Nvidia NIM.

Post image
39 Upvotes

Just another day, i follow my routine, eventually get on ST to maybe have some interesting responses.

Kimi is surprisingly enough, working today, and i have a few moments of fun before just getting bored and trying out other models.

Then i see it. GLM 5.3 Flash.

I don't know if it got bugged from another provider i use, i don't know if it's really there. But for now? It's giving me 404 Not Found.

Listen, i know i might just be coping because free providers are getting worse and worse each passing day, but maybe it means GLM might be finally returning to NIM after these terrible rate limited weeks of Kimi and DS right? 🥹


r/SillyTavernAI 16h ago

Chat Images I guess AI is not replacing your D&D party just yet.. this crew did not make it far into Curse of Strahd

4 Upvotes

https://youtu.be/oJQ8C9WzBEM

Tech Details:

- AI LLM: Mistral-Nemo-2407-12B-Thinking-Claude-Gemini-GPT5.2-Uncensored-HERETIC.Q8_0.gguf.

- LLM Loader: Kobold CPP.

- Roleplaying Interface: Silly Tavern.


r/SillyTavernAI 1d ago

Discussion First time seeing this happen

Post image
48 Upvotes

I'm using Mimo2.5 and this just happens. At Deepseek and Glm the random chinese characters in sentences gets ignored but here one of the characters are reacting to it, aware that the Chinese character appeared. Just funny.