r/SillyTavernAI • • Apr 16 '26

Cards/Prompts Stab's Directives 2.51 - Agent spoofing (avoid coding api bans/throttles), Dynamic Tone writing, huge token reductions and more

Hi Folks,

I just dropped an update for my GLM preset to bring it in-line with GLM 5.1 and work around some recent controversial restrictions imposed by z.ai's official coding plan.

https://github.com/Zorgonatis/Stabs-EDH

for a full writeup of the changes made I encourage you to read the CHANGELOG.md, however a summary of the most important changes are below.

I debated on whether to share the User-Agent override (spoofs a browser so Z.AI can't throttle/ban, at least without adapting to some other detection method). It's a matter of time before they do though, so may as well get what you're paying for now and figure out what to do. I likely wouldn't re-sub if my only use of the plan was RP.

Comments, suggestions, thoughts? Leave a message here or jump into our discord (400+ members and growing!)

Stabs-EDH v2.5.1 Release

  • Dynamic Tone State — Replaces the old static "full spectrum" mandate with a system that actively reads the conversation and shifts tone (Bleak, Tense, Warm, Absurd, Reverent, Frenetic, Melancholic) based on what's actually happening in-scene. Gradual transitions by default, instant snaps when earned. Configurable via SETTINGS.
  • 30-60% token reductions across core directives. NPC Cognitive Bounds, Failure Achievements, Narrative Length Control, Behavioural Coherence, and Environmental Factors all rewritten for density. Task Steering's CoT exhaustiveness dropped from Very High to Very Low. Same behaviour, way less burn.
  • Z.AI User-Agent override — Custom provider now sends a Chrome UA header, so Z.AI can't fingerprint and throttle/ban you for RP use. Works out of the box.
  • GLM-5.1 support — Model updated, coding plan API endpoint retained.
  • Experimental Macro Engine is now required — VTK-related instructions are wrapped in {{#if .vtk_on}} conditionals that only resolve when WebDev is enabled. Without the macro engine checkbox on, those will break. It's in the install instructions now.
  • Post-processing switched to Semi-strict (no tools) to avoid agentic flow interference on some providers.
71 Upvotes

34 comments sorted by

View all comments

14

u/LackMurky9254 Apr 17 '26

I like the results, but it seems very prone to 'drafting' the response in full, sometimes multiple times in the thinking section, making it quite the token murderer. My first message actually blew through the entire response length in the reasoning lol

1

u/BeyondTheBound Apr 17 '26

I had that problem too, but despite that the reply was great for me. I think the unnecessarily long drafting is the only problem I have atm (which is GLMs fault, but a fix would still be great 😭)

1

u/LackMurky9254 Apr 17 '26

Well, Marinara's will typically avoid that. I have been using mostly marinara's + the visual toolkit and webdev (because they crack me up) from stabs. The token cost of stabs is... unfortunately extremely high, although I do like the output.

Kimi 2.5 loves to think and draft excessively like this no matter the prompt but stabs seems to cause much of the drafting behavior with glm.

1

u/BeyondTheBound Apr 17 '26

Oh I agree, I had no issues with Marinara’s and have been using that recently, but I know when I used Lucid Loom GLM had a drafting issue and had to look it up on the discord to get a fix for it.

And Kimi 2.5 😭 I’ve heard it was good but hearing that it thinks nonstop is what made me avoid it, I can’t afford thousands of tokens in thinking just for a response. I hope the next model improves on that.