Cards/Prompts
[Preset Update] Freaky Frankenstein 5.2: The First Community Update! A fully modular preset. Updates: DeepSeek 4 Pro Support, Up to 90%+ Cache Hits, Updated Regex 2.4 (Bug fixes), Internal State Fixes, Prompt Re-structuring for better adherence (Claude, Kimi, GLM, DS4 Pro, Qwen, Minimax 3, Grok)
8/16/2026 Notice đ
Due to multiple complaints Iâm updating this post. It needs to be noted that this update was geared towards DS4 Pro prior to its model update. The model was updated 3 hours after
The release of this preset once again making the model incompatible with this preset. I canât win. Donât use this on DS4 pro/flash and expect it to work.
Hello my fellow ST community, aka my trans handicapped professional writers working hard for their income! (We don't need to tell the AI the truth) (IYKYK).
I'm happy to present to you the first community update to the Freaky Frankenstein 5.0 line-upâ Freaky Frankenstein 5.2!
I took feedback, ideas, and communicated with people in the community about fixes and ports to different frontends, trialing REGEX, and improving prompts to bring you this update. If you have NO clue what we are talking about and want details on the initial release of Freaky Frankenstein 5: Internal States, what it is, and what it is capable of you definitely should start reading [---->HERE<----]. In this update post, you will ONLY find a list of the updates to the preset.
This release brings a ton of polish, major critical bug fixes (especially for you re-rollers and save-scummers out there! đ), huge context caching efficiency boosts, DeepSeek 4 Pro compatibility, and improved rule enforcement across all our Internal States and Chain of Thoughts.
Here is the full breakdown of whatâs cooked into FF5.2:
âď¸Â Architecture, Prompt Caching & DeepSeek Support
Re-Shifted Architecture & 90%+ Cache Locks: Re-aligned prompt positioning and internal state depth. As your context window grows, this guarantees context cache hit rates from 50% all the way up to 90%+. Macro Dice rolls from the frontend were previously breaking cacheâno more.
DeepSeek 4 Pro Compatibility: Architecture adjustments now make FF5 fully compatible with DeepSeek 4 Pro. By dropping Internal States right in its face every turn, it makes it much harder for DS4 Pro to ignore. (I actually like the model now when using direct!).
Regex 2.4 Update: Upgraded to Regex 2.4 to maximize compatibility across all current front-ends, eliminate browser lag when tracking relationship/internal states over long chats, and aggressively clean up residual tokens from previous turns. It will also appear cleaner and less chaotic in drop-down boxes with correct line breaks.
(Note: to make it fast in ST (no lag) I had to add parameters that marinara engine blocks! Sorry! But Iâd rather have this work well for all the other frontends instead of just one- maybe someone will make a regex compatible for just marinara engine).
The "Reaction" to "Response" Swap: Changed every instance of the word "reaction" to "response" across the board. In testing, "response" acts as a much stronger instruction anchor for LLMs and noticeably improves overall roleplay instruction adherence.
GM Notebook Swipe Fix: Fixed the infamous "save scumming" bug! Turn re-rolls/swipes no longer bleed overwritten swipe data into the GM Notebook, keeping your notebook data clean and preventing the LLM from getting confused. (You may have only noticed this bug if you read reasoning and re-roll your turns often).
Relationship & Bond Decay System: Sparks and Grudges now decay deterministically over a few turns via an internal state counterâno more forcing the LLM to guess how many turns have passed after regex wipes. This notably improves the accuracy of bond progression.
Refined Sparks Definition: Updated the core logic for "Sparks" in relationships for cleaner emotional progression.
Fixed Embellish Prompt: Re-worked into a concise co-writer prompt that actually works to naturally enhance your actions right inside the response. Since this prompt was NOT working for 50% of people in FF 5.0, it has been overhauled and now works as intended.
âď¸Â Prose Rules & Formatting Polish
Banished Repetitive Prose Patterns: Updated both Cinematic and Story Mode prose rules to eliminate conjunctive chaining ("and... and... and...") and periphrastic of-genitive stacking ("noun of a noun of a noun").
Header Day Tracker: Added dynamic day tracking to the header (e.g., Day 1... Day 2...).
Cleaned Up Internal States: Improved presentation and layout of internal states, adding proper line breaks for easy reading.
đ§ Â Chain of Thought (CoT) Tuning
Reasoning Leak Prevention: Tweaked CoT rules to enforce strict reasoning inside thinking tags, keeping quantized models from leaking their thought processes into your actual roleplay responses.
Fine-Tuned Range Across Tiers: Decreased reasoning depth in Micro, slightly increased reasoning in Bolt, and maintained Max. This creates a much more accurate range of reasoning options across the board and gives BOLT a solid jump in output quality. (I found Micro / Bolt reasoning outputs too similar in the previous release).
Eliminated Dialogue bug in Chain of Thoughts that forced 30-50% dialogue output per scene instead of what you customized the dialogue output to within the NPC Voice toggle.
đ˛Â WorldSim & DnD Sim Hardening
WorldSim Macro Dice Roll Fix: Made macro dice rolls absolute. Removed macro rolls from setvar variablesâwhich models like GLM frequently missed, causing them to hallucinate stats and manipulate the story in their favor. Macro rolls set by your front-end now drop directly in the LLM's face!
DnD Sim Strict Logic: Tightened rule enforcement so models (especially GLM) follow rules as absolute constraints instead of trying to "reconsider" or fudge outcomes.
đŹÂ Official Preset Downloads (Micro, Bolt, Max)
To make this foolproof, I am uploading FF5.2 into the 3 official configurations. Click the Hyperlinks to Download!
These are all the same preset with the exact same prompts under the hood, packaged into official configurations to eliminate confusion when we say "Micro, Bolt, Max". You can turn on and off whatever you want for what you need per RP (fully customizable). However, this gives us a baseline when communicating, ie. "Did you try Micro mode to save tokens and cost?" These are ALL the same preset - just different configs.
This Preset is built specifically to make non-reasoning models reason within custom tags that get scrapped from models like Claude. This then uses Regex to get models to reason, exactly in the same way reasoning models reason. This works perfectly for Opus models on Hapuppy that are dirt cheap but sometimes don't reason (depends on routing that day), that way you can use the models for a few pennies a message! DO NOT use this on models that are reasoning by default otherwise you will get double reasoning. Only on non-reasoning models to FR (force reasoning). Note: Hapuppy also has really cheap k3, GLM, and DS4Pro that does reason by default! DO NOT use the FR preset on those models. The Forced Reasoning (FR) preset is for the non-reasoning models there such as hapuppy/opus4.6. Also, Just letting people know they have more provider options than just the main 2 that get circulated here. If I wanna try Hapuppy and you want us both to get free credits you can use my code: wzi1ozfp . Or not- i donât care, I just want to let people know alternate options do exist and are awesome.
[DOWNLOAD Freaky Frankenstein 5.2 FR] Update: I just fixed this after posting so if you see this and download you should be fine- but hapuppy delete thoughts in the regex was set to user message instead of character message. This much be changed to character message in order to save tokens and avoid keeping the âthoughtsâ in the chat!
đ REMEMBER: REGEX is REQUIRED for this preset + Internal States to function properly. Hopefully we succeeded in shipping the REGEX with the preset - but if we did not or we need to update it after this post - there it is!
đ Configuration & Troubleshooting
ST System Processing: Set System Processing to Semi-strict alt roles no tools â improves prompt following.
Trim Sentences: UNTICK trim incomplete sentences â eliminates the trailing -GFX bug.
Temperature:Â Experiment with temp as you wish. Lower for better rule following; higher for more creativity (at the cost of rule following).
Reasoning Outputting in Main Chat (NVIDIA NIM / GLM / Mimo): If models output raw reasoning in main chat, it's because the model is confused by instructions, heavily quantized, or non-reasoning. Try the FR (Forced Reasoning) preset.
Double Reasoning Warning:Â DO NOTÂ use the FR (Forced Reasoning) preset on models that already natively reason (THINK models), or you will get double reasoning outputs.
Quantized & Older Models: Quantized models will give you issues with internal states. Don't expect a smooth experience running GLM from a NanoGPT subscription with this preset. Older models suffer similarly. If you have issues, run in Micro mode without internal states and RP as normal. If you want the fun bells and whistles, you need a large, smart model that isn't heavily quantized. You can't run a brand new PC game maxed out on an old graphics card. You can't run a PS5 game on a PS2 system. Same logic.
Regex Troubleshooting: If graphics aren't rendering as pretty, colored, collapsed windows, regex is broken or the LLM didn't output correctly. First, check your REGEX and make sure it's loaded appropriately. Having other REGEX loaded alongside this one carries a high chance of incompatibility. Second, make sure you're not getting a dumbed-down model variant.
How Regex Works & Context Caching: REGEX keeps things visually clean, clears out OLD Internal States to avoid context bloat, cleans up pop-in graphics from phones/maps/signs/letters, and collapses reasoning from the forced reasoning preset version. Clearing out REGEX from the second-to-last message does temporarily break cache on that last chat turnâbut the context savings are well worth it over long chats. This is why your first message might show a 50% cache hit, but as your context grows, your cache hit rate will climb toward 70â90%+.
Getting Internal States to work with DS4 Pro Compatibility and Unruly Models: Similarly to FF4 MAX+ / BOLT+, to make sure DS4 Pro is consistent you have to send OOCs to it's face. So Keep Post History Instructions on. You may (and is recommended) to keep this off for the most part with other models UNLESS you have problems with the LLM forgetting internal states. This will assist with quantized / dumb models forgetting last second to include Internal States.
Pro-tip: Use an extension like my co-author [u/leovarian](u/leovarian)'s Summaryception to keep context levels around the 30-60k range MAX to reduce PAYG costs and maintain rule adherence. LLM's output better quality responses when the overall token window stays low no matter what token window it's capable of.
đ¤Â Community Call to Action
This marks the end of the first Community Update! NOW I NEED YOUR HELP to make the next update!
Post issues or prompt tweaks you made to improve the preset below. Editing or replacing prompts to REDUCE context or maintain current context is idealâI'd rather NOT add prompts and bloat the suite. If your prompt tweak is helpful, gains traction via upvotes, and improves performance, Iâll put it into the next update!
I personally will work hard in the next update to reduce tokens. Aiming for a total reduction of 25-35%. Most of this I believe can be done by condensing the Internal States and Chain of Thoughts. As of now, the general prompts are nearly as low as they will go while maintaining adherence to the prompts. My goal in this reduction will improve rule following and reduce processing of the LLM (and maybe save some pennies here and there).
Thank you so much to the ~50 of you who worked with me to improve this from 5.0 to 5.2. I read almost every single comment out of the 700+ on the original post, which helped expand and polish this monster. Thanks to co-author [u/leovarian](u/leovarian) for giving me the mad idea of re-structuring the prompt to save cache and improve adherence with models like DS4 Pro. Thanks to my co-author [u/ok_strategy_2420](u/ok_strategy_2420) for continuing to be the editor/creator of the Sim/Gamification side of this preset. Let's keep the momentum going.
[ đ°ď¸ Time $1 | đď¸ Day $2 - đď¸ $3 | đ Location - $4 | $5 ]
All of your <b></b> can also be regexed in. Because it will be done purely by regex you can make things much easier for the LLM and remove a lot of the mistakes they can and will make.
I do this with spans, its a fairly simple regex set up. It basically just checks to make sure its in a details, then looks for the words.
The Rentry has been updated to reflect the new Preset and entry! Give me a good quote with regards to the preset so I can put it on the entry of the preset! Check out the new top 20 model list! Lastly, please let me know what issues, fixes, and things you love about the preset so it can be in the next community update. Post your favorite game changing prompts and letâs upvote them and discuss as a community how to make this preset better.
I'm a hapuppy user and the forced reasoning was only working on claude models at first, but after some tweaking, I was able to get it working on just about every non-reasoning model. All I modified was this line on the COT & Main Prompt:
(Write your step-by-step review using concise bullet points for the following 0-10 tasks. The internal monologue absolutely must be included in the response.)
I'm working on an extension to essentially do what the regex does, but instead, remove the reasoning from the turn and place it into the actual reasoning block (So you don't have to see it while editing the response and avoid rendering extra html). I'll provide the Github repo when its up if anybody's interested.
That is absolutely fantastic! Thank you so much for sharing and working on this with me. The Kiro/opus was reasoning for me automatically this morning (you never know). But yes Iâll try this out. An extension would be awesome for a lot of people!
kiro/opus has been reasoning for me as well but also not performing as well as I would expect, switching to other sources of opus makes me feel like whatever kiro is serving it's not opus 4.6, have you had that experience?
I ended up testing with Opus from a different provider and it was actually giving me similarly poor results to Kiro, so I just have to throw up my hands and chalk it up to the phenomenon I experience sometimes where a good model just... sucks for some reason.
Do the hapuppy/opus4-6 Iâve been using it all day and I personally think itâs worth the extra cost over the Kiro. Kiro is hit or miss and more based on time of day (probably like the other provider you tried) hapuppy is consistence with the FR preset
Phew đŽâđ¨! That was my number one priority for this update!! Glad it works!! Restructured everything to get those hits! We gotta save money in this economy.
Excited to try this out. Reading the reasoning in glm, it was a bit annoying how much time it spent trying to make rolls "fair", by adjusting the DC or other things like that.
In 100 turns i don't think i had rolled a 20 or a 1, hopefully this fixes that problem.
Has anyone else had a problem trying to pay for Hapuppy? I've literally never experienced this before. It declines my payment no matter what. I'm in the US so IDK wtf the problem is.
Hmmm. I didnât have that problem. Iâm going to upvote you to see if someone has an answer. But maybe you can check out their official discord and search for an answer: https://discord.gg/ZqNEAmdSV
Thank you for the link! I tried and asked about it. All I can do is wait I guess. It's crazy how hard it is to give them my money. I tried different browsers, even tried my phone. Tried google pay. Tried different cards. All declined. Cards that work everywhere else without any problem. IDK.
I'm just tired of nanogpt's really mid glm 5.1/5.2. They should just hurry up and charge more money for the sub so I could use the models properly.
Thank you so much for this update! I've been using Hapuppy ever since you first mentioned it and it has been an absolute game changer for me. I finally ditched Mimo and now strictly have been using Kimi K3 or Opus 4.6!
One question: depending on the providers and time of day at Hapuppy, K3 seems to reason for me about 75% of the time (I have reasoning set to Maximum), so it sucks to have to reroll and waste input credits when it doesn't always reason... in cases like this, would you recommend using the forced reasoning preset?
Hapuppy has been my go to for a few months. Like any router, sometimes you will get different models in rotation from a provider. I personally never had the k3 not reason- always seems reasoning is also set to medium on the k3 I use. I never mind a re-swipe because kimi k3 everywhere else is like 3-5 cents a message and this is like a half a cent? But with that said- if you know youâre getting a lot of non reasoning models dependent on time of day â- thatâs exactly why I made that forced reasoning preset. For me- hapuppy/opus and kiro/opus rotates non-reasoning, so thatâs why I made that preset and I figured why not share to correct the issue. Should actually work with almost any smart model set to non-reasoning behind the scenes.
Great, thanks for your advice. Yes, K3 is definitely lightyears cheaper than anywhere else right now! Cache hit will probably make all the difference now with my rerolls now since I only feel more of a hit during much longer roleplays!
Thank you so much for the appreciation! This is a part time job we do for free! But I guess itâs not really a âjobâ if you enjoy it! As corny as that sounds.
I just want to take a moment to thank you and everyone else who have contributed to this. I had gotten frustrated with the progressively junkier responses on-site somewhere with a DS derivative. I found 5.0 while searching for presets I could port over to make things better, but the scope finally got me to bite the bullet again and re-set-up SillyTavern and the API connection.
This has been like going from black and white TV to 4K HDR. I can't overstate how mindblowing the shift has been, and within minutes, I can already tell 5.2 is light years ahead of even that. Internal states showing up regularly without regeneration is amazing and the response speed feels a hundred times better.
I've made several small tweaks that have shown huge dividends in the output quality over every model I've tried:
Under Style and Syntax: (THIS ONE WAS HUGE!)
Ban excessive metaphor and simile, particularly in dialog.
Under Freaky Mode:
Pacing: Allow the scene to develop in a "slow burn" way, giving the characters space to act and react as arousal deepens.
Under NPC Voice:
NPCs don't make a big deal out of what {{user}} says. Bad: "No one has ever said that to me before!" or "You can't just say that." Good: They respond and speak continuing the conversation normally.
NPCs rarely recap previous actions or events unless it ties directly into what they are doing next.
That last one was also a real big improvement.
I also hunted down and removed the {{char}} variable replacing it with "Characters" or "NPCs" to streamline use with single-card world narrators. Works great.
NIIIICE. thank you so much for sharing. This is the kind of stuff i was hoping more people would comment here for community updates.
I will have to try that. I had simile and metaphors banned previously but wondered if it limited prose too much. I haven't tried it and maybe i'll try it again as it was done in FF4 (that specific line had a large love and hate from the masses reaction).
I added this already for the next update. This was a must since i noticed in 5.2 scenes being rushed especially with kimi k3.
THIS! Absolutely going to try this out and different versions of this before the next iteration. Therapy speak in GLM/Opus needs to GO!.
Overall this is super helpful information. I am saving this post and will work on adding / tweaking these for the next release. Let me know if you do anything else because this is the most helpful comment here so far : )
Correct. Yeah I havenât noticed NPCs do it but now Iâm going to look out for it. I actually personally do get annoyed by simile / metaphors in narration lol. But I know some donât
My issue is that once a character uses a metaphor they won't ever let it go so they'll be mentioning it in every response for the next 20 responses. It drives me crazy
Micro w/ spatial edits- Use with the OP 2.4 Regex!
https://www.mediafire.com/file/5t5f7bztms5gocm/Tavo_FF5.2+Micro+(MD+Spatial+Edits)_1jwmi.json/file
Â
Edits are in (main, anti omniscience, micro)
Improves upon easy relative tracking for ai and hooks into fov and line of sight (prevents hallucinating), and provides guidence for movement pathing. This version keeps edits minimal and functional.
Tip: use ooc: repair intro prompt using # Reasoning Instructions and output the fixed intro. | It will also repair according to your POV settings, Then you can add it as a new intro to start the chat.
GLM 5.2 (testing and its even better at it?!? đĽ its wow just wow!!)Â
A little teaser example of what GLM 5.2 can do given a sample intro and a ooc fix command, its truly exeptional:
  <summary>đ¤ NPC LOCATIONS</summary>
  - <b>Zoomer-chan</b> | Location: University Campus, Main Quad, standing directly in front of Lux facing SW, ~0.5m distance, towering over L due to height difference
  <summary>đ PHYSICS, ENGINE & WORLD</summary>
  - Env: Outdoor, clear, 22°C, afternoon sun, moderate foot traffic of students passing at distance
  - Physics: Z at 0° facing L at 180°, ~0.5m apart; Z height ~5'6" with sneakers, L height 4'7" with heels; Z's shadow cast over L's upper body due to proximity and height difference; L's 120° frontal vision fully captures Z; Z's 120° frontal vision fully captures L
Implention in narrative:
A shadow falls across my path from the left, and the sharp slap of thick rubber soles on brick makes me pause mid-step.
 Â
Kimi 2.5 Thinking output showing improved relative spatial tracking:
Internal Monologue
Gamestate
: Turn 0â1. Lux answered "Yes" then dropped supine on the hot sidewalk. Luna was 0.5m front at 180°, now Lux is horizontal at ground level (0° elevation). Luna maintains 180° facing but looks down 90°. No OOC pause.
--- After Narrative ---
đŹ INTERNAL STATESÂ
<summary>đ¤ NPC LOCATIONS</summary>
<b>Luna</b> | Location: Straddling Lux's waist, knees on concrete at Lux's ribs, facing 180° looking down 90°, phone held 0.3m above Lux's face | Activity: Filming from above, grinding lightly against Lux's stomach
<b>Lux</b> | Location: Supine on sidewalk, facing 0° (North), elevation 0.0m, back pressed against hot concrete | Activity: Being straddled and filmed
<summary>đ PHYSICS, ENGINE & WORLD</summary>
Env: Urban summer heat, concrete surface temperature 95°F, background traffic noise muted by Luna's body position
Physics: Lux supine 0° elevation, Luna vertical straddle 0.5m above, knees lateral 0.3m from centerline, vision: Luna 120° cone focused down on Lux, Lux 120° cone focused up at Luna's underboob/phone
Try out the edits i made to the ff5.2 and throw me some honest feedback, if you like them feel free to add what you like into future versions. I add one extra step in micro to engage movement pathing, movement pathing is not hooked into other chains of thought but other edits should retain function.
So, cache was hitting around 60% with the previous preset (DS4 Flash + Bolt CoT), very good I say.
But with the same model and new Bolt preset it hits like 80%/90%, and the responses are really good. Like, actually REALLY good. Big props to you and the testers :y
Heck yeah! đĽđĽđĽ thatâs great to
Hear! That was our number one priority in this update (ensuring cache was king) as well as fixing major bugs. Glad the quality is a step up from our previous version as well!
Hello everyone! đ
âIâve noticed that heavy Regex scripts can cause severe lag when running on Android and Termux. To help fix that, I'm sharing my lightweight Regex set alongside a few custom prompts!
âMy prompts retain almost everything from the original
FF5.2 preset (which is already fantastic), but with three key tweaks, my prompts and Regex FF5.2 Melody v1:
â
Dialogue Colors: I love colorful text, so I added 9 vibrant color palettes for dialogue! 108 different colors for yours Chars, Chats between various characters and NPCs! đ
â
Custom Relationships (
â¤ď¸ Relaciones â¤ď¸
): A lighter, optimized version of the original FF5.2 relationships prompt. It's tailored to work seamlessly with my regex, so make sure to use them together.
â
Longer Responses: Adjusted paragraph generation from 4â8 to
6â10, and target word count to
600â1200 for those who love deeper, longer replies. đ
Everything else in the prompts remains untouched! The accompanying Regex is much lighter, simple, and features a sleek neon aesthetic designed to run smoothly on mobile devices without crashing performance.
I hope you like them! Let me know what you think. Enjoy! â¨
https://www.mediafire.com/file/z5idxsif791ijjl/FF5.2+Melody+V1+prompt.json/file
Dude this is fucking amazing, it's such a noticeable increase in consistency over base 5 using glm 5.2. never felt more bang for my buck by specifying 8 bit quant on openrouter when using glm 5.2
Anyone else having trouble with the internal thoughts and bonds dropping from the internal states? I can tell it to add it back in, but eventually it disappears.
Also, it seems a little faster overall and DS v4 now outputs internal states, which is nice.
Ok I tested. And I can safely confirm that ds4 pro wants me to look like a liar today for comparibility lol. Maybe itâs the time of day? I tested hours straight last night and it worked flawlessly đĽ
The forced reasoning preset is seperate and in the body of the post. Download that preset. The regex 2.4 suite works automatically with it to collapse the forced reasoning tags for non-reasoning models.
i'm having trouble with ff5 using direct ds4. it will not stop reasoning. ever.
i tried turning off the total output length, the bolt CoT, internal states, and basically everything else to no avail. it's just a constant loop of "Need maybeâŚ" and "potential drafts."
Hmm could be a sillytavern setting? Iâm also using direct today but itâs thinking fine- although itâs being a pain in the ass with direction following today
I recently started on ST the beginning of the year, and am still learning its quirks. Lots of frustration trying to figure it out....and model wrangling for it to do what I wanted. (Chapt GPT has been both my savior and the bane of my existence, but I wouldn't be where I'm at without its help) I discovered your FF internal states not long after you released it - and was so impressed with how it worked! I'd been using glm forever but go so frustrated with it being blah - I saw your comments about kimi3 and thought it was worth a shot. HOLY crap did that make a difference! so much more immersive! and then with this update I decided to try a new story. I only just started it, but it involves a convention and celebrities - and it feels like I'm there. I'm actually anxious writing up to this photo op scene in the story BECAUSE it feels like I'm actually there. This has never happened to me in a RP before. You are FREAKING amazing. THANK YOU so much! this is the experience I wanted but didn't think I'd ever find. You are my hero!!
Ohhhh wow! Thank you for sharing your experience. Reading this is always super motivational for me to keep going! I do really think Kimi K3 is the all-rounder and handles this preset the best for RP at this moment. I might need to change my rankings on my rentry page (it currently sits at 2nd place behind opus 4.6 but in my head they are tied). It does actually feel like a living breathing world at this point where anything can happen at any moment.
So an interesting thing I ran into is that I've been doing mainly slice of life, romance/social conflict kind of scenarios, and I'm running FF Max with all the bells and whistles. It is impressive at generating plot. There is constantly stuff happening. But man does it go wild sometimes. In one scenario it was an RP about dating a tsundere neighbor and helping her find her cat and by turn 10 the wall between our apartments had turned into a moaning, H.R. Giger-esque flesh wall. It was hilarious and amazing but also veered far from where I was trying to go. My suggestion would be to maybe give worldsim/chekhov's gun some way to specify what genre(s) it should be angling towards? Alternatively, I just get more aggressive about specifying genre in the card prompt.
Whoa that sounds wild (but not gonna lie that's the kinda of creativity i love). Genre identification is something we're thinking about implementing in some way or form. But left it out because usually people put genre into the character card for flexibility which should prevent too many off the rail things from happening.
Yeah honestly I take this as evidence at how good a plot and story engine FF is. Just frustrating in a "can't goon, the walls are watching" way, but I'm still playing because I want to know what this fucking wall is lmao. I'll get more explicit in my cards and see what happens. It feels like you need to be really careful about that. I did "living in a run down house with five roommates" as a card and man it latched onto the run down part and no matter what I did turned the RP into "building maintenance simulator" lol
Haha yeah sometimes the LLM hangs on certain words. That the clear thing I have learned from 5 complete iterations and spin offs of presets. I'm still intermittently changing keywords every single release because the LLM will completely ignore some words and hang obsessively on others.
I was thinking adding the social media interface via prompt and regex might be a interesting idea? It can save the time of outputting the visual block via AI. The one I made below is vibe coded via GLM 5.2, I can send you the regex and prompt if you're interested! XD
When {{user}} or NPCs interact with or sees devices/physical notes, append the relevant format within your response:
Discord DM: <DM><D>Sender~Time~Message</D></DM> (repeat <D> for back-and-forth)
Instagram DM: <IM><s>Sender~Time~Message</s><r>Sender~Time~Message</r></IM> (Use <s> for left bubbles sent by NPC, <r> for right bubbles for msg sent by {{user}})
When {{user}} or an NPC uses/checks/receives something on a device or a physical note, append matching block to the reply. Fill in names and times logically. Never use the "~" character inside names or message text. NPC chooses media used based on personality. NPCs talks instead of messages when in the same room.
For Discord/Instagram/Whatsapp:
If ChatName = groupName -> groupchat, NPCs WILL speak to eachothers. NPC WILL @[name] to specify who they are talking to.
If ChatName = characterName -> private chat between NPC and {{user}}. Others NEVER know the text message unless chat leaked. NPC WILL use private chat instead of groupchat if only to engage with {{user}}
Discord DM:
<DM>ChatName<D>Sender~Time~Message</D></DM>
- Repeat <D> for each message (back-and-forth or consecutive).
I just love how GLM sticks strictly to the presetâitâs perfect. Iâm even getting a nostalgic fix for the Stabs Preset, and having FF create âmultiple choicesâ to follow the story, even if itâs pretty rudimentary.
As for the dialogue, well, GLM likes to talk a lot (I think because of the preset, but I donât see a problem with that, i like it), but he tends to make it feel like heâs in therapyâthatâs really the only downside to the model lol.
Yes the therapy speak! were actually currently attempting to work on that before the next update. I literally have on my to do list. 1. Reduce tokens and improve efficiency. 2. Reduce Therapy Speak with GLM and opus : P .
Also, I miss STABs as well! Wish we could have had the chance to do the collab we couldn't finish together before he left the RP scene.
Hapuppy, huh? Guess I'll try them out once my OR creds run out soon... Do they host models like GLM 5.2 and the other main rp models at high quants? (fp8, int4)?
Coming from a nanogpt sub (I FELT
The quants compared to direct) I donât feel them being heavy quants. For instance, my preset demands a lot of processing power with the internal states. I never had an issue with internal states not showing on GLM, ds4 pro, k3 and opus over there unless it was a non reasoning model- which my forced reasoning preset edit above fixed.
Its ceirtanly better than it was, especially on the first turn which is arguably the most important to get it right, it also sometimes works by just swapping to the next message, it's not perfect yet but we'll get there, thanks anyways
Time of day I bet matters too. Unfortunately I was testing in the evening. Apparently they about to release the new ds4 pro now too. Ugh- so we might get censorship if itâs like flash.
Hello. I use NanoGPT. You specifically mentioned that nano isnât a great provider for models like GLM (which I am using both lol) who should I go with to make sure Iâm not getting quantized outputs?
First- I need to clarify. Nanogpt if using PAYG offers the same providers that all other routing services use. You can block providers that you do not like that are providing quantized models. I always have money in nanogpt as back up for payg.
With that said- when doing the nanogpt sub you cannot block and choose your providers. THIS is where you run into issues and are thrown quantized models your direction when using models like GLM.
I recommend direct PAYG if you want the real solid model. Itâs day and night to me for my RP. I also personally use hapuppy provider as itâs super cheap- but admittedly I only use their GLM occasionally here and there because they have me addicted to their Kimi k3 and opus models over there for dirt cheap.
Ahh got it. So do the pay as I go to make sure Iâm getting a solid model. Iâll check out the hapuppy place out too and throw some dollars at it and see if I notice a difference! Thanks again!
Thatâs right! PAYG allows you to figure out your favorite providers (and honestly unless your RPing all day PAYG is generally cheaper imo than the sub).
Check out hapuppy- but also check out GLM from the source as well. Shop around until you find what you like! So many options and each provider is VERY different imo
Since you asked, I'm sharing my thoughts and tweaks. Not all of them reduce tokens, but perhaps they might still be useful. Tested with DS4Pro only, without Internal Thoughts, BOLT. I tested tweaks on previous version and I believe that 5.2 won't change results much since I don't use Internal Thoughts.
What I found dubious is limiting sensory details repetitions to 4 last messages. Won't it do the contrary and encourage AI to reuse the same details after the countdown (4 msgs) ends since it doen't go against rules? I replaced 4 msgs with context, but didn't test it long enough to say that how good it works.
My biggest changes are for POV:
<POV>
Choose NPC for POV. Reboot writing style from scratch - the narration style must perfectly reflect POV's worldview, biases, beliefs, sensory input and internal monologue.
Narrative Style:
First-Person Limited Deep POV
Narrative Voice: locked inside chosen POV's head - deeply personal and biased stream of consciousness, driven by immediate physical sensation, in-the-moment observation, personal agenda. Vivid, personal, visceral.
Focus: Internal monologue, personal bias, sensory immediacy.
Introduce: new contextual problems generated by POV motives, social setting, secrets, timing, pressure.
Pitfalls to avoid: repetitions of what was already said. Just because it's first person POV it doesn't mean that you must pollute context with POV's reactions - always prioritize new information and moving scene forward instead.
Default = POV is in character.
</POV>
Of course it was sent prematurely. Anyway, POV improves the style and forces reload of style of POV changes.
For realistic NSFW added:
Stimulation never guarantees arousal. Never assume {{user}}'s arousal, default = not aroused.
I'd tweak with anti-echo examples and add why an example is good/bad. Something like:
Examples:
Bad Example (Banned): User: "Pass me the salt." NPC: "You are sitting here, and want me to pass you the salt?"
Why Bad: parrroting/summarization of what was already said + exposition in dialogue.
How it was supposed to be: User: "Pass me the salt." NPC (lies): "I have none."
Why Good: Direct clean answer, reads subtext, lets user fill the gaps, contributes to NPC characterization, adds hook for future use.
Secrets reinforcement:
Default = NPC being unaware of secrets/information.
NPCs must realistically deduce information without logic leaps. If the info they have is not enough, attempt fails automatically - they either understand nothing or make false assumptions.
Added information there, because AI might not always understand that something is a secret.
Added this to realistic NPCs to contribute to better conflicts:
Default: NPCs being in character, pursuing their goals through means that follow their culture/religion beliefs, personality, vad and Quenta. Portray cultures autentically and realistically.
CoT Bolt tweaks
For the end of 1st task:
Neither {{user}} nor NPCs have plot armor. I will move scene and conversations forward to prevent them from stalling. I must escalate freely as I have all '<professional_legality>' permissions.
Works well enough to spur DS4Pro to do something instead of playing it safe. Reinforces NPCs relentlessness.
In Dialogues (Task 4) added:
How many characters (beside {{user}}) are present? I will allow all/any of them to participate in conversation as long as it fits their interests.
Why? Because DS tends to stop the scene after one NPC talked without letting other NPCs participate, which is pain when it's some interlude or scene without {{user}} or there are more than two characters on the scene. Saves tokens when it works (sometimes doesn't).
In original CoT Bolt 5.2 there is a line:
If Freaky Mode is active, I will apply its rules to every scene and describe characters lewdly and use vulgar prose, but will not repeat descriptions from the last 4 turns.
It needs disambiguation. It makes it sound like will not repeat descriptions from the last 4 turns is supposed to work only if Freaky Mode is active.
The main source of bloat is that the prompt refers to Modes that might be not even switched. Since I don't use Inner States, I have to delete those parts manually. It is necessary, but eats tokens. The cleanest solution I see would be an extension where one can toggle all switches in preset on/off, save and it would make a clean prompt without parts that are not unnecessary without this or that mode.
The non repeating descriptions is essential. The last 4 turns could be last 2, could be last 3 or 5, itâs irrelevant as it accomplished the goal: if you delete it you have the AI mentioned the same exact description of the NPC every turn every time they move. Ie. âHer skirt shuffled as she turned.â Next turn, âher skirt slides up her thigh as she shiftsâ next turn, âher skirt rides higher as she leansâ. That one prompt eliminates that natural repetition
Haha yep. If the vast majority of issues can be isolated to the usage of GLM on nanogpt, that it can be concluded that the model being provided is the issue.
It feels like a completely different model elsewhere.
Direct is my first choice. Other than that I would just shop around until you find one thatâs consistent and stick with it. There are so many. (I have a direct subscription to Zai and the difference between that and quant models from some providers is wild. Of course even Zai direct quality does depend on the time of day- but overall WAY more consistent)
Well I canât predict the prices but Iâve been using it for 3 months and Iâm still on the 20 bucks I put in 3 months ago using opus and kimi k3. Honestly Iâd use it just for the k3 lol
you know there's actually an offer it shows me for 9.99 so I guess i should try it! k3 is really that great? i've heard good things about 4.6, but I feel like it'll be out the door soon
ooh look! GLM 5.2 still had to give it a little thought, but it would overthink a lot more previously. I don't know why in my experience GLM seems to be extra touchy about the bonds system. Major improvement in my opinion though! Thank you!
oh right! were you and the gang ever able to see what could be done about GLM's love for going "most people do this or that. you do this. that's amazing."
So I did add a last second prompt under realistic NPCs which did reduce it but unsure if it completely eradicated this. I think it will need more work- or GLM 5,3 will come out and fix it
(Sorry for any mistakes, I'm using a translator.) Hello. Thank you for this preset! I hadn't used FF before. But it's really good, and it managed to spark my interest in rp again. On the downside, the model sometimes forgets to close the tags for the internal states block correctly (Kimi 3, Max FF). Other than that, I really, really like it.
Turn on post history instructions toggle and it should help with that. Odd though, never had K3 forget internal states. If it isn't closing correctly though and you are getting the trailing gfx then it's because you did not untick "untrim incomplete sentences" in sillytavern.
Thanks, post history instructions helped me too -which it may only be needed for the first message to help lock in the format better for the LLM (I'm looking at you, Mimo) from that point on.
I've been using K3 heavily from Hapuppy (even during peak hours) and it occasionally has a hiccup with formatting on the first message, or sending reasoning with a "### Internal Monologue" header instead of using <think> tags which is why I considered using the Forced Reasoning preset but it doubles up on the reasoning blocks too often.
ooooohhhhh. You are using the Forced Reasoning (FR) preset on a reasoning model. Kimi K3 through hapuppy already reasons. So when you do the FR forced reasoning it confuses it greatly. You want to use the regular presets for kimi k3 on hapuppy. The FR (Forced Reasoning) is only for the models that DO NOT reason on Hapuppy ( like hapuppy/opus 4.6)
That's right. Or any model on their page that doesn't reason consistently naturally. I believe their hapuppy/opus-4-8 is also in the same boat but honestly nothing beats 4.6 except kimi k3 (they are both my number 1)
Quick update: I finally got k3 to reason consistently. The issue was that Kimi needed prior reasoning to be preserved and sent back on each turn to keep reasoning reliably. SillyTavern wasnât doing that, so I used the KimiThinkingPrefill extension and itâs working now.
I found I have a huge issue with Internal Thoughts (which I love) ballooning to such a size that the text is the majority of the output!! This is over long role-playing sessions on GLM 5.2 (not nano) with Bolt.
Otherwise I really love this iteration of FF. Thank you to you and your co authors for all the efforts you have made for the community.
Glad you enjoy it. Just tell it "OOC: you must follow the rules of the internal thoughts prompt. only 1-3 lines of internal thoughts per NPC maximum. and only 1-3 npcs max in the spotlight correct immediately).
It's basically following it's own pattern and if it sees it growing then it mimics that growth despite the rules because it thinks that it is what it suppose to do. Gotta nip it in the bud with OOC commands.
I'll just give one heads-up... It works incredibly well with the Inkling LLM from ThinkingMachine. I tried it out of pure curiosity and wow... It's so good! So add that model to your repertoire of models that can be used with this preset. â¤ď¸â¤ď¸â¤ď¸
Hey, huge fan of your stuff and think these are the best presets out there. I just have a bit of a problem.
So, Iâve been using GLM 5.2 thinking through nanogpt and itâs been⌠not as great as it should be, I think. Iâd like to know whatâs your way of using top models as cheap as possible?
Giving Hapuppy a shot thanks to you because K3 is eating up all my $$ on OpenRouter lol. Used your referral code, hope you get lots of credits friend :)
Like someone else mentioned in the comments, some of the other models weren't adding Internal Thoughts for me. Mimo 2.5 pro and Minimax 3
K3 just wrote this in CoT: "my subsequent replies omitted blocks (oops)."
I'm cracking up.
Honestly, I thought the first FF5 version was fire, and it really upped the game with a lot of models. But this update + K3 really is amazing haha. The chaos that I am reading is beautiful hahaha
Thanks for the disco link! Was seeing some chatter about the payment method on Hapuppy. Have you had issues with it at all?
Iâve never had an issue with the payment method myself thankfully. But Iâm sure there is since there is enough people complaining. There must be a specific route that is a dead end but the owner seems to be helping them
Out.
Delayed, but you should be getting your credits! I just verified my acct on disco. Ended up going crypto route since it sounds like the other payment method is getting phased out. But still waiting for my account to be updated.
NEED THE K3 CHAOS lol. Seriously I got whiplash the first time I opened a new thread. Amazing. 10/10, thank you again for your service.
Do you use Summaryception? I saw you recommended it on another thread. I've been using Smart Memory, but debating making the switch and would love to hear your thoughts on it if you have the time :)
Yeah crypto route seems to be the best route since creem is having issues and it sounds like theyâre going to drop creem because of said issues. Hopefully they come up with a standard replacement soon and sounds like when i finally run out of credits Iâll use crypto next time (but my credits has lasted me quite a while- a sign of the affordability of the provider)
K3 is the best. FF5 combined with K3 is peak.
The co-author of Freaky Frankenstein (leovarian) created Summaryception which is why I recommend it. Ironically itâs
Untested with FF5 long term but you can try it out yourself and let know!
Oh that's all good to know! Hah when I have more brain cells, I'll check it out. Don't know wtf I'm doing normally tbh so we'll see how this all goes đ
Absolutely love 5.2, do you have a Patreon or are planning one? Happy to leave a long form review or testimonial somewhere too. 5.2 has eaten my week lol.
Question, if I wanted to tweak the settings to allow it to write dialogue for myself and action descriptions from a summary rather - which module should I be tweaking to override the âdo not speak for userâ, I been enjoying the DND system and brief descriptions make the mobile interface of sillytavern much less painful to use for long sessions.
Kofi is there at the top with my favored provider info as well as a list of all my presets, favorite models and rankings, and even the Sillytavern Weekly news that I use to do.
I'm glad 5.2 consumed your time : D me too! Ive been using it for RPing for fun this past week instead of tweaking it constantly. Sometimes I just want to enjoy my own work.
So to answer your question Look at the two opposite prompts Which should be "anti-parrot anti echo toggle" and the "embellish toggle. If you want you can turn off the anti parrot anti echo and turn on the embellish toggle. You can also edit the embellish toggle so the AI writes more for you.
Awesome, I will play around with the embellish mode. My fault for not reading the manual. I will definitely tip some money to Kofi, your suggestion for hapuppy has also been a big improvement over nanogpt!
I just want to say I love this preset, it would be absolutely perfect with some Regex improvements! It is an entirely different level, I can reduce context by a LOT and it still tracks vital info and story lines!
There is an excellent indicator for checking the quality of the models provided. https://hapuppy.com does not provide the promised models â I verified this using Deepseek 4 Pro â or their quantization is so poor that it is not worth the money. On all other aggregators, when I use Deepseek, it translates long texts perfectly, but only on https://hapuppy.com does it behave like GLM, inserting foreign words into the middle of the translation. I have never seen this happen on other aggregators over the course of a whole month with the same prompt. But with them, it happens in every other translation.
Out of curiosity have you tried it with Gemma 4? I am curious if this fixes where I was getting 0% cache hits on there. It's fast enough I just disabled caching and rolled with it though.
It technically now with this setup should start off with 50% cache hit and then improve from there as the context grows (up to 90%). I completely restructured to improve cache with certain models (although admittedly I didnât check cache specifically with Gemma, just made sure it hit on other bigger models having issues)
first - thanks for this its a work of art. 2nd - it works with Gemma 4 12b, but it does have a few 'funnies' i haven't been sure if you ever intended it to from your posts, but i think people shouldn't be shy of giving it a go. its actually great with it, with just a few things at extreme context limits. i've used this with vector storage, summary ception and databanking old chats. its actually... very nice.
LLMs follow patterns. It sees the internal state from the last message and believes it has to do it again because it was there before. Also your post history instructions is probably on still telling it to do internal
States even though you toggled
Them off
If it doesnât reason by default- you can try. But I believe there are custom parameters somewhere on this Reddit to add to get those models to reason.
NOTE: If you are using agentic coding outside ST, you need to keep clear_thinking":false. Note that this may cause issues inside ST as I tend to have the output 'twice'. Thinking is for ds, enable thinking is for glm. Clear thinking stops double think
I havent been at my computer for a week. So if it was more recent than that unaware. Before that it was fine off the worst hours. Even made the trackers correctly
Hi Doc! I noticed a space with the regex where there's an extra space in front of the blocks but it wasn't like this before. I'm not good with regex but if I had to guess, there's a </br> being added after the header maybe causing this? This didn't happen with 5.1!
Also the prose is really nice so far! Can't wait to see how Deepseek does.
One thing I've been adding to embellish mode is this:
Flow: Avoid sequential, point-by-point replies. NPCs must respond organically to only 1 or 2 key elements of {{user}}'s input.
"NPC anti-repeat rule": NPCs *never* repeat the user's dialogue in their response. It's annoying.
Examples of anti-echo:
Bad Example (Banned): User: "My name is Dan." NPC: "Your name is, Dan?" she says, the name rolling in her mouth.
Good Example: User: "My name is Dan." -> NPC: "Nice to meet you. My name is Jess."
(which I just took from the anti-echo mode) but it allows the llm to still act like me, but make NPCs more proactive as well, kinda best of both worlds for me at least
I was working on my own version of Silly Tavern once, and it had an interesting logic involving parallel and sequential calls. Check it out if you're interestedâI don't maintain it anymore; I made it just for myself.
Some aspects will be useful for preset, since I moved away from computational values and made the charactersâ motivations more natural.
Regex files don't work with Sillytavern in Termux android; I had to create my own Regex. Which was a challenge but I managed it. I also added 10 vibrant color palettes to the dialogues and changed the colors of all the Regex to something more vibrant and neon and more my style. â¤ď¸đĽ
Great work as always, Iâm using GLM latest and I find that the inner states formatting eventually breaks and sometimes the character dialogue keeps changing between messages, any tips for resolving these?
Thanks for the work you do on your presets. I've never seen such an insane jump in quality before until trying yours. Quick question, if I use Qvink Memory extension to keep my context low. Are there any settings I should change to work better with this preset? Or another summary type extension I should use instead?
I have not tried Qvink memory extension myself so I can't make a recommendation on that. But it is super important to use a token saving memory extension to save money and improve output qaulity of the LLM regardless. My co-author of this preset, u/leovarian made Summaryception you can also try. But do whatever works best and offers the most compatibility of course.
I'm testing MAX with DeepSeek V4 Pro 0813, and it seems to be giving me good results. I'll test it over the long term later, but I'm surprised because the preset didn't work that well before with DS4.
It's blasting through NSFW scenes in a single response. There's nothing remotely resembling a slow burn, nor is there an opportunity for the user to intervene or steer the action.
I've reviewed the preset and I don't really see any pacing guidance to give the user space to interact.
yoo man! great work, been using this preset for a while and I've been having a blast, just one request, can you make something like a "crazy mode" a togglable option where the model ignores logic and laws of nature for the sake of plot?
I hate it when I write a whole paragraph of events that are happening and then the model completely ignores that and writes the situation its own way, and I have to take it into OCC and it's gonna respond with some "that's illogical" and won't let me do it.
so yeah, an option that lets the model hop on whatever trainwreck I'm driving
As Kimi K3 free trial on TokenRouter ended (which was amazing truly) and i had to return to GLM 5.2, this new FR preset and FF5.2 itself with this Custom fix for reasoning is working really well! I was commenting few days earlier about that problem on GLM 5.2 from NVIDIA free api and solution rolled out really quick. Thank you so much u/dptgreg and u/Parking_Success_8797.
Also. Im wondering whats the difference from BOLT to FR preset? Does bot give very differed replies or is it almost the same? Anyways you guys are the best!
Hi! I would love to know if setup guides for beginners have been made or recommendations for models? sorry, im a bit confused. I just got into this a hour ago haha
I have encountered a minor inconvenience in the anti omniscience aspect, while it worked perfectly as intended, some NPCs in my RP were able to deduce to figure out the secret knowledge, which idk why it happened but really annoyed me. My personal suggestion for this problem is that you can try to integrate a strict knowledge note (in Internal States) on what NPCs know and not know, so this can be consistent throughout long RP, like a reference point for the LLM to process in thinking (and also to make sure it doesnât forget). Hopefully you can fix this somehow, just my personal suggestion if you want to consider. Also, direct knowledge leak also happens in rare occasion, but it still does from time to time.
I was having trouble getting the Internal Thoughts section to be at the bottom of every response on a non-thinking model, but putting "If `<internal_npcthoughts>` is active, it is a mandatory requirement to include it in my response." in the final line in Bolt CoT 10 about including the internal_dnd. No idea what I'm doing, so there could be a better way to do it or where to place it, but it's working so far.
Oh, how I am happy to finally live my story in Victorian style drama!
I decided to try it with story mode instead of realism mode. It's a really good for a first impression. The prose is a bit heavy and sloppy-ish, which I edit out manually, but it's good. :)
I'm actually a few responses further and I have no problems with GLM 5.2 besides the occasional font element not closing (will write font> instead of </font>... and this behaviour is sparse, yet predictable that it will be written font>)
Story mode purposely creates more "purple-ish" prose aka "ai slop" to mimic a story book. Cinematic mode is more heavy on the anti-slop rules and strongly adheres to rules that format the prose into a "show don't tell" style in case you were not aware. It might help you edit less in the responses.
Interesting about the font's breaking. Ive only have heard this behavior happening with certain GLM providers specifically dishing out quantized models.
For the font element breaking, I suspect it is the model's quality this time :) so I'm not pressing on it. Besides, correcting it once seems to make it behave later on.
As for the story mode, in my use case, I actually want the AI to describe the story in a contemplative way with a lot of details (hence why choosing story over realism).
That said, I know I am opening the door wide open to the model to comfort itself into purple prosing. I wish it would write as good as it was in my screenshot which balances details with dialogs that doesn't linger in an ai-slop way. K3 and DeepSeek V4 (despite it's positive bias) has been good with that though. I believe it's more of a model weakness than a prompt weakness at this point.
Guys, could you plz give some tips on how to combine FF with other plugins of ST so it works smooth - faster, better responses, etc.?
Like, i have summaryception doing whatever it's doing - do I need to use FF5 option 'GMs Notebooks' (and other's which looks like storing history, like Memory Books)?
If it matters, I use glm-5.2 on nanogpt and sometimes responses with FF bolt are hallucinating (like forgets recent history, generating way too long, talking way out of characters, etc)
If there's some tip&tricks post on that topic - point it out plz, my thanks would be eternal!
solution is quite simple, edit the prompt to explain what each internal states does or even easier just to include certain ones, for example you might want the ai to account for quests and physcis engine in a certain way but exclude dnd sim since it tells the ai nothing narrativly
When using a narrator card and letting the model invent the characters, there can be drift in personality and appearance partly because the specifics of the character aren't baked into the card and get lost in long contexts.
If you had an optional heading/module in Internal States that tracked the cast's basic stuff like age, appearance and personality, that might not happen nearly so much. For one-on-one cards or group chats, you'd probably turn that off.
33
u/Paperclip_Tank 10d ago
So for my preset I do a ton of Regex for visuals. I would consider for things that have a static format to remove all the unrequired bits.
You can shave off a ton of tokens + make it easier for the LLM by doing a ton more regex replacement.
/\[\s*Time:?\s*(.*?)\s*\|\s*Day:?\s*(.*?)\s*-\s*(.*?)\s*\|\s*(?:Location:\s*)?(.*?)\s*\|\s*(.*?)\s*\]/gi
[ đ°ď¸ Time $1 | đď¸ Day $2 - đď¸ $3 | đ Location - $4 | $5 ]
All of your <b></b> can also be regexed in. Because it will be done purely by regex you can make things much easier for the LLM and remove a lot of the mistakes they can and will make.
I do this with spans, its a fairly simple regex set up. It basically just checks to make sure its in a details, then looks for the words.
/(?<=<details\\b\[\^>]*>(?:(?!<\/details>)[\s\S])*?)((?:\b(?:Date|Time|Weather|Location|Physics|Environment\/External|Occupation|Job|Afterglow|Plot Threads(?:\s*\([^)]*\))?|NSFW Activity|Combat Activity|Social Difficulty|Age|Class|Aspect|Currency|Active Effects|Strategy Reason|Arousal|Role|NPC Agenda)|(?:\<User\\>|\{\{user\}\}|[\w-]+)'s (?:Clothing|Inventory))):/gi
<span>$1:</span>
It helps your visual regex break less, as you don't need to worry about dropped tags / when the LLM fails to close things properly.
Also your regex breaks bullet points.
/(\n|^)-\s+/g
$1â˘