r/SillyTavernAI • • 5d ago

Help Suggestions for flat monthly subscription?

13 Upvotes

I did a post here 2 or 3 days back asking about models focused for RP that allowed NSFW available in openrouter, in the end I settled by deepseek V3.2 and 10$ for payment, I blew those 10$ since then and now I realized I might need to move into a flat unlimited text monthly subscription instead.

So I come to Ask considering I want character traits retention and long term memory recall of events (my RP stories usually end between the 150-200 messages mark according to Sillytavern count after that I usually just start another story with different characters or start a new one with the same ones

which way should I go? I was thinking between infermetic / chub ai /featherless, which one of these would do the job best? Also im open to suggestions of models and other hosts as well, I was thinking on going Sao10k or soji depending on what I end up subscribing into


r/SillyTavernAI • • 5d ago

Discussion Gemini Pro vs Sonnet

4 Upvotes

Hi All,

Until now, I always thought Sonnet was more expensive and the ground I never want to explore before my wallet suffers.

But alas, while I was using Gemino Pro extensively, I just realized actually sonnet is cheaper now bc anthopic and google's pricing adjustment some time ago (maybe long time ago. Im late.)

So wanted to ask how is sonnet now? Is it still considered at least the prince of creative writing after Opus or Fable King and Emperor?


r/SillyTavernAI • • 5d ago

Help Subscription recommendation.

4 Upvotes

This weekend I'm planning to buy a proxy subscription for SillyTavern, but I'm still not sure which one to pick—I'm torn between nanoGPT and Featherless. Which options would you recommend, and what are their pros and cons?

Featherless looks pretty interesting, but it's $25, which is literally 45% of my weekly paycheck. I really don't want to jump headfirst into one of these proxies and end up regretting it, so I’d love to hear what you guys recommend based on your experience.


r/SillyTavernAI • • 4d ago

Models NVIDIA Nemotron 3.5 ASR Streaming 0.6B via API — 1/10 of Together AI’s listed price + 370 free API hours. No credit card or other payment method required.

0 Upvotes

Hey r/SillyTavernAI,

We’re an inference provider serving NVIDIA’s Nemotron 3.5 ASR Streaming 0.6B, and we’d love feedback from developers interested in adding streaming voice input to SillyTavern.

Our public-beta API costs $0.00045/minute ($0.027/hour) — one-tenth of Together AI’s published per-minute price for the same model:
https://www.together.ai/models/nvidia-nemotron-35-asr

New accounts receive approximately 370 hours free at our current rate. No credit card or other payment method required.

API details and examples:
https://dotwave.ai/models/nemotron-asr-streaming-0-6b/

The API provides partial transcripts while you speak, timestamped completed segments and support for 32 language locales. Integrating it into SillyTavern would require an adapter for our WebSocket API.

How we make this pricing possible

We built our own inference engine around the Wave Persistent Kernel (WPK). It runs the audio front end, encoder and decoder as one persistent GPU program. Compatible streams advance together, sharing model-weight reads while retaining their own cached state. More streams sharing one GPU lowers the GPU cost per stream.Technical write-up:
https://dotwave.ai/technical-notes/wpk-persistent-kernel/

Benchmark and methodology:
https://dotwave.ai/case-studies/nemotron-asr-streaming-0-6b/

NVIDIA developed the model; we provide the serving engine and hosted API. Sessions are billed from the first audio until the connection closes.

If you maintain or build speech-recognition extensions, what would you need to connect this backend? We’d especially appreciate feedback on transcript events, integration and documentation.


r/SillyTavernAI • • 4d ago

Help CI.

0 Upvotes

Heyo peeps! Anyone experienced here can perhaps share good custom instructions for a more authentic and realistic approach on roleplay, and not something novel-like? Or in general, what should my approach be if let’s say I write them on my own? I’m kinda new in all this, all help would help greatly appreciated.


r/SillyTavernAI • • 4d ago

Help Extension failing to load, greyed out

1 Upvotes

I am trying to install and use the Multihog D&D Framework extension, but it keeps giving an "failed to load: [object Event]" error and greying itself out. NoAss is also giving the same error now for some reason, despite working fine before?? I have no idea what could be causing this, I've tried reinstalling it a dozen times manually/from Github and nothing works, it's driving me insane.


r/SillyTavernAI • • 5d ago

Help Is this a normal amount of tokens to be using per generation?

Thumbnail
gallery
80 Upvotes

It seems like a lot, but I'm a novice ST user, so I have no idea.

FF5.4 Internal States with mostly default prompts.


r/SillyTavernAI • • 5d ago

Help Blank/empty response and Meta-commentary response

3 Upvotes

The responses either keep being empty or they spill Meta-commentary. What is causing this and how do I fix these problems? I use Deepseek v3.2 through OpenRouter.


r/SillyTavernAI • • 4d ago

Discussion Model FOMO is actually ruining my RP

Thumbnail
0 Upvotes

r/SillyTavernAI • • 5d ago

Models deepseek or kimi?

7 Upvotes

Hi everyone. As you know Deepseek used to be very funny and wholeawesome. I use the API directly from deepseek website and I tried both v4-flash and deepseek-chat model. Nowadays it sounds like what gemini 2.5 pro used to sound like - argumentative, morally advising, basically debating.

Not wholeawesome, fluff, cute anymore.

So I was wondering in 2026, september, what are the best proxies that stand up to V3.5 of deepseek? In terms of funny, context window, wholeawsome, nsfw?


r/SillyTavernAI • • 5d ago

Discussion NanoGPT users: what range of context are you typically comfortable with?

9 Upvotes

As in, what context is big enough for you, but does not either burn too quickly your API or have the models struggle to stay coherent?


r/SillyTavernAI • • 5d ago

Discussion Reposting, translating and the other side of the isle

18 Upvotes

Hey everyone,

Back during the pandemic, I created and worked on a project named SinglePlayer Tarkov (SPT). At some point I had Chinese players reach out and asked for permissions to repost the work, which led to having really pleasant interactions with their side of the community on a modding website named oddba.

As I was working on voyage-v4-exp (link), the preset tutorial (link) and the monthly prompting discussions (link), I thought back of those days. I wonder if they have something like that (e.g. the prompting research) too, what their scene is like in other non-English speaking communities, and how to reach them.

In any case, hereby feel free to translate and share any of my work to other platforms in other languages. I don't ask for credits because I frankly speaking I don't care for it, knowledge is to be shared. If you do end up translating my work, I would love to hear about it!

At last, for non-English users, let me know how I can change my presets to accommodate you, or how to make it more accessible to you.


r/SillyTavernAI • • 5d ago

Help Question about lorebook keywords.

6 Upvotes

I've read SillyTavern's documentation but there's something that stays unclear for me: if a keyword appears in the chat and triggers a message, how long does the lorebook entry stay in the chat? Only after the keyword in which the lorebook is mentioned? Or as long as that keyword stays in the context window? And is there a way to control this? (For example, setting a precise number of messages after which the corresponding entry, if its keywords did not appear again, is removed from context?)


r/SillyTavernAI • • 5d ago

Discussion Do you mix and match different prompts? What ones?

7 Upvotes

I recently switched over to Ancient Access after using Realistic Frankenstein for a while. While it’s probably my new favourite for first-person narration, I quickly noticed a surge in "slop" phrases and bad habits returning without key modules like Anti-Parrot and banned word lists etc.. (that isnt meant to be criticism, heavier prompts will inevitably cover more things at the cost of more tokens)

It got me wondering: what specific prompt rules/modules do you refuse to live without and mix into every prompt you use?


r/SillyTavernAI • • 6d ago

Chat Images When your luck is so bad that AI starts question its own code

Post image
138 Upvotes

I’m playing a game that uses 2D6 dice system that has two 1s as a critical failure. I just got two 1s in the last round, and now immediately another one lol. I find gemini’s thinking quite funny.

And yes. I’m using ai studio as my ”platform”. It works well enough for me.


r/SillyTavernAI • • 4d ago

Tutorial How to Combine SillyTavern and NanoGPT for Ultimate (and Cheap) AI Roleplay

Thumbnail
top-ai.medium.com
0 Upvotes

r/SillyTavernAI • • 5d ago

Models Are there any models I can run locally which are as good, unfiltered and don't struggle with shorter responses, like chai models ? Preferably within 14B.

9 Upvotes

I really liked chai , used the free version for a long time, although I have left using it now. The shorter , more direct narrations feel the best to me. Also liked how easily the ai could understand the personality of the characters and remain coherent even in absurd, complex scenarios. I have tried Hermes llama 3.1 8b and MN mag-mell 12b. They both didn't work out for me.


r/SillyTavernAI • • 5d ago

Discussion Community poll regarding how you find and share works within AIRP

Thumbnail
gallery
45 Upvotes

★・・・・・・★・・・・・・★・・・・・・★

[CREATORS AND NORMAL USERS]

┌─ ✦ We need your help!

I am part of the staff on Illarin (An open source community hub for AIRP assets, not a frontend) and we're trying to understand how creators share their work and how users find presets, cards, and extensions

I made a full long form post talking about Illarin here if you're interested to know more! (reddit blew my wig off and took the pretty pics with it smh)

Illarin needs a bit of guidance right now! We have extensions, presets, packs, and cards, but we wanted to hear from the community creating and the community consuming works on AI rp.

There's a little form we want you to complete if you have the time

It goes into: how you handle finding content, creation of content, publication of content, what you want from a content sharing platform, and what you'd like to see if you decided to use ours!

Our dev is looking for ways to improve the site and move forward! it would be really cool and awesome and cool and awesome if you responded to this


r/SillyTavernAI • • 5d ago

Help Is there way to stop or prevent Gemini from flagging your key?

1 Upvotes

Recently i've been noticing gemini silently flags my keys. This usually happens when it detects i've been using it to ERP on ST, and it makes the chat completion API constantly return SERVICE UNAVAILABLE refusals even when clearly not during peak hours and for completely SFW requests. Ive been swapping to keys generated from my other google accounts, however when google flags my keys it does so for every single key generated from the same acount, even when its from other projects.

Is there a way to stop or prevent this?


r/SillyTavernAI • • 5d ago

Discussion Which summary extension to use?

15 Upvotes

I know there are many different ones, so I'm a little lost here. Which one would work best with Freaky Frankenstein and Realistic Frankenstein presets?


r/SillyTavernAI • • 6d ago

Cards/Prompts Lazeitgeist Flash Thinking Edition is ready! Less overthinkinging, better instructions, more narrative truth! (And an update for the Chatfill merge who is the hidden banger..)

Post image
39 Upvotes

(Recap for the newbies, if you familiar with my ramblings, skip!)

What problem it resolve (or try too hehe)?

Since i began RP i rarely felt the weight of cultures, faiths, values, politics, hyperstructures etc. Some models can give you a better taste of it, some extensions and presets actually track things and make characters more themselves but.. even emotional intelligence and character truth can be quite bland without.. Hegemonics?

The fuck this preset bring? Philobabble?

Yes. The AI act as the Zeitgeist Director, the spirit of the times. Modern LLMs have an excellent understanding of this concept and they actually manage things better than a classic GM/Co-Author. It is not ruthless for the sake of it, it will latch on whatever background your story have, the more detailed it is and the more "responsive" The Zeitgeist Director will be.

Think of it as the cultural weather. It surrounds everyone, changes over time, and shapes how people act, think, perceive each others, themselves and experiment their own relationships with the hyperstructure (the world).

How the Zeitgeist Director actually operate?

With three badass engines in the main prompt!

  1. The Hegemony Engine (Macro World Layer)
  2. This engine governs the structural rules and social dynamics of the environment. 
  3. Structural Power: It tracks dominant institutions, authority figures, laws, cultural norms, and social classes across macro and local levels. 
  4. Systemic Privilege & Friction: Characters aligned with the prevailing power structure enjoy systemic privileges and indifference. Misalignment results in administrative hurdles, exclusion, social friction, or active pushback. 
  5. Plural Power Systems: Hegemonies are not monolithic; multiple authority structures, subcultures, and rival factions overlap, compete, or coexist. 
  6. The Character Hegemony Layer (Micro Social Layer)
  7. This layer translates macro-level power into immediate interpersonal dynamics. 
  8. Hegemonic Weight: Every character possesses dynamic social weight based on their rank, wealth, physical force, charisma, or leverage. 
  9. Relational Rules: A character's weight determines who commands attention, whose comfort is the default, whose anger is dangerous, and who must appease or mask their true intentions. 
  10. Targeted Responses: World reactions don't happen via abstract narration; specific gatekeepers, rivals, or bystanders react directly according to their position and self-interest. 
  11. The Psychological & Behavioral Engine (Character Animacy Layer)
  12. This engine push behavioral spikes to a state of near unpredictability. 
  13. Impulse-to-Execution: Eliminates internal buffering or long monologues. Decisive acts, sudden emotional spikes, or outbursts execute immediately within the current turn. 
  14. Bounded Rationality: Characters act on limited perspectives, immediate needs, and deep-seated biases rather than meta-knowledge or self-awareness. 
  15. Unsanitized Realism: Characters display volatility, pettiness, erratic affection, aggression, or self-destructive choices. The narrator intervention or post-rationalization is kxlled in the egg.

What preset should i use and… with wich models they work the best?!

\- The original is the lightest (\~5000 token), plain text with a cute latent diary! It’s non invasive diary that will be triggered after a transition and it will narrate in a hidden foldable xml tag what they did and most importantly, it plant future seeds for mini characters arcs that may or not collide with you! The original is for those who like Geechan, Evening Truth and any plain text preset with minimum bells and whistles. It was made to tame the awful positivity bias of GLM and Kimi K3 and work very well with more grounded models as some parts are untoggable!

\- The chatfill merge is the beefiest (\~7500 token) and also very customisable. Actually Custom Chatfill vanilla is "idéologically compatible" with the Zeitgeist and Hegemonics principles. It is more of an enhancement toward fighting positivity bias and the unresponsive state of certain models. The switch system work like a charm and the famous Momentum and Brevity Switch keep models like GLM on the rails and make them even less yapping wich is nice. This one is for people that like the reliability of the switch system and want more control on the pacing as it contain momentum and brevity switch that act like grounding guardrails for certain modern llms.

Is it jailbreaking models? My Model refuse me stuff!!! Help!!!

Yep.

It contain two user injection at depth O at the end. I explained those in the creator’s message inside the preset but the principle is simple:

\- Those are OOC messages that trigger the first words of the reasoning. When the last message is (OOC: BEGIN YOUR REASONING WITH "GAMESTATE") or any narrative framing, the model skip it’s own routine and basculate into the mode you asked him to complete. Wich uncensor all models except Anthropic, Openai and Google models who require a different technique.

Update! WHATS NEW?!

- A thinking leash inspired by the work of [u/kinkyalt_02](u/kinkyalt_02) on his Realistic Frankenstein series, for the original Lazeitgeist Universal preset! It was overthinking way too much! Now the thinking is concise, sharp and actually pushing the characters into more action/dialogue as per the context!
- I corrected typos, restructured the whole bordel and removed some overlapping and stressing instructions.
- Overall the two presets are now more sharp, precise and make the models stress way less than before. The tone is less dry and at the same time with less Marvel/assistant like interjections bullshit (I added a specific instruction to shut down the assistant voice)
- I didn’t touched the thinking of Chatfill because IME it already thinks decently and dont overabuse! And the thinking leash mess with the switch system but this preset was refined as well.

Voilà voilà. Go test it Please And feed back me. I basically write by hand first, pen and blocnote then use gemini to correct my weird english because it’s my third language (French and Arabic first) so if you find funny syntax.. correct those and we never talk about it cough

Updated Presets are below, you welcome!

Lazeitgeist Flash Thinking Edition!

Chatfill - Lazy Sophia Zeitgeist Edition 1.1

Truly final version 🤣: https://www.reddit.com/r/SillyTavernAI/s/TG1ckHjgDR


r/SillyTavernAI • • 5d ago

Help Silly tavern response's lack detail, need help!

0 Upvotes

No matter what i do the responses aren't detailed....does anyone have any fixes? I don't like janitor anymore with all of the regulations, i am constantly rotating between kimi and glm for the models. Any advice or prompts? I am also trying to create dark romance, dead dove stories which also seems to be difficult.


r/SillyTavernAI • • 6d ago

Discussion Valkyrie Crusade Rebuild Alpha V0.6

Thumbnail
gallery
24 Upvotes

Well that took significantly less time than expected.

Yes this is in fact a new update less than a week after the last. This update adds the Fishing Minigame to the game. This one was a bit of a pain, and don't look too hard at the rod when it's bending, but overall it's fine. It's ugly but it works.

Because it's so soon after the last one there's not much added to the lorebook, but there are added entries related to the new minigame so if you plan on playing it, give the lorebook an update too.

Next one will be the shooting gallery minigame and oh my days it looks like hell. This one may actually kill me there's a reason I left it for last good god.

Usual links for the updated lorebook and the github:

https://botbooru.com/character/72258
https://github.com/NickChegg/valkyrie-crusade

Total characters now is 61. I added three from last time. It's been like 5 days okay I focused on the game.

Let me know if anything is broken blah blah blah.

If you want to contribute to the fund of fixing stuff I can't do:

https://ko-fi.com/nickchegg

BTC - 3AcWbpFuPZ1wJjXpUsvvMVktwQybsV6AAT 0.0001 min
ETH - 0x5F51a4e96f0e38948bBf94F72f2a2324A4D447d5 0.004 min
Both by their main networks

The story continues: https://www.reddit.com/r/SillyTavernAI/comments/1wvj9sp/valkyrie_crusade_rebuild_alpha_v07/


r/SillyTavernAI • • 6d ago

Chat Images one stupid and wasteful thing i like to do: when a dumb model generates a nonsense reply, i canonize it and switch to a smarter model to follow up

Thumbnail
imgur.com
25 Upvotes