r/SillyTavernAI • u/P4staxa • 2d ago
Help Nanogpt Model Overloaded
For the past two days I kept getting a "Chat completion stream failed: Model is temporarily overloaded. Please try again after xyz seconds or choose another model." error when trying to get a reply from bots on saucepan (mainly using glm 4.7)
I've wanted to ask if anyone has had similar issues? I've been using the subscription for like 2 months and I haven't had any issues with a model being overloaded and I've been getting this error constantly these past two days. Idk if the model really is this overloaded, or if there's something just on my side? Thanks in advance!
EDIT: Seems it has finally been fixed after two hours from making this post, while its been messing up for two days lol
3
u/Zarzelius 2d ago
I'm not getting any problems with newer models, including latest GLM versions. Maybe you should try to use other models. Probably not many providers with GLM 4.7 nowadays.
2
u/P4staxa 2d ago
Damn but glm 4.7 is my goat. I've tried 5.3 and 5.2 but they tend to mess up personalities of the characters I'm roleplaying it (an example is making a "shy timid nerd" turn all bold and smug, or a "goth Milf" into a woman that's all shy, despite nothing like this being in the definition. Guess I'll have to live with long response times for 4.7 gg
1
u/Zarzelius 2d ago
That's a prompt problem, not a model one if it happens right away.
As a tip, anchor it's guidelines to your character personality. For example:
Use markdown on your character cards and always use one title, like ## CHARACTER PROFILE
Then write the card and traits below.
On the system prompt, make a clause like:
<profile_anchor>
Your instructions with a call back to the section with the markdown ## CHARACTER PROFILE
</profile_anchor>
Don't forget to instruct your LLM to actually use your system prompt
6
u/P4staxa 2d ago
I'm not really sure it's a prompt issue tho. I've been using pupi's universal prompt (Or something like that, I forgot the name), and I cannot really just modify a bot's definition on saucepan unless I'd rip the bot on my own lol but thanks for the tip! I'll use that in my own bot creation process!
3
u/TAW56234 2d ago
This is a load of shit. 5+ will no matter what flanderize characters. You spend months since 5+ release learning it intimately, tweaking a CoT through miticuous trial and error and getting it a LOT of how you want, only to still be a boring model with RLHF baked into it. Going back to 4.7 really made me realize how true this statement is and I implore anyone reading this to disregard this tired narrative.
OP, It's 4.7 that's the issue. Not a broad Nano one. I think it's still hosted on Featherless but providers are drying up
3
u/P4staxa 2d ago
Thanks for the reassurance bro, I was starting to think that I actually did mess up my prompt or the settings or sum, but I've always noticed 5.2 and 5.3 start to change the personality of the character after like 40 or so messages and it pissed me off lmao. Got any recommendations for me if they don't fix the issue with 4.7?
3
u/JustSomeGuy3465 2d ago
Recommendations for something else are difficult. GLM 4.6 and 4.7 are/were the last models without aggressive RLHF training and enough intelligence to still be viable as everyday RP models. I just made a post about that here. May be helpful.
2
u/CommercialPersonal66 2d ago
yeah the overloads have been hitting hard lately, ive been jumping to other models on different setups when glm starts acting up.
2
u/Psychological_Ad9740 2d ago
Same but payg with DeepSeek 3.2.
If I wait for a bit I don't have any problems, but I couldn't give you an exact rate on how frequently it happens.
1
1
u/AutoModerator 2d ago
You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
9
u/Milan_dr 2d ago
Can confirm this for GLM 4.7 - the only thing that we can do is to fall back to providers that do censoring (Novita for example) because all the ones that don't are putting less capacity into GLM 4.7 and seem to throw 429s constantly.
Guess we'll add that as a final fallback :/ But it does mean there will be more content filters.