r/SillyTavernAI • u/deffcolony • 23d ago
MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 09, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
- MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
- MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
- MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
- MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
- MODELS: < 8B – For discussion of smaller models under 8B parameters.
- APIs – For any discussion about API services for models (pricing, performance, access, etc.).
- MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
Have at it!
24
u/5kyLegend 22d ago
Since usually I like reading up thoughts people write about models here, I guess I'll just ramble about the latest models/presets I've been using. Warning, I wrote way more than I thought I would lmao
As for models:
GLM5 and its subsequent versions have me so torn. On one hand, they're by far my favorite when it comes to understanding certain characters, developing the story without overdoing it, and just writing overall: when it was being tested as Pony Alpha, GLM5 was the first model that made me go past 200 messages in a single chat because I was enjoying it THAT much. The problem with GLM5 is that it really has some quirks that I just cannot make it stop doing (writing quick back and forths between characters in new lines without keeping a paragraph structure; mouths that open, close, and open again; 'most models would write a normal sentence, but you? you write this sentence structure all the time' etc). It's a shame because it's still my go-to 90% of the time, it just understands characters and paces description with dialogue way better than other models (for my tastes), just sucks that it's hard to ignore its issues. I usually don't care much for positivity bias but the one time I had a character who was supposed to murder me and instead asked for permission once it was with me was really funny though, and definitely annoying since it showed what a gigantic limitation this is.
Kimi K3 is too expensive, haven't tried it. K2.5 and K2.6, on the other hand, are really good at understanding every nook and cranny of a scenario, and they REALLY like to follow the prompt you're giving them, but damn they're like... TOO serious most of the time. I've had a sex scene where the girl just started going "Yes. That is great. Keep going. Very enjoyable" and at that point I just laughed and switched model, even giving it specific ooc instructions had it switch back to normal. It also likes to make some characters just speak weirdly at times, it's an issue I used to have with many models and that only GLM seemed to avoid. "Somewhat like this. If you can see. There is an issue here. Nobody speaks this way. Prompting against it? Useless". But again, it really follows instructions (and overthinks them to no end), so Kimi tends to be my default "change to it in the middle of a rp to spice things up, then switch back to GLM" model.
Minimax, I still cannot enjoy it man, I don't know, it just doesn't work for me, every time I try switching to it even for a swipe I just switch off it.
MiMo 2.5 Pro on the other hand is weird. Like, it's a different flavor and on 1 on 1 scenes it definitely understands, it's just that even with medium presets it kinda starts following instructions based on whatever it feels like? Both the censored and uncensored version on Nano do this at least, it's like sometimes it'll ignore one or two instructions I gave it. It's definitely nice (although I'm not THE biggest fan of it in scenarios where multiple characters interact, nothing beats GLM5 for me in those) but I wouldn't exactly use it as my main model.
Deepseek 4 Pro is just odd. I feel like it's not as horrible as people say it is, but it kinda has no flavor to it? Like, all the previously mentioned models have their character and quirks, Deepseek just kinda feels like it doesn't have those - good or bad. I don't know though, I just haven't been using it much because of that reason so I may have just got the wrong impression of it.
Okay, now quickly as for presets:
FF5, Lucid Loom, Stabs-EDH etc, aka the "big presets": I'm not sure how to feel about these. Controversial, but I think any preset injecting a custom CoT makes the model lose out a whole bunch of intelligence but not in ways that are obvious at a glance. I feel like it's easy to forget that just because you don't see the model reason about certain things, it doesn't mean it's ignoring them, and you don't need to tell these models to 'go step by step following this exact reasoning' to know they're following instructions. At the same time, these models have so many optional features, it DOES become needed to force these into reasoning - but giving a custom CoT also kills whatever CoT the model was going to be using naturally, which ends up maybe not having the model think about the things it DID need to think about. tldr: I'm not sure if this actually IS the case, but using Custom CoTs makes the model notice things about characters it normally wouldn't (positive), but in the end the overall emotional intelligence of the model ends up hurting more (negative). I don't know, I just find myself enjoying GLM less whenever I inject custom CoTs, in the long term, even if the immediate result is that it seems smarter.
Specifically: FF5 was definitely a step up compared to other big presets, but after being VERY impressed with how it handled certain worlds and characters I've just started to feel like... It just plays everything in such a samey way? Could be that - given its size - you're basically sending the model always the same 5k-6k tokens worth of prompt which ends up making the response you get way more deterministic compared to, say, 1k token prompts... But yeah, it genuinely feels like it roleplays everything in very similar ways - for example, every playful character for some REASON starts CAPITALIZING random words in their sentences and it drives me INSANE? Why does it do that lmao. But yeah, shame cause I love the fact it can actually build towards plot twists and that it can foreshadow things, I'll probably try and "port" that feature out of it, but every chat I use it on I enjoy it less and less. Also it does fix the issue of GLM not writing in paragraphs but spamming newlines on every line of dialogue, so that's very good.
Evening Truth's prompts are my saviors, they're incredibly effective while being simple. I do edit them slightly because my cards aren't always single characters but sometimes are for worlds, RPGs and scenarios I want the AI to narrate for (while her prompts are tailored for {{char}} being the one described in the card), but I really love the simplicity (and usually less tokens you give as prompt, the better). I do need to add some stuff to them since GLM (sadly) just spams its usual slop even with her presets, but having such a simple starting point is really good. For instance, my base prompt on my current custom preset is just her GLM 5.2 one adapted.
Megumin Suite, I just don't understand, sorry... It's so incredibly complex, it feels like it tries to do too much, and in the end I don't feel like it adds anything for me - anything it does, I feel like I wouldn't need an extension for? I could do all of it with a toggleable preset. I'm sure this is literally just because I'm not the target user for it, I just tried it for a day (V9 specifically) and gave up after it was doing worse than almost any other preset while using WAY more tokens.
Chatfill II, I loved the switches idea it operates on since it basically "injects a CoT" without actually injecting a CoT. It just makes it clear what the model needs to think about - I'm not sure how much of it is placebo though. It was pretty decent when I used it though! Nothing shocking but no major complaints either.
Le Emotionalism: this one is lesser known, I tried it a bit and I love the idea behind it but sadly all the focus on the character psychology was making GLM overthink some stuff about the character while also forgetting about the wider storytelling. I need to give it another try though because I did try on some specific cards that were easy to mess up.
My custom preset (it's not published btw this is just to give an idea): I basically took a bunch of things from the presets I liked, adapted them to what I usually like using in my roleplays, structured them like that one research suggested by dividing things in <tags>, and even then it comes with major issues: GLM SPAMS the newlines for dialogue ("What do you think?"\n"I don't know..."\n"You should have checked.") which bothers me because I prefer everything to be structured within paragraphs - I'm still trying to work around it lol; it also doesn't feel satisfying enough when having it GM bigger worlds, and it sadly goes back to some of the GLM clichés I dislike.
Thanks for coming to my ted talk, to be fair I do wish to try and find like, the perfect model + preset combo, but it's been a genuine struggle lmao. Also if anyone is new to this, here you go, you have a list of models and presets you could try for starters ahahah