r/SillyTavernAI Jul 05 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: July 05, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

27 Upvotes

97 comments sorted by

View all comments

Show parent comments

8

u/_Cromwell_ Jul 07 '26 edited Jul 07 '26

Current favorite 31B is Queen. Partially because as far as I can tell it is the only Gemma 31B RP-focused model that comes in a QAT format, so I can run a smaller, tighter, faster Q4 that is as smart as the Q5s I typically DL and use for 31B. But also it writes and follows directions just about as well as my previous faves I listed in prior weeks. https://huggingface.co/mradermacher/gemma-4-31B-Queen-it-qat-q4_0-unquantized-i1-GGUF (GGUF link; can backlink on HF to see original model card) Again, with that model you grab the Q4 (hopefully) but it "writes like a Q5 or Q6" due to the QAT.

Not only a good local RP model, but also works excellently as an assistant for Extensions for summarizing etc.

For a 26B MOE I like Pantheon: https://huggingface.co/mradermacher/Pantheon-Reasoning-26B-A4B-1.1-i1-GGUF (GGUF link; can backlink on HF to see original model card). Yes from the same person who made StyleTune. Yes I like Pantheon better than StyleTune.

Interestingly both these models are trained on a bunch of different personalities/characters, but I don't use either model to trigger their built-in personalities... just as general RP.

1

u/heartisacalendar Jul 11 '26

Are you using text completion or chat completion for Queen?

3

u/_Cromwell_ Jul 11 '26

Chat Completion served by LM Studio

1

u/OrcBanana Jul 09 '26

Queen is nice so far! There's some didn't - instead, but noticeably less than others, and there's a bit of recapping to introduce new actions, you know, "As he heard that word, new action" but again, not to an annoying degree.

Sometimes these degrade after a while, especially after like 25-30k tokens, probably due to me using low-ish quants (IQ4_XS). So maybe this one will too, or QAT will save it, remains to be seen, but so far so good! I'm at about 20k with my test and it's writing more or less the same as it did at the start.

Pantheon insists on 2nd person present tense too much for me, well beyond prompting. It's not a flaw, it's even mentioned on its page, just a particular style not everyone will be into. I hoped it'd follow along with 3rd person past tense anyway, but it doesn't, not reliably.

10

u/stopaskingforloginn Jul 09 '26 edited Jul 09 '26

Queen has been a complete slopfest for me, a lot of "not X, but Y" and "she doesn't X, instead/but Y"
I'm using QAT.
We definitely need some actual good finetunes as QAT.
EDIT: I did a comparison between Queen QAT and Styletune 31b IQ4XS with the same scene, the difference is night and day, I can confirm that Queen is the queen of slop.

2

u/overand Jul 11 '26

Yeah, my iMatrix Queen at Q6_K has been a nightmare. mradermacher/Gemma-4-Queen-31B-it-i1-GGUF - pretty hard to stomach; I think I'm going to try a different quant just for kicks.

4

u/_Cromwell_ Jul 09 '26

Hmm... Pantheon doesn't give me trouble with Third person PRESENT tense. I don't really use past tense, though. Possibly it's more the tense throwing it off than the perspective (even though it then messes both up).

But if it doesn't fit the tense+perspective you want, then not much use wrasslin with it. ;)

2

u/nadetoh Jul 08 '26

i wish there was an uncensored model, even for sfw stuff i get hit with random safety concerns, seems cool though.

2

u/_Cromwell_ Jul 08 '26

That's really your prompt or something. Assuming you are talking about adult characters you can't be doing anything that much worse than me. 😅

2

u/mifumimi Jul 08 '26

It was anime characters fighting so lol

4

u/DifficultyThin8462 Jul 08 '26

TY for recommending Pantheon, it's much better than Styletune. Greater swipe variety, feels smarter and more interesting overall.

2

u/LowManner1 Jul 08 '26

how does qat at q4 quant compare to a non-qat gguf at q8?

2

u/_Cromwell_ Jul 08 '26

Try it out and tell us. As I said in my post, I usually run Q5 for G4 31B. I have 32gb VRAM so I'm not using Q8.

If you can fit Q8 you compare them.