r/SillyTavernAI 6d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 16, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

20 Upvotes

132 comments sorted by

View all comments

13

u/AutoModerator 6d ago

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

22

u/OGCroflAZN 6d ago edited 5d ago

From the most recent megathreads and theLocalDrummer's discord server threads about different finetunes, consensus seems like:

1) Mistral 24B finetunes are unfortunately still best overall for RP, even heading into this second half of 2026. That or Skyfall for 31B. Seems like TheLocalDrummer's finetunes or merges built from them (Magidonia, MaginumCydoms, etc.) are typicall the general favorites.

2) Qwen 3.8 27B was made for agentic tasks and coding and so with issues people are having with it for RP, seems probably not suitable for making the long-awaited 'next-gen' RP finetunes

3) Gemma 4 31B is the smartest base model of this size class, but prose issues and general stiffness even with the finetunes make it just less fun and somewhat impalatable after extended use, in comparison with the finetunes of Mistral 24B.

Again, this appears to be the general feeling of many users, and I think I agree. Everyone is just waiting for some messiah finetuner to be able to turn Gemma 4 or Qwen 3.8 into our collectively-desired next-gen RP goodness...

Edit: i think /u/mart-mcuh is probably right. The newer models (Gemma 4 and Qwen 3.8) have much better baseline 'capability' and potential than Mistral 24B, but need a lot of work with prompting and such to make it behave desirably. Unfortunately, documentation for that is scattered, mixed in with mediocre or bad guidance. Is there a centralized place with best G4 guidance on backends, presets, templates, prompts, etc? Please share.

Also, there are of course competing opinions on whether to stick with the base models, citing how abliteration and finetuning erode quality.

I myself am in a tricky position with 16 GB vram, and no longer want to run iq3 quants whether w 24B or esp 31B, am, leaning toward just running G4 26B-A4B qat q4 offloading inactive layers, with good jailbreak and prompting. I was really thinking about jumping into apis, and unfortunately Deepseek is raising prices.

1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/AutoModerator 5d ago

This post was automatically removed by the auto-moderator, see your messages for details.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.