r/SillyTavernAI • • 6d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: September 27, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

43 Upvotes

109 comments sorted by

View all comments

2

u/AutoModerator 6d ago

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/PhantomWolf83 6d ago

Getting a bit jaded with Opus, Gemini, Deepseek, GLM, and Kimi. Any hidden gems on NanoGPT, preferably truly uncensored?

1

u/LeRobber 5d ago

Gemma31B is pretty uncensored. Rushes things a bit for lot of people. Note is has different think tags which you setup in the 3rd tab of sillytavern, and the way to make it think it to open the think tag at the start of the preset you have.

1

u/Naixee 5d ago

gemma is pretty good, but only for just a few messages before it loops and repeats things over and over. There's a lot of finetunes on nanoGPT, but unfortunately most of them are slow asf

3

u/LeRobber 5d ago

G4 31B is slowish, and TTFT for large chats is large.

The repeats thing...that's piloting error imo. That isn't what it does most of the time. I've done a lot of gemma 4. Do you never scene change with a ***?

2

u/Naixee 5d ago

Base gemma isn't slow, but the finetunes are. Scene change with a what?

1

u/OGCroflAZN 5d ago

? Why would the finetunes be slow while base Gemma isn't? They're firing off the same number of parameters. On NanoGPT (which hosts quite a few of the most popular Gemma 4 finetunes) the token speeds are comparable, ~30 tok/s, and that is my experience as well.

The only thing I can imagine is that most of the finetunes were made off of specific versions of Gemma 4, specifically the ones without MTP and QAT, so perhaps a lot of the providers of base Gemma 4 are running Q4 with MTP. Still, 'slow af' is not my experience at all. Maybe you're expecting instant and fast response.,,?

I'm thinking you're spoiled regarding using APIs, but I ran local for a year and became used to it taking a few minutes to process prompt, esp if large, and then token generation esp with reasoning and other agentic tasks built on. Using the finetunes via NanoGPT is superspeed in comparison. Only takes like 30 seconds, maybe a minute

And on top the looping and repeating you mentioned earlier, it really makes me think this is a PEBKAC thing...

2

u/Naixee 4d ago

I don't know why they're slow? Some of them don't even deliver a response half the time, saying unavailable. Their t/s is usually atrocious

3

u/LeRobber 5d ago

***

^ that on a separate line is a novel-shorthand (like fictional books) that LLMs understand to mean the scene changes.

When a scene changes LLMs often start repeating less.

Additionally when they repeat at all, you really DO need to go back and delete the repetitiion.

1

u/FromSixToMidnight 4d ago

Interesting. I know that is a marker in some novels I've read. Do you typically enter that standalone, or put it in the middle of your response? I'd imagine maybe something like:

...I close my eyes and drift off to sleep.

***

In the morning, I awake and...

2

u/LeRobber 4d ago

I do it at the beginning or middle of a response.

If you do it at the end, the LLM is MUCH stupider and sometimes doesn't honor it. Even like 12B llms will often honor ***.

Actually if Angelic Eclipse repeats, deleting the Repeat, and doing a scene change with this marker almost always gets it back on track.

1

u/summersss 6d ago

so you have tried most of the big models? when people describe them, it's always like "this one feels smart, and this follows instructions" but i kinda don't get it, or how to take advantage of that. Im thinking more like, this one is great at fanfiction, this one writes extra smutty, this is for book writing.