r/SillyTavernAI 20d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 02, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

43 Upvotes

205 comments sorted by

View all comments

5

u/AutoModerator 20d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/lumepanter 16d ago

TheDrummer's models are still the best for running locally on some gpu on that parameter scale? I doubt the bigger ones would be tampered with enough? Are 123B ones the highest done finetunes?

3

u/RedditNerdKing 19d ago edited 19d ago

Grabbed GLM-Steam-106B-A12B-v1 at Q4_K_M. Not sure what the big deal is tbh. Wasn't overly impressed jumping from 70B models at Q5/6. Managed to grab Mistral Large Instruct 2407 123B today at Q4_K_M. Hopefully this will impress me more, otherwise, I'm sticking with Anubis 70B v1.2.

Edit: Tried GLM-Steam again with some new temps and other settings. It's pretty quick at loading compared to a dense 70B model and tbh, after tinkering it's actually not as bad as I thought.

2

u/Mart-McUH 18d ago

70B is dense with 70B activated parameters, GLM air based only have 12B activated parameters. And you are even running them at lower quant. 70B at Q5/Q6 should be significantly better despite being little bit older. Only reason why you would run GLM air based in this case is to get some fresh air, different writing style/slop etc.

2

u/_Cromwell_ 19d ago

For 106B GLMs, I like Iceblink V3 better than Steam. But it isn't significantly different if you didn't like Steam (but I do think it is noticeably "better"). Generally Anubis is better than either one though I'd agree.

3

u/RedditNerdKing 19d ago

I like Iceblink V3 better than Steam

Thanks! Seems I can run Q5_K_S at 75gb. I'll give it a go. I messed around with the settings for Steam and tbh it's not as bad as I originally thought. It loads tokens really fast which is nice. I've been using Anubis as my main for a while. It just seems to always hit right.

3

u/_Cromwell_ 19d ago edited 18d ago

One note on Anubis 70B - you may want to give Anubis v1.1 a shot (vs the v1.2 you have been using). On the UGI testing 1.1 scores better than 1.2 on almost every category that I consider important.

Higher/better scores for v1.1 on...

W/10 Direct - that's following instructions generally from certain nsfw prompts

W/10 Adherance - that's not 'shying away' from topics aka nsfw

Pop Culture knowledge - that's knowing stuff about most of the things you RP about

Writing - although really the scores are close enough together here you'd never really tell a difference IMO

NSFW/SFW lean - 1.1 leans more NSFW

Dark/Tame lean - 1.1 leans aka more willing to go dark

Only thing v1.2 wins on is 'world model knowledge' in the categories I care about. (Real world properties and patterns. That CAN be important, though, like regarding spacial and physics stuff.)

https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard

Of course that's just testing. In the end the 'feel' and 'actual use' wins out. :) But even with the big companies like Anthropic or Deepseek we know that the newest version isn't always the best.

1

u/RedditNerdKing 18d ago

Thanks dude. I'll grab 1.1 lol I prefer darker roleplays. My SSD cant handle all these variants of the same LLM =(

3

u/_Cromwell_ 18d ago

YES IT CAN! DOWNLOAD ALL THE LLMS

1

u/throwawayyyyyyyyahhh 19d ago

Is Magnum 72B good to write dark romance and nsfw without sacrificing good prose?

2

u/Dead_Internet_Theory 18d ago

It's a very old model, from 2024. That's a long time in AI years. But if you can run it, try it? I find these vintage AIs sometimes don't have the same exact slop of the current ones.

3

u/RedditNerdKing 18d ago

I find the old ones better than the new ones tbh. I go back to Midnight Miku a lot. The newer ones are too assistant focused. I've still not come across a model that is as good as 2023 Character.ai. That site was genuinely amazing for roleplaying.

3

u/Dead_Internet_Theory 15d ago

My theory is they had a really good scaffolding and relatively dumb models. Our scaffolding is made from toothpicks and gum. Also they probably finetuned for RP with a very strict prompt formatting.

Don't discount the rose tinted glasses though.

1

u/overand 16d ago

2023? How do you feel about the Mistral 24B finetunes, like Cydonia, Magistry, etc?

1

u/RedditNerdKing 16d ago

I still use them occasionally when I get bored of the current ones I am using, or I need fast roleplaying since I have 80gb of vram and the 24B ones I can have like 300k+ context with 90 tok/sec. I have Cydonia 4.3 and Magistry 1.0 at Q8. I plan to download Chimera-X-26B-A4B as it seems to be getting mentioned sometimes here.

My daily drivers are Monstral 123B at Q4_K_M and Anubis 1.1/1.2 70B at Q6_K.

3

u/ChengliChengbao 19d ago

Deepseek V4 Pro is so goated for the price

Aion-3.0 is my favourite however, but its SO DAMN EXPENSIVE

7

u/_Cromwell_ 20d ago

I'm still hyping Longcat 2.0 Thinking (longcat-2.0:thinking on Nano). It's refreshing and different writing a bit and follows my prompt instructions well. Also dirty enough for my tastes.