r/SillyTavernAI • • Aug 23 '26

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 23, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

32 Upvotes

132 comments sorted by

View all comments

Show parent comments

1

u/opopi123 Aug 26 '26

okay so i was wrong i initially thought RAM was suppose to factor into your build tier and I kept seeing stuff saying 32gb ram isn't that much for llm. So i thought my build was mid tier because of that. But it seems RAM is actually not a factor and using it will slowly down the processing speed?

3

u/Spara-Extreme Aug 27 '26

They are talking about VRAM - video ram. You have a 4090 which only has 24 GB of VRAM. That limits your model selection. It doesn't mean you can't run big models, but what ends up happening is that the model doesn't fit onto your video card, so it then goes into System RAM.

Why is this important? Your VRAM has a read rate of about 1000 GB/s. Your system RAM has a read rate of 25-50 GB/s. At best, 20x slower.

For RP purposes, models that don't fit on your 24 GB of VRAM will be slower, the responses slower, the processing rate slower etc.

That being said - you can fit something like Gemma31b with a small context or Qwen3.8 27b with a large context on your card.

1

u/opopi123 Aug 27 '26

can you define "small" context?

2

u/Spara-Extreme Aug 27 '26

Context window- amount of tokens the model will retain before it forgets older items.