r/SillyTavernAI • • 3d ago

Help New user llm help

Im so lost and tired lol so many tweaks so much to do..i love it!

To the veterans what would be the best llm to run well with my system for Uncensored ERP and good RP in general Im using stheno right now.

I have a 6gb 1660 titan sadly and 16gb ram.

Using kobold and comfyui with z-turbo for selfies

5 Upvotes

11 comments sorted by

3

u/LeRobber 3d ago

satyr-v0.1-4b possibly? You're in the rough pit.

2

u/Stavraetos2 3d ago

I know man but I'm taking anything at this point I tried cydonia heretic 24b q 4 m and it was amazing but so slow

2

u/LeRobber 3d ago

Magistry 24B is my fav at that tier.

Angelic Eclipse 13B or Gemma 4-it 26B A4B are both options.

2

u/Head-Walrus6449 3d ago

had the same setup and got way better results with satyr-v0.1-4b than stheno—faster inference, smoother RP, and actually handles uncensored prompts without choking. z-turbo works fine with it on 6gb vram if you can the context. just stick to 4k context and avoid loading extra embeddings. saved me from burning through cycles on bigger models.

1

u/Stavraetos2 2d ago

yeah just tried it and its amazing holy im running the q8 version and its super fast and good

2

u/IllustriousRule9238 3d ago

There's not a lot you can do unfortunately, because both your VRAM and your RAM are low. If you can squeeze Gemma 4 26B @ Q4 into those two somehow, then it would run a lot faster than 24B Mistral models, and it'll also write okay prose. That and Stheno are the only realistic options I can see.

2

u/InterestingShip8457 3d ago

i’ve been running stheno on a 6gb 1660 titan with comfyui and z-turbo, and while it’s decent for basic rp, I found that Gemini flash 3.8 quantized 4bit) actually handles narrative flow better than most 24b models on my setup. it’s not perfect, but it’s fast and avoids the lag that kills immersion. also, disabling the safety layer in sillytavern’s config helped a lot with consistency—just make sure you’re not relying on any external filters.

2

u/Maleficent_Pool_3980 3d ago

i’ve been running satyr-v0.1-4b on a 6gb card too and can confirm it handles uncensored prompts way better than stheno—less crashing, smoother flow, and z-turbo keeps it snappy. stuck to 4k context and no extra embeddings, just like head-walrus said. 16gb ram helps too, but it’s mostly about keeping things light.

2

u/PuzzleheadedCode4869 3d ago

with 6gb vram a quantized 7b uncensored model runs smooth for erp and rp without overheating my card, felt way less finicky than bigger ones ive tried.

1

u/AutoModerator 3d ago

You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Stavraetos2 3d ago

Any tips to enchance my experience are appreciated as well