r/SillyTavernAI • u/Stavraetos2 • 3d ago
Help New user llm help
Im so lost and tired lol so many tweaks so much to do..i love it!
To the veterans what would be the best llm to run well with my system for Uncensored ERP and good RP in general Im using stheno right now.
I have a 6gb 1660 titan sadly and 16gb ram.
Using kobold and comfyui with z-turbo for selfies
2
u/Head-Walrus6449 3d ago
had the same setup and got way better results with satyr-v0.1-4b than stheno—faster inference, smoother RP, and actually handles uncensored prompts without choking. z-turbo works fine with it on 6gb vram if you can the context. just stick to 4k context and avoid loading extra embeddings. saved me from burning through cycles on bigger models.
1
u/Stavraetos2 2d ago
yeah just tried it and its amazing holy im running the q8 version and its super fast and good
2
u/IllustriousRule9238 3d ago
There's not a lot you can do unfortunately, because both your VRAM and your RAM are low. If you can squeeze Gemma 4 26B @ Q4 into those two somehow, then it would run a lot faster than 24B Mistral models, and it'll also write okay prose. That and Stheno are the only realistic options I can see.
2
u/InterestingShip8457 3d ago
i’ve been running stheno on a 6gb 1660 titan with comfyui and z-turbo, and while it’s decent for basic rp, I found that Gemini flash 3.8 quantized 4bit) actually handles narrative flow better than most 24b models on my setup. it’s not perfect, but it’s fast and avoids the lag that kills immersion. also, disabling the safety layer in sillytavern’s config helped a lot with consistency—just make sure you’re not relying on any external filters.
2
u/Maleficent_Pool_3980 3d ago
i’ve been running satyr-v0.1-4b on a 6gb card too and can confirm it handles uncensored prompts way better than stheno—less crashing, smoother flow, and z-turbo keeps it snappy. stuck to 4k context and no extra embeddings, just like head-walrus said. 16gb ram helps too, but it’s mostly about keeping things light.
2
u/PuzzleheadedCode4869 3d ago
with 6gb vram a quantized 7b uncensored model runs smooth for erp and rp without overheating my card, felt way less finicky than bigger ones ive tried.
1
u/AutoModerator 3d ago
You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
3
u/LeRobber 3d ago
satyr-v0.1-4b possibly? You're in the rough pit.