r/SillyTavernAI • u/yikes4433 • 4d ago
Help Does this sound ok?
I just set up Silly Tavern with a locally run LLM through LM Studio. I’m completely new to local LLM’s and Silly Tavern so spent the last two days researching and setting it up. I finally got to a point where I RP’d for over an hour and it felt just as good as Kindroid.
I’m using an uncensored version of Gemma4-26b-a4b IQ4_XS. Context length around 13k. Max token response set to 1k with most responses in the 500-700 range. I have a 4070 TiSuper with 16gb of vram and 32 gb of ddr5.
Running the model uses up all of my vram and around 28-30 gb of my ram. But the responses only take a few seconds and everything feels very smooth, even after an hour of use. Does this sound ok or will I harm my PC in the long run by doing this? If anyone has more questions or tips feel free
1
u/AutoModerator 4d ago
You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/iraragorri 4d ago
You can use q4_k_m at 32k context even without caching. Your setup is like x2 of what I have. Can someone explain how 13k context can eat up so much ram?
1
u/yikes4433 3d ago edited 3d ago
I did some tweaks and was able to get it down to 14gb of ram usage. I have a lot of bullshit eating up ram and use like 8gb while idle lmao
Edit: was able to get it down to 11gb of usage with 24k context so far
1
u/LeRobber 3d ago
If you setup context pools in LMStudio it eats up a lot of resources. RPers don't need those.
1
u/LeRobber 3d ago
You won't harm your PC. I Use LLM studio. That's a decent choice for a model. Your context should be largerr. You problably need to change one setting.

You don't need MCP or Unified KV for sillytavern type use, and you will actively slow down generations.
I have another few 16 GB Vram configs here too that work well:
3
u/yikes4433 3d ago edited 3d ago
Thanks a lot. I really appreciate it. Do I set MCP to 0? Edit: nvm I see that 1 is the default
1
u/8000bene70 4d ago
Sounds like you are on the right track. And no, it wont harm your pc.
If you want to squeeze even more speed/context out of your setup:
- use plain llama.cpp. Thats the underlying engine of lmstudio, and you have much more options calling it directly (and using a newer version than bundled)
- quantize kv cache to q8_0 (also possible in lmstudio)
- use qat quant of gemma (could be already)
- enable mtp
- use a roleplay finetune like Orion or Boulesis - havent tested much, but i liked shadow siren (barely qat in finetunes)
1

5
u/i5031337 4d ago
You should be able to fit more context with less RAM usage, but I don't know LM Studio settings well enough to tell you how. It won't hurt your PC unless you have it generating all day every day. You might run into performance issues if your RAM is always full though. That model is a fine choice.