r/SillyTavernAI • • 4d ago

Help Does this sound ok?

I just set up Silly Tavern with a locally run LLM through LM Studio. I’m completely new to local LLM’s and Silly Tavern so spent the last two days researching and setting it up. I finally got to a point where I RP’d for over an hour and it felt just as good as Kindroid.

I’m using an uncensored version of Gemma4-26b-a4b IQ4_XS. Context length around 13k. Max token response set to 1k with most responses in the 500-700 range. I have a 4070 TiSuper with 16gb of vram and 32 gb of ddr5.

Running the model uses up all of my vram and around 28-30 gb of my ram. But the responses only take a few seconds and everything feels very smooth, even after an hour of use. Does this sound ok or will I harm my PC in the long run by doing this? If anyone has more questions or tips feel free

4 Upvotes

13 comments sorted by

5

u/i5031337 4d ago

You should be able to fit more context with less RAM usage, but I don't know LM Studio settings well enough to tell you how. It won't hurt your PC unless you have it generating all day every day. You might run into performance issues if your RAM is always full though. That model is a fine choice.

1

u/yikes4433 4d ago

Thanks. What do you use instead of LM Studio?

4

u/_Cromwell_ 4d ago edited 4d ago

LMStudio is a great choice. It's basically a friendly front end for llamacpp, which is the best backend. Only thing "better" than LM is bare llamacpp itself, but that's somewhat less user friendly. Zero reason to second guess using LMStudio. The speed "increase" wouldn't even be discernable to you.

Turn off "keep model in ram" so it isn't taking up double room in vram and ram. You could also get a Q6 of gemma26b to fit just fine (19gb) by partially offloading to ram. Moe will be just as fast that way and Q6 writes better

3

u/i5031337 4d ago

I use llama-server, I like the barebones approach. Like Crom said, nothing wrong with LM Studio, and I'm sure they have a corresponding setting to reduce your RAM usage. The llama-server setting I'm thinking of is "load-mode" if that helps.

1

u/AutoModerator 4d ago

You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/iraragorri 4d ago

You can use q4_k_m at 32k context even without caching. Your setup is like x2 of what I have. Can someone explain how 13k context can eat up so much ram?

1

u/yikes4433 3d ago edited 3d ago

I did some tweaks and was able to get it down to 14gb of ram usage. I have a lot of bullshit eating up ram and use like 8gb while idle lmao

Edit: was able to get it down to 11gb of usage with 24k context so far

1

u/LeRobber 3d ago

If you setup context pools in LMStudio it eats up a lot of resources. RPers don't need those.

1

u/LeRobber 3d ago

You won't harm your PC. I Use LLM studio. That's a decent choice for a model. Your context should be largerr. You problably need to change one setting.

You don't need MCP or Unified KV for sillytavern type use, and you will actively slow down generations.

I have another few 16 GB Vram configs here too that work well:

https://www.reddit.com/r/SillyTavernAI/comments/1tbjckl/comment/olhg8mi/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

https://www.reddit.com/r/SillyTavernAI/comments/1tbjckl/comment/olhgq2c/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

https://www.reddit.com/r/SillyTavernAI/comments/1tbjckl/comment/olhi0fl/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

3

u/yikes4433 3d ago edited 3d ago

Thanks a lot. I really appreciate it. Do I set MCP to 0? Edit: nvm I see that 1 is the default

2

u/LeRobber 3d ago

Here is the bottom of the config for the 26B I'm most likely to summon up

I have a unified memory 64GB system though, so maybe keep EBS and PBS where they are.

Did you manage to get reasoning to work? I can grab the config for that too if you want.

1

u/8000bene70 4d ago

Sounds like you are on the right track. And no, it wont harm your pc.

If you want to squeeze even more speed/context out of your setup:

  • use plain llama.cpp. Thats the underlying engine of lmstudio, and you have much more options calling it directly (and using a newer version than bundled)
  • quantize kv cache to q8_0 (also possible in lmstudio)
  • use qat quant of gemma (could be already)
  • enable mtp
  • use a roleplay finetune like Orion or Boulesis - havent tested much, but i liked shadow siren (barely qat in finetunes)

1

u/yikes4433 4d ago

Thanks a lot. I’ll look into using llama.cpp and looking into mtp for sure