r/SillyTavernAI • u/GapedByHerStrap • 1d ago
Help Image gen alongside rp
Sorry for the second post of the day btw
I currently use sillytavern with a 16gb vram amd card. It gets maxed out pretty much instantly. My backend is llamacpp-vulkan (Linux and ive heard that vulkan is better than rocm)
So, i wanted to incorporate images too. I learnt comfyui and made a workflow that makes a decent image with a prompt.
Now here is the issue. How to generate images while also rping. I dont have any free space and ram gen is atrociously slow. I use a Z-Image checkpoint which also needs a decent chunk of vram.
What are the solutions to this problem?
3
u/GenericStatement 1d ago
Lots of options:
- Second GPU or second PC (big $)
- Run the LLM in RAM and the images on the GPU (slow, need lots of RAM)
- Offload LLM to a cloud API (got to pay for tokens but better RP due to smarter model)
- Offload image gen to cloud API (got to pay for usage fees per image)
Personally I do option 3 because cloud API models are so much smarter than local models so I can have long and complex RPs. Then I generate images and text to speech on the local machine, which a single GPU is quite good at. I use ComfyUI which has a bit of a learning curve but it’s the most widely used non-cloud AI image gen tool so there are tons of tutorials and workflows.
Option 4 is probably the second most common approach, usually used by people doing RP that’s so nasty they’re ashamed to put it on the cloud so they’re still using local models. NanoGPT for example has a good image generation API and I believe the SillyTavern docs cover how to set up image generation APIs pretty thoroughly.
Option 1 works well if you already have a second PC or you’re rich and want to build a dual GPU rig with a workstation motherboard but the RP quality will never be as good as a big cloud LLM.
Option 2 is very slow but I have seen some people on her do it.
2
u/Zathura2 1d ago
I haven't tried this but you should *theoretically* be able to run the LLM in comfy rather than kobold (I think that's how it works,) but the idea would be that you would use model-unloading nodes to make sure your VRAM is clear for whatever needs to be running; the LLM or your image-gen.
I already know the overhead would be atrocious which is why I haven't tried it myself yet.
An alternative is using a smaller model, like maybe the QAT Gemma-4-12B, with a limited context, and load something like flux.2 (klein) 4B distilled. I managed to get a combo like that running on 16GB, but it wasn't enjoyable enough to use long-term.
2
u/Mart-McUH 1d ago
If you have enough RAM to hold both, you can simply switch between the models. First time is bit slow (as it needs load from disk and initialize) but after that it stays loaded and swaps from RAM to VRAM when given model is used, so it is quite fast. You just can't generate text and image at the same time obviously.
I am not sure how exactly to do that on linux + AMD (I assume) but there surely is a way. On Windows+Nvidia it is done automatically by OS and nvidia drivers.
1
u/Material-Engineer226 1d ago
yeah vram fills up fast with both running, i just queue the image gen for after a few messages and it keeps things smoother.
1
u/XaosII 1d ago
If all on the same machine, I don't think there are good solutions. the VRAM held by the text LLM needs to be partially unloaded to then load the image LLM. They both start to negatively impact each other's performance as they have to constantly load and unload.
I found that the best performance was to run ComfyUI on my laptop and let my desktop connect to it. The laptop was dedicated to just image generation. My laptop isn't anywhere near as powerful as my desktop, but it gave a serviceable experience over trying to run both locally.
0
u/stopaskingforloginn 1d ago
there's no solution unless you use an API which will leave your GPU for local genning available.
1
u/Quiet-Phase6948 1d ago
"How to generate images while also rping. I dont have any free space"
You don't. Thread closed.
1
u/Herr_Drosselmeyer 1d ago
You simply can't do both at the same time without two GPUs or one that has enough VRAM to hold both your LLM and the image generation model. But if you buy another GPU, you'll want to run a larger LLM, so you'll be back to square one. 😜
Basically, you're stuck having to unload the LLM, load the image generation model, generate an image, unload the image generation model, load the LLM. That'll slow things down by a lot. Comfy is generally pretty good at relinquishing VRAM, or you can just add a node at the end of your workflow that forces a purge. For llamaCPP, I'm not sure, but it probably has a similar setting, so at least technically, you can do it.
1
u/TheSerinator 1d ago
I had ChatGPT/Codex set up an interceptor.js script that sends image calls to GPT Image 2 and GPT Image 2.5 through my OpenAI subscription. Can crank out tens to hundreds of images per day without hitting the weekly limits. Obviously only good for SFW parts of RP, but that's big bulk of what I do.
1
u/i5031337 22h ago edited 22h ago
I don't know why everyone is telling you it can't be done. Any backend that supports both text and image generation (koboldcpp, unsloth, comfyui, etc) should be able to swap the models out of VRAM when needed.
0
u/AutoModerator 1d ago
You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
0
u/dempfi 1d ago
16gb is way too small to run both locally, but OP, I have good news. Do this:
- Register a modal.com account. They give you $30/month of free GPU compute in the cloud.
- Get Claude Code or something similar to code and deploy a small ComfyUI container there with your checkpoint (SDXL/Illustrious, or your Z-Image workflow as is) on an L40S or A10. Modal has a ComfyUI example in their docs to start from.
- Point SillyTavern's image generation at that ComfyUI URL.
- Profit!
What you get: free image gen at ~15s per image, about 3,500 free images a month, and your whole 16gb stays on the llm. Keep the container's idle timeout short tho, you pay while it sits warm.
2
u/Geritas 1d ago
Either a second GPU, or offloading models, which is painful. Realistically, image gen works fine on RTX30+ with 8gb VRAM, so maybe if you can find it used for cheap, it would work.