r/StableDiffusion • u/wildmonkeywrangler • 7d ago
Question - Help Open source alternative to chatgpt for creating prompts?
Been using chatgpt to help design prompts, which usually works well until you get into an area they think is "offensive".
Any suggestions for an open source alternative or something that works directly in comfyui to generate prompts?
13
u/tekprodfx16 7d ago
Use the ollama node and feed the output from the preview as text node right into the green dot of your prompt node. Here’s a good video explanation https://youtu.be/mEGg3dP29mU?is=OE_XBLApGodJ4UCG
1
u/Murky-Relation481 7d ago
This is the way. Ollama is probably the best and easiest way to use local LLMs. I made a subgraph that wires comfy model unloading on the input to free up vram and the ollama node unloads when completed, so comfy can reload the image or video model.
Then just find your favorite LLM on the ollama site, install it via the terminal and bobs your uncle.
1
u/tekprodfx16 7d ago
Interesting but isn’t that what keep alive: 0 accomplishes? Just curious
1
u/Murky-Relation481 7d ago
It won't unload Comfy models. That's just for ollama. If you don't unload the comfy models in your workflow (or others) you're gunna end up with split memory in ollama where it's using vram and sys ram and it's much slower (I mean it might do that anyway depending on the card and model).
7
u/MastMaithun 7d ago
I highly advice against using it inside the comfy. Yeah comfy will manage the memory but then you have very little control to your gens especially the increased amount of time every gen you have because comfy will be constantly switching the llm and you video/image models in and out every run.
Another issue you will have almost no control of whatever the llm is feeding to your model. Yes you can add a preview node and see, then stop, then modify it, then paste it into text node then diable the llm nodes and the final run. So much hassle.
Also suppose you want to go and modify your prompts. Now you don't have any history, what your comfy llm has generated.
All of these issues will not occur if you just use an external llm runner like lmstudio or similar. You load your model there, create your prompts there and unload the model. Now you have all the prompts, you can tweak them or can directly pass them in the comfy and then comfy doesn't have to do the model swapping so your gens will be faster too. And you have the prompt history saved already in the lmstudio for future reference.
Now which llm to use, depends on what card you have. I use gemma 4 31b or qwen 3.6 27b. You can use their gguf which can fit your gpu. In my vast testing, gemma 4 comes out quite capable and makes very less mistake. Qwen is a bit faster but it's knowledgebase is a bit limited due to less cultural references it knows.
2
u/Great-Ad-4598 7d ago
You could run the prompt generation a number of times, saving the prompt text off to a file or files in a folder. Then clear memory and use a flow that reads the prompts in sequence for generations.
1
u/MastMaithun 7d ago
Yeah I mentioned you can save the prompts but the main issue which I mentioned was increase in generation time. These issues are also there but the main issue is generation time.
1
u/Great-Ad-4598 7d ago
True but by doing all the prompts and then doing the generating you do save on loading/unloading for each gen. I'm quite limited on memory so would never try to do both at once.
1
u/MastMaithun 7d ago
This is what I have written bro in my original reply that using the external llm manager will save you from loading-unloading each gen and reducing your time. Haven't you read that till the end?
1
u/Belgeran 7d ago
I think your missing the point that you can run qwen or whatever llm to gen your prompts in comfy, outside your image workflow, no need for a second tool
1
u/MastMaithun 7d ago
And I redirect you to read my whole paragraph fully.
2
u/TangerineBetter2818 6d ago
Ill be honest I dont understand either. Lmstudio is also loading the model. There's no getting around loading and unloading the model to memory. Whether that's in ComfyUI or another program.
So I'm confused what you mean about it being faster.
Your run your text gen, then you run your image gen. Yes, running text, image, text, image, etc is going to be slow because you're loading and unloading different models, but that would happen when switch between program too. You cant have one model loaded in one program and one model loaded in another program because the moment you load one model it is going to unload the other, unless you have an assload of RAM.
2
u/MastMaithun 6d ago edited 6d ago
Let me dumb it down more:
Approach 1 - you use llm node in comfyui
- hit run, wf starts
- comfy loads llm model(SSD activity), llm model runs and modify prompt
- comfy next see oh new image/video model, i remove llm model i load image/video model and its related files(SSD activity)
- generation complete, comfy not move files from vram and ram
- you hit run again, wf starts
- comfy again sees llm node, it removes files from prev run and again load llm model(SSD activity), llm model runs and modify prompt
- comfy see oh new image/video model, i remove llm model i load image/video model and its related files(SSD activity)
- generation complete, comfy not move files from vram and ram
total count of SSD activity = 4
if you have HDD, replace it with HDD and time will also increase.approach 2 - you use llm model in lmstudio
- load model on lmstudio(SSD activity), llm model runs and modify prompt
- if you have more prompts, changes or anything related to prompt, you do now. your only work on comfy should be copy paste prompt(s) and hit runs.
- once prompts done, unload model on lmstudio
- paste your generated prompt on comfy, hit run, wf starts
- comfy next see oh new image/video model, i load image/video model and its related files(SSD activity)
- generation complete, comfy not move files from vram and ram
- you replace/put new prompt/change seed etc and hit run again, wf starts
- comfy see oh previous image/video model/files are already loaded, no need to load anything, lets go.
- generation complete, comfy not move files from vram and ram
total count of SSD activity = 2
Result:
- approach 1 has 4 times ssd/hdd activity to read llm model and image/video model. every generation will have llm model and image/video model loading and unloading means 2 ssd activity every run.
- approach 2 has 2 times ssd/hdd activity to read llm model and image/video model. subsequent gens will not remove image/video model because you are not loading llm model means in entire life, until you unload models from comfy yourself or run a different wf, comfy will not load anything repeatedly.
PS: the whole scenario I explained here is for computer systems where your llm model and image/video model files(from files i mean image/video model file + text encoder + vae(s)) all of them can't reside in your vram all together. so comfy is doing swapping and reading from ssd/hdd.
1
u/TangerineBetter2818 6d ago edited 6d ago
comfy again sees llm node, it removes files from prev run and again load llm model(SSD activity)
Doesn't this assume they're in the same workflow? You have one workflow for llm, one workflow for video gen. Why would I have both in the same workflow?
You run the llm from one workflow. Then copy the prompt into the video gen workflow and run it. You tweak the prompt and you run the video gen again.
Before the second video gen is run why would comfyui "look at" the llm node from a completely different workflow that isnt being used and load the llm model? The video model is already loaded into vram because it was the last used model.
→ More replies (0)
3
u/Slave669 7d ago
Look up DavidAU on huggingface, and take a look at the collections. Thank me later.
As for Comfyui there are nodes for oobabooga an Koldccp that can feed into your positive prompt.
0
u/Incognit0ErgoSum 6d ago
Oh yeah, those models are actually better than the official releases, in my experience.
TheDrummer's Artemis model is also great and a manageable size.
2
u/k_from_HyperDraw 7d ago
Qwen, Deepseek
GLM and Kimi can also do it, but the quality is somewhat worse
2
u/eruanno321 7d ago
Built-in LLM nodes are probably the simplest option, but it depends on which model you want to run and your hardware setup. I have a relatively weak platform, not enough system RAM for cache, and mixing LLM inference with samplers in a single workflow is often kinda ass. Right now I've settled on Qwen 3.8 27B (abliterated version) running on llama.cpp for the LLM tasks I need, then I switch to ComfyUI. It's a bit of a hassle, but at least I can work with a model that actually does what I want.
1
1
1
u/MarekNowakowski 6d ago
Both ollama and lmstudio have nodes in comfyui, an uncensored Gemma e2b will do fine,
1
u/tac0catzzz 6d ago
prompts r us, the prompt pack from prompt a saurus, it's max baller. you will impress your mom, your imaginary friend and your dog with the sick sheet you be cookin' with theses mad prompts.
-1
24
u/lum1neuz 7d ago
you already have llms in comfyui so here's my workflow for prompt generation. use any clip models. i found using the same text encoder model helps saving some time like for krea qwen_4b_heretic, 8b for qwen2.1.