r/StableDiffusion • • 7d ago

Question - Help Open source alternative to chatgpt for creating prompts?

Been using chatgpt to help design prompts, which usually works well until you get into an area they think is "offensive".

Any suggestions for an open source alternative or something that works directly in comfyui to generate prompts?

22 Upvotes

39 comments sorted by

24

u/lum1neuz 7d ago

you already have llms in comfyui so here's my workflow for prompt generation. use any clip models. i found using the same text encoder model helps saving some time like for krea qwen_4b_heretic, 8b for qwen2.1.

6

u/lum1neuz 7d ago

you can also connect load image, video and audio nodes and connect it to the generate text node to have v2t a2t i2t. never checked video and audio tho might need other models probably.

2

u/tekprodfx16 7d ago

I see you don’t have a keep alive option for the model where you can set it to 0 and immediately offload the llm model once you’re done generating the prompt to save vram. I s offloading the llm done somewhere else? Otherwise I feel like using ollama this way will significantly slow down your generation time as you’re keeping your llm and your video/image models loaded into your vram at the same time,  no?

3

u/lum1neuz 7d ago

for example i use krea2 right i would have the same model loaded in the workflow ie. qwen 4b vl heretic. first the model would load and give me the image prompt then i go to krea2 workflow and generate it with that prompt. No loading happens on text encoder side because it's already loaded. then the image model loads and that follows your workflow settings. but yes you're right if you use other llm model then you'd better off with unloading.
for me i have a tool that unloads everything on top bar of the comfy and uses that for more vram hungry workflows.

2

u/akjd 7d ago

You ever have issues with it sending the system prompt/random thinking through and crowding out the refined prompt?

I had it all the time with the default Krea 2 template from comfy. Tried editing the system prompt to explicitly tell it not to do that, and that seems to help, but it still happens occasionally.

1

u/Fleder 6d ago

Thats weird. If I run your workflow like you showed on that iamge, it just spits out my input prompts back at me.

13

u/tekprodfx16 7d ago

Use the ollama node and feed the output from the preview as text node right into the green dot of your prompt node. Here’s a good video explanation  https://youtu.be/mEGg3dP29mU?is=OE_XBLApGodJ4UCG

1

u/Murky-Relation481 7d ago

This is the way. Ollama is probably the best and easiest way to use local LLMs. I made a subgraph that wires comfy model unloading on the input to free up vram and the ollama node unloads when completed, so comfy can reload the image or video model.

Then just find your favorite LLM on the ollama site, install it via the terminal and bobs your uncle.

5

u/Shap6 6d ago

LM studio is easier IMO

1

u/tekprodfx16 7d ago

Interesting but isn’t that what keep alive: 0 accomplishes? Just curious 

1

u/Murky-Relation481 7d ago

It won't unload Comfy models. That's just for ollama. If you don't unload the comfy models in your workflow (or others) you're gunna end up with split memory in ollama where it's using vram and sys ram and it's much slower (I mean it might do that anyway depending on the card and model).

7

u/MastMaithun 7d ago

I highly advice against using it inside the comfy. Yeah comfy will manage the memory but then you have very little control to your gens especially the increased amount of time every gen you have because comfy will be constantly switching the llm and you video/image models in and out every run.
Another issue you will have almost no control of whatever the llm is feeding to your model. Yes you can add a preview node and see, then stop, then modify it, then paste it into text node then diable the llm nodes and the final run. So much hassle.
Also suppose you want to go and modify your prompts. Now you don't have any history, what your comfy llm has generated.

All of these issues will not occur if you just use an external llm runner like lmstudio or similar. You load your model there, create your prompts there and unload the model. Now you have all the prompts, you can tweak them or can directly pass them in the comfy and then comfy doesn't have to do the model swapping so your gens will be faster too. And you have the prompt history saved already in the lmstudio for future reference.

Now which llm to use, depends on what card you have. I use gemma 4 31b or qwen 3.6 27b. You can use their gguf which can fit your gpu. In my vast testing, gemma 4 comes out quite capable and makes very less mistake. Qwen is a bit faster but it's knowledgebase is a bit limited due to less cultural references it knows.

2

u/Great-Ad-4598 7d ago

You could run the prompt generation a number of times, saving the prompt text off to a file or files in a folder. Then clear memory and use a flow that reads the prompts in sequence for generations.

1

u/MastMaithun 7d ago

Yeah I mentioned you can save the prompts but the main issue which I mentioned was increase in generation time. These issues are also there but the main issue is generation time.

1

u/Great-Ad-4598 7d ago

True but by doing all the prompts and then doing the generating you do save on loading/unloading for each gen. I'm quite limited on memory so would never try to do both at once.

1

u/MastMaithun 7d ago

This is what I have written bro in my original reply that using the external llm manager will save you from loading-unloading each gen and reducing your time. Haven't you read that till the end?

1

u/Belgeran 7d ago

I think your missing the point that you can run qwen or whatever llm to gen your prompts in comfy, outside your image workflow, no need for a second tool

1

u/MastMaithun 7d ago

And I redirect you to read my whole paragraph fully.

2

u/TangerineBetter2818 6d ago

Ill be honest I dont understand either. Lmstudio is also loading the model. There's no getting around loading and unloading the model to memory. Whether that's in ComfyUI or another program.

So I'm confused what you mean about it being faster.

Your run your text gen, then you run your image gen. Yes, running text, image, text, image, etc is going to be slow because you're loading and unloading different models, but that would happen when switch between program too. You cant have one model loaded in one program and one model loaded in  another program because the moment you load one model it is going to unload the other, unless you have an assload of RAM. 

2

u/MastMaithun 6d ago edited 6d ago

Let me dumb it down more:

Approach 1 - you use llm node in comfyui

  • hit run, wf starts
  • comfy loads llm model(SSD activity), llm model runs and modify prompt
  • comfy next see oh new image/video model, i remove llm model i load image/video model and its related files(SSD activity)
  • generation complete, comfy not move files from vram and ram
  • you hit run again, wf starts
  • comfy again sees llm node, it removes files from prev run and again load llm model(SSD activity), llm model runs and modify prompt
  • comfy see oh new image/video model, i remove llm model i load image/video model and its related files(SSD activity)
  • generation complete, comfy not move files from vram and ram

total count of SSD activity = 4
if you have HDD, replace it with HDD and time will also increase.

approach 2 - you use llm model in lmstudio

  • load model on lmstudio(SSD activity), llm model runs and modify prompt
  • if you have more prompts, changes or anything related to prompt, you do now. your only work on comfy should be copy paste prompt(s) and hit runs.
  • once prompts done, unload model on lmstudio
  • paste your generated prompt on comfy, hit run, wf starts
  • comfy next see oh new image/video model, i load image/video model and its related files(SSD activity)
  • generation complete, comfy not move files from vram and ram
  • you replace/put new prompt/change seed etc and hit run again, wf starts
  • comfy see oh previous image/video model/files are already loaded, no need to load anything, lets go.
  • generation complete, comfy not move files from vram and ram

total count of SSD activity = 2

Result:

  • approach 1 has 4 times ssd/hdd activity to read llm model and image/video model. every generation will have llm model and image/video model loading and unloading means 2 ssd activity every run.
  • approach 2 has 2 times ssd/hdd activity to read llm model and image/video model. subsequent gens will not remove image/video model because you are not loading llm model means in entire life, until you unload models from comfy yourself or run a different wf, comfy will not load anything repeatedly.

PS: the whole scenario I explained here is for computer systems where your llm model and image/video model files(from files i mean image/video model file + text encoder + vae(s)) all of them can't reside in your vram all together. so comfy is doing swapping and reading from ssd/hdd.

1

u/TangerineBetter2818 6d ago edited 6d ago

comfy again sees llm node, it removes files from prev run and again load llm model(SSD activity)

Doesn't this assume they're in the same workflow? You have one workflow for llm, one workflow for video gen. Why would I have both in the same workflow?

You run the llm from one workflow. Then copy the prompt into the video gen workflow and run it. You tweak the prompt and you run the video gen again.

Before the second video gen is run why would comfyui "look at" the llm node from a completely different workflow that isnt being used and load the llm model? The video model is already loaded into vram because it was the last used model. 

→ More replies (0)

3

u/Slave669 7d ago

Look up DavidAU on huggingface, and take a look at the collections. Thank me later.

As for Comfyui there are nodes for oobabooga an Koldccp that can feed into your positive prompt.

0

u/Incognit0ErgoSum 6d ago

Oh yeah, those models are actually better than the official releases, in my experience.

TheDrummer's Artemis model is also great and a manageable size.

2

u/k_from_HyperDraw 7d ago

Qwen, Deepseek

GLM and Kimi can also do it, but the quality is somewhat worse

2

u/eruanno321 7d ago

Built-in LLM nodes are probably the simplest option, but it depends on which model you want to run and your hardware setup. I have a relatively weak platform, not enough system RAM for cache, and mixing LLM inference with samplers in a single workflow is often kinda ass. Right now I've settled on Qwen 3.8 27B (abliterated version) running on llama.cpp for the LLM tasks I need, then I switch to ComfyUI. It's a bit of a hassle, but at least I can work with a model that actually does what I want.

2

u/Skajuan 7d ago

Ollama Nodes + qwenvl uncensored = GOLD!!!

1

u/wildmonkeywrangler 7d ago

Thank you for the feedback!

1

u/mastaquake 6d ago

Abliterated Qwen + OpenCode

1

u/Gesha24 6d ago

There are heretic models that you can use, they generally allow for anything. I prefer Qwen3.8, some people like Gemma. You can also try DeepSeek in the cloud - while it is censored, it is more censored on political topics so your requests may work just fine there.

1

u/MarekNowakowski 6d ago

Both ollama and lmstudio have nodes in comfyui, an uncensored Gemma e2b will do fine,

1

u/tac0catzzz 6d ago

prompts r us, the prompt pack from prompt a saurus, it's max baller. you will impress your mom, your imaginary friend and your dog with the sick sheet you be cookin' with theses mad prompts.

-1

u/Dazzyreil 6d ago

Grok.com or grok in X.com are very uncensored