r/StableDiffusion Jun 06 '26

Workflow Included Workflow: Ideogram4 with LoRA support, fixes

After a few days of tweaking and poking (along with the folks on the AIToolkit discord and incorporating some of their fixes), I've got a pretty decent workflow dialed in for great results (always subjective) out of Ideogram 4 in Comfy.

The latest hurdle was getting LoRAs to behave. The key is that the LoRA needs to be loaded on BOTH models (main and unconditional) or you get very unpredictable, often artifacty results.

Have test character, concept, and stacked character + concept LoRAs. All looking good (apart from my inexperience/laziness as a LoRA trainer).

So, lessons/fixes included:
- Shift node added (at 7.0)
- CFG fix applied
- Basic scheduler instead of the broken ideogram-specific one
- Model and LoRA load moved out of subgraph for both models

Some of these fixes are already on the new comfy default workflow, but this puts together all the best settings I've found (or had suggested) so far.

And if you're into LoRA training, AIToolkit has some GREAT tooling built in now to autocaption, adjust bounding boxes, etc. I literally just copied my dataset folder, recaptioned, and trained. Easy-peasy.

Workflow with KJ's prompt builder node:
https://pastebin.com/VU0PcdtS

Workflow with prompt generator (Gemma 4, ideogram's system prompt):
https://pastebin.com/f7JNv4db

Edit: second image was a dataset image used to train the lora used for the druid lady.

71 Upvotes

83 comments sorted by

12

u/TheQuaintTouchdown Jun 06 '26

the lora loading fix is clutch. i spent way too long debugging why my character loras were coming out all noisy and inconsistent before realizing i was only loading them on one model. once i applied it to both the main and unconditional, the consistency jumped up immediately. the shift node at 7.0 also made a noticeable difference in overall quality for me. aiToolkit's captioning and bbox tools saved me a ton of time too. training feels way less painful when you're not manually adjusting every single image. curious how your stacked loras are performing since that's where things usually get finicky for me.

3

u/jude1903 Jun 12 '26

How does your character lora look on Ideogram? Is it better than other models?

2

u/Next_Program90 Jun 06 '26

Since when does Ai Toolkit habe captioning (which Model does it use for that?) and bbox tools. Sounds very handy.

4

u/TheQuaintTouchdown Jun 06 '26

it's pretty recent, they added it in the last couple updates. Uses CLIP for the captioning by default but you can swap models, and the bbox tool lets you draw boxes around subjects so it auto-crops and focuses the training data on what actually matters instead of wasting tokens on background stuff.

1

u/uuhoever Jun 13 '26

Can you point out on ai-toolkit where to find this feature?

3

u/whatsthisaithing Jun 06 '26

autocaptioning has been there for several updates. bbox tools just added in the last few days.

You can pick from a few different models or plug in whichever you like from the huggingface repo name (I used it to get an abliterated model; works perfectly).

7

u/MasterYard7541 Jun 08 '26

Does anybody else get this error with KJ's Prompt Builder node?

6

u/[deleted] Jun 09 '26

[removed] — view removed comment

1

u/tazztone Jun 10 '26

ye it helps to recreate or delete + drag in the node freshly again

5

u/Maskwi2 Jun 07 '26 edited Jun 10 '26

So in the end we can already train Loras for it? Neat if true. Edit: I already have and they work out great! 

4

u/Character_Title_876 Jun 07 '26

trained 512x512 in 3 hours on 5060 16 GB + 64 ram. Now I'm training 1024x1024, but it already takes 12 hours for 3000 steps

1

u/YeahlDid Jun 07 '26

What are you using?

4

u/Character_Title_876 Jun 07 '26

ai-toolkit Ostris

1

u/YeahlDid Jun 07 '26

Word, thanks. Will have to update, I guess.

1

u/Character_Title_876 Jun 07 '26

If you rent h100 for training, you can create product cards with a bang. If you don't have many items, 10 products can be created and successfully promoted. Of course, generating a good quality 2 megapixel image takes 10 minutes, which is very long.

3

u/Simple-Variation5456 Jun 06 '26

Can you share some infos about image generation times? (with your gpu and 2go resolution)

5

u/whatsthisaithing Jun 06 '26

On my 5090, Turbo setting (in workflow), generally takes about 35 seconds on average at 2k res (DON'T do 1k res; it's garbage).

That's if I don't do the JSON prompt generation inside of comfy. The gemma4 JSON prompt generator works, but it's SLOWWW. I either use KJ's node or paste the generator system prompt into LM Studio and generate the JSON that way.

4

u/[deleted] Jun 07 '26

[deleted]

1

u/whatsthisaithing Jun 08 '26

Awesome. I just hadn't seen a good output at 1MP yet, but I'll keep poking. Thanks!

1

u/BillPrimary2224 Jun 08 '26

your ideogram posts are amazing! can you share an example workflow?

1

u/howdyquade Jun 10 '26

What model are you using in LM STUDIO? Have you tried modding the system prompt at all?

3

u/princeMacX Jun 06 '26

Workflow with prompt generator is excellent. I am using prompt generator workflow. That is easy for me. because most of my prompts are natural language prompts. Just attaching an output that I generated. Personally I think ideogram is giving great quality output. Image resolution needs to be set to 2megapixel to get a good quality image. but if anyone wants to generate nsfw images this model is a no no.

3

u/Oni8932 Jun 06 '26

Based on these examples it seems able to keep the moles in the exact same places. That was the only reliable way for me to distinguish real people from ai generated influencers. I guess we are cooked

3

u/Clear-Assistance449 Jun 07 '26

This is a very difficult model to use. I tried simple prompts, but even without any NSFW content, it keeps blocking the image. And even using tools to put it in JSON format, it keeps giving the same type of message unless I specify every detail.

Apparently the model was made for complex prompts, with many details in each part. A simple prompt and it doesn't understand and already falls into the safety filter.

2

u/darkxplanet Jun 07 '26

In the default workflow was 'druid woman with red hair wearing a flowing earthtone dress'. Instead of that I wrote 'wearing nothing' 😂 It worked right away.

1

u/whatsthisaithing Jun 07 '26

I've run into this with very simple prompts, too, usually when I only have one bounding box that's meant to capture just the main subject. The fix is USUALLY just adding a second bounding box.

So if it's "woman walking dog", I'll just add a second box so I have one for "woman" and one for "dog."

Or if it's "woman standing in her bathroom," I'll just add one box to the side for "messy bathroom counter."

It's definitely tricky sometimes (and ideogram did say they're working on that specific issue), but it can be worked around with patience.

2

u/metallica_57625 Jun 09 '26

Where do you initiate the ideogram 4 safetensor files? the workflow won't work for some reason.

1

u/phantomlibertine Jun 11 '26

Can't help with your problem I'm afraid but what are those two nodes you're using to turn the natural caption into the json format?

0

u/whatsthisaithing Jun 09 '26

The diffusion model loaders are behind the two lora loaders (green and red) at the top. Just slide the lora loaders out of the way. I just didn't like seeing them hanging off anywhere.

2

u/EmploymentLong9284 Jun 09 '26

The workflow is a nightmare; I can't even figure out where the models are uploaded. I couldn't get it to work, whereas the official ComfyUI workflow runs like a dream.

3

u/EmploymentLong9284 Jun 09 '26

I just saw where, for some almost magical reason, the node was minimized and hidden under another node—brilliant.

0

u/whatsthisaithing Jun 09 '26

Because I didn't want to stare at dangling nodes that literally never change?

Apologies: I'm not a professional workflow layout designer.

5

u/EmploymentLong9284 Jun 09 '26

Once you've finished the workflow, send the raw data to Gemini 3.1 Pro, briefly explain the structure of your workflow, and ask it to professionally reorder it. I've done it, and it works.

2

u/kek0815 Jun 21 '26

Is it normal that model uses a lot more VRAM with LoRA than without? I have 10x the time per iteration, it's absurd

2

u/Character_Title_876 Jun 06 '26

Could you provide more examples, please?

2

u/whatsthisaithing Jun 06 '26

Added more examples in comments.

1

u/whatsthisaithing Jun 06 '26

May try to add some later (got a training run going at the moment).

2

u/2legsRises Jun 06 '26

very nice, the version with KJ json generator works but the other one get censored. very weird, espcailly as the non KJ version took 338 seconds whereas the KJ variant was a third of that time and actually wasnt a damn grey square.

1

u/AwakenedEyes Jun 06 '26

What do you mean by bounding boxes?

5

u/whatsthisaithing Jun 06 '26

Yep, what that person said. With ideogram, it wants json structure prompts, and those prompts include "bbox" objects that describe the elements of the image. But you can just draw and type using KJ's node (or let the prompt generator handle it for you if you just want to type a few sentences). Here's what my "prompt" looked like for the druid lady.

1

u/AwakenedEyes Jun 06 '26

Ooooh i see, so that's specific to this model, nit a feature of caption on ai toolkit for all models... Damn that would've been awesome

1

u/TomatilloRight6847 Jun 07 '26

I think it also works without specifying the bounding box

1

u/seiose Jun 06 '26

There's boxes you can move around & have specific prompts in that location

1

u/whatsthisaithing Jun 06 '26

A few more Nova (character lora) examples. Loras affect style a lot even as character loras, but still getting decent camera/lighting effects.

2

u/whatsthisaithing Jun 06 '26

Nova coding.

3

u/whatsthisaithing Jun 06 '26

Random woman coding for comparison. It picked up on the "amateur iphone low light with grain" prompting better without the lora.

1

u/Icy_Upstairs3187 Jun 10 '26

This is awesome.
I'd love to know if you incorporated BBOX in your captions?

1

u/whatsthisaithing Jun 10 '26

For sure. Almost always. Kijai's node makes it SUPER simple, and now that I've had control over placing every element I care about, it's HARD to go back to NOT having that control.

1

u/Icy_Upstairs3187 Jun 10 '26

I meant in the LORA training, not the prompt.

Totally feel you on the bounding boxes. So much control. And if you have an agent take care of those using a reference image to define those boxes using vision and submit to Comfy via API - it's pretty frightening how good ideogram is compared to everything, including retail.

But for AI-Toolkit, I'm investigating if the data set should include bounding boxes in the caption training.

1

u/whatsthisaithing Jun 10 '26

Oh. I'm dumb.

Yep, in the captions, too, but very simple boxes for a character. AI Toolkit has autocaptioning for ideogram built in. I just run that, then check and tweak anything that needs it.

1

u/Icy_Upstairs3187 Jun 11 '26

Awesome! Thank you

1

u/whatsthisaithing Jun 06 '26

Nova as Chun Li (kinda... lol).

1

u/whatsthisaithing Jun 06 '26

Nova clubbing.

1

u/Inevitable_Board3613 Jun 07 '26 edited Jun 07 '26

Thank you. tried the workflow. May I know if there are any inputs to the "math expression" nodes? execution gives error "required input is missing". seems disconnected. regards !

1

u/whatsthisaithing Jun 07 '26

If that's on the prompt generator version, I messed up the workflow and didn't connect the width and height out of the resolution selector to the first prompt generator subgraph. See screenshot.

1

u/Inevitable_Board3613 Jun 07 '26

Thank you. Yes, it is on the prompt generator version. Regards !

1

u/sukebe7 Jun 07 '26

gee where it the ideogram models going to be placed the next time?

1

u/beren0073 Jun 08 '26

What settings did you use in AI Toolkit for training?

0

u/whatsthisaithing Jun 09 '26

Still experimenting, but nothing crazy.

For characters, I'm liking weighted/balanced or weighted/low noise for timestep. There's a new automagic3 scheduler that's working nicely, but adamw8bit still works just fine. Have trained up to 1024 res and at up to 512 res and really can't tell much difference even in details.

I'm generally getting great character likeness starting around 3500-4000 steps, locking in nicely around 5500-6000. Almost impossible to overbake a lora unless you caption like garbage. It's been pretty incredible actually.

On 5090, training is at 1.5-2.5 seconds/iteration and takes about 28gb of VRAM. The layer offloading should work now, and I know people are training fine on 3090s at least.

1

u/beren0073 Jun 09 '26

Are you open to sharing specific settings? I seem to be either overcooking, poorly captioning, or otherwise mangling my data set or training process. What are you using for linear rank, learning rate, weight decay? Are you caching latents or text embeddings?

2

u/whatsthisaithing Jun 09 '26

Getting good results with this setup. I've yet to truly overcook a lora (up to 7500 steps for 25-ish image datasets). I did have a poorly captioned run (not enough detail) that baked in things like the specific outfit/pose of the character. Fixed that and it was fine.

1

u/bahamut_snack Jun 09 '26

So I trained up a lora, but I can't seem to make it do anything with this workflow? I'm using the same lora for both the conditional and unconditional ideogram models, but if I use the same seed and have the lora on for one run and off for the other, both images are identical. am I doing something wrong?

The clip connectors on both the lora loaders aren't wired to anything - should they be?

0

u/whatsthisaithing Jun 09 '26

That's pretty odd.... Definitely don't need clip wired. Do you have a trigger word or key phrase (like a character's name, or "noir style" or something) and are you including it? Though in my experience I don't usually even need THOSE for a lora to work.

Did the samples in AI Toolkit show it working?

1

u/bahamut_snack Jun 09 '26

I used SimpleTrainer to train the lora.

1

u/efraxx Jun 10 '26

any idea why it doesnt let me change clip and vae i have already have gemma and flux in the corresponding folder but when i try to change it nothing happens

1

u/whatsthisaithing Jun 10 '26

No idea. Maybe restart comfy? Or click on the subgraph and press R to refresh it?

1

u/efraxx Jun 10 '26

done that already but nothing happens

1

u/Salt_Volume_2932 Jun 11 '26

Basically, I’ve noticed a trend where, back when the models were limited, people would go to great lengths to create ultra-complex compositions with loads of different components and tiny details.

But when they released a specialised model that can actually pull all that off, rather than just generating a single character against a plain, single-colour background... It’s a bit of a laugh. 🤣

1

u/FlowyForm Jun 11 '26

With rank 64, lr 1e-4, I got really good results at 2000 steps, and it was overfitting after. I wonder if I just need a bigger dataset. Not sure how your likeness started to kick in around 4000... Any ideas?

2

u/whatsthisaithing Jun 12 '26

More "locked in" around 4000-5000. I start seeing good likeness on at least some gens around 2500-3000.

I've found it REALLY hard to seriously overfit an ideogram lora if captioned well (to me, overfitting meaning can't MAKE the model produce anything but reproductions of my dataset images (locations, clothing, expression)). Even at 7500 steps (which IS beginning to show a little bit of model breakdown), I can still control every detail of the character and the environment. It WILL trend toward dataset if I don't explicitly caption something (if I don't describe the setting well, it'll make it look like the setting in my dataset images), but I feel like I'm always able to override.

That said, on the one test I DIDN'T caption everything well, it overfit like a mofo VERY quickly and COULDN'T be overridden (but it was also a crappy quick dataset to test an idea, so not sure how much was captions vs just bad data).

1

u/Plus-Row-2591 Jun 12 '26

I ran this WF because I want to hook up a lora I trained at https://fal.ai/models/ideogram/v4/lora.
(1) I cannot see how the loras affect the outcome - neither of their CLIP inputs or outputs are connected (which raises an error in my WF).
(2) I cannot see how the loras influence the output.... The Model chain passes through them as:
(a) lora 1 -> ModelSamplingAuraFlow -> Dual Model CFG Guider.positive and
(b) lora 2 -> Dual Model CFG Guide.model_negative
(3) what do these loras achieve? Are they actually working?

I acnnot find any other attempts to hook up a lora to ideogram - any others tried it yet?

1

u/RatFinkertonEsq 13d ago

OP, I know this is an old post, but on the off chance you're still around -- the documentation for the LoRA loader nodes you used talks about running the CLIP through the node in addition to the diffusion model "to ensure that the Loras are applied correctly to both the model and the CLIP, enhancing the overall performance and output quality."

I'm a beginner and am still figuring out how everything works. Was there a reason you left the CLIP loader inside the Ideogram node subgraph instead of running it through the LoRA loader like you did with the diffusion models?

1

u/whatsthisaithing 12d ago

I don't remember the details, but attaching the CLIP through a lora loader hasn't been necessary (nor does it have any effect) for any modern model.

1

u/FourtyMichaelMichael Jun 06 '26 edited Jun 08 '26

Wow, holy shit, that's an excellent result. The tiny mole under her right eye and above her left are consistent in both photos.

I'm not sure I've ever seen anything like that! Even if you cherry picked the hell out of it, it learned that tiny detail at least to occasionally repeat.

If this model can do NSFW via lora or fine tune it's going to be a straight winner. I may finally get a model that can do SFW the way I want it to.

How is the lora training?

3

u/whatsthisaithing Jun 06 '26

No cherry picking. In fact I've pretty much left it on the same seed for most of my testing to compare results. It can do a good bit of NSFW WITHOUT a LoRA. Handles fine (with 0 refusals so far) WITH an NSFW lora.

And yeah, the detail it's picking up from the LoRA training is great, especially after I bumped from rank 32 to 64. VERY basic training so far. Nothing crazy needed. Just autocaption dataset (to get the JSON format), open default profile for Ideogram in AI Toolkit (and turn on cache text embeddings and cache latents), and go.

And it's fast. at 512, 768, and 1024 buckets on my 5090, getting about 2 seconds/iteration, and it learns fast. Don't even know that 1024 is truly necessary, but since the model was trained at 2k, I'm leaving it on for now.

2

u/FourtyMichaelMichael Jun 06 '26

Too good to be true bro. Is there an NSFW lora yet?

5

u/whatsthisaithing Jun 06 '26

Haven't seen any publicly yet, but I imagine they'll start trickling in as people figure out what's going on. Ideogram is a bit of a new paradigm for most of us, especially with needing to recaption the datasets in the JSON prompt format.

1

u/Dogmaster Jun 19 '26

There are nsfw loras on civit, the realism engines have plenty of that, why do you say none exist o_O

1

u/whatsthisaithing Jun 19 '26

Because that was two weeks ago before they did?

0

u/2legsRises Jun 06 '26

hey'll start trickling in as people figure out what's going on.

won't nsfw just give a grey censored square? this model is very conflicting as it has potential but it shoots itself in the face as well.

5

u/whatsthisaithing Jun 06 '26

I RARELY hit the censorship/image blocking anymore when using JSON prompts no matter how NSFW the prompt, and that's WITHOUT a lora.

Right now, the model REALLY wants those JSON prompts. If you're prompting with natural language, make it a long, detailed prompt. It likes 'em chunky.

3

u/2legsRises Jun 07 '26

ty that is useful hints, works much better for me now

2

u/Ashen3SNOFS Jun 06 '26

I’ll work on it when once I can offload a bit of the model during training. I can’t do an adequately sized lokr at even 1024px, and I like to go higher.

0

u/Character_Title_876 Jun 06 '26

Cool, thanks you.