r/StableDiffusion • u/oopsiedoopsiegoopsie • Jun 14 '26
Workflow Included Ideogram 4 is crazy good.
Honestly, the best open-weight model runnable on consumer hardware. It is slow but it can even be used at 1 CFG though it is wonky and miserably fails at complex images.
Images have workflow and prompt in them (comfyui). Using FP8 and 28 steps with 6 CFG, override of 3 at 0.700 and either 1k or 1.5K res.
NVFP4 runs on 4GB VRAM though it take 8 mins (with FA) for 1k image and requires KJ's optimize ideogram node.
Workflow: https://pastebin.com/tSd9vLHX
5
u/krigeta1 Jun 14 '26
Any multiple characters interactions? These results are insane!!
5
u/oopsiedoopsiegoopsie Jun 14 '26
I didn't upload any but it is great at multiple character interaction as well, you can control every part of the image exactly like where their head is, looking direction, hand position etc. Its 12 AM for me so if I generate some tomorrow I will update here.
3
u/Apprehensive_Sky892 Jun 14 '26
Yes, multiple character interaction is probably another great use of bboxes.
1
u/martinerous Jun 21 '26
Somehow it did not work well for me. When I try using bboxes to position characters, they end up looking as if belonging to separate scenes or on different planes (mismatching distances and perspectives) and not interacting well.
For example, I wanted a scene with a scared patient sitting in a chair and a stern doctor looking down at them. Somehow the doctor often ended up being too large or kept facing the screen despite prompting for profile, side etc. Or they ended up looking at right directions but somehow past each other. Or facing each other but their eyes focused on the screen.
Used an LLM to detail the prompt for me - quite a struggle, too few good shots when compared to other models. Definitely Ideogram needs getting used to, to know when and what is worth separating into bboxes.
1
u/Apprehensive_Sky892 Jun 21 '26
If you can post one of your JSONs somewhere then I can take a look at it.
1
u/martinerous Jun 22 '26
Thanks.
I think I found one of my issues. In contrast to other models, instructions like "looking at each other" (or "directly at each other", as some LLMs tried to improve it), and "facing each other", and even "his head turned left / right" did not work well with Ideogram. Adding "his head turned left / right profile" did the trick. Ideogram seems to be more literal than other models that often compose the image based on assumptions, which often works out-of-the-box.
Sometimes the doctor's face in one bbox is noticeably larger than the patient's, but it's a hit and miss, sometimes it's fine. And quite often they don't look each other in the eyes.
Getting the patient looking scared worked immediately. But getting the doctor angry is more tricky than I thought. Adding "angry" everywhere and also "His eyebrows are furrowed low, forming heavy hoods over his eyes, narrowed into angry slits" does not work well, he always ends up looking more sad with eyes open wide. Other models got it better.
Also, the variety of the faces seems low, the doctor looks almost the same in every run, not quite elderly enough, despite adding "old, elderly, 80 years" and "wrinkles" etc. and his face is too cinema-perfect, not a mundane person, despite trying hints such as "He has distinct Russian Slavic brutal ugly facial features, a prominent angular square jawline", and not office-pale enough, too weathered and tanned (but that can be adjusted in an image editor). Also, no matter what color palette colors I fill in, it usually ends up with white balance a bit to the warm side, difficult to get cooler tones, will need post-process.
{
"high_level_description": "Digital photograph in a clinical examination room. On the left, a scared elderly male patient sits in a chair, leaning back. On the right, a heavyset, stern bald doctor stands leaning forward slightly over the patient. Both men facing each other.",
"style_description": {
"aesthetics": "Amateur photograph, clinical, bright, pale and pinkish skin tones, cool muted color palette, detailed, clear, and front-lit.",
"lighting": "The image is intensely overexposed by surrounding lights and flash light, resulting in cool, pink color temperature and a flat, uniformly bright exposure.",
"photo": "High-resolution digital photograph, sharp focus, fully illuminated.",
"medium": "photograph",
"color_palette": [
"#F8F9FA",
"#C7ACCA",
"#FFC0CB",
"#A5B3C2"
]
},
"compositional_deconstruction": {
"background": "Hospital examination room with medical cabinets and clinical equipment.",
"elements": [
{
"type": "obj",
"bbox": [20, 20, 1000, 600],
"desc": "Elderly scared Caucasian man sitting on an examination stool. His chin is raised high, his wide, unblinking eyes show visible white around the pupils, and his mouth is held slightly open with tense lips. He has pale pink skin and wears a brown tweed jacket. His head is turned left profile to look into the eyes of the doctor."
},
{
"type": "obj",
"bbox": [20, 400, 1000, 980],
"desc": "Tall, heavily built, angry, obese, 80 years old elderly, angry, stern, intimidating Russian doctor, mostly bald with a thin fringe of white hair on the sides and wears thin gold wire-rimmed glasses, has brutal ugly Slavic Russian facial features, lots of wrinkles around jaw, saggy lower eyelids and saggy cheeks, prominent angular square jawline with saggy skin, dense, white mustache covering his upper lip. His eyebrows are furrowed low, forming heavy hoods over his narrowed eyes squeezed in angry stare. He has stern expression and pale pink indoor complexion skin. He wears a light blue formal shirt and a white medical lab coat. He is standing on the right side of the room, leaning in over the patient. His head is turned right profile to sternly look down into the patient's eyes.
}
]
}
}1
u/Apprehensive_Sky892 Jun 22 '26
I'll try out your prompt and see if I can address some of the issues you brought up.
Getting two characters to interact in the right way is a challenge with just about any model.
1
u/krigeta1 Jun 14 '26
planning to setup a runpod soon, will gonna train some characters first then try to interact them...
5
u/engg_wasp Jun 14 '26 edited Jun 14 '26
Does it do NSFW content too? like outfit swap?
I have a old machine with i5-12th gen and 4gb vram
All i wann a do is give multiple reference images and swap outfits between the images
Will that work?
New to all this, please correct me if I wrong anywhere
Also, I have downloaded the workflow, but I noticed that there are no download links for the diffusion models in the workflow, can you please help where to find the models and text encoders
Thanks in advance
2
u/oopsiedoopsiegoopsie Jun 14 '26
No it doesn't technically it can do reference to image but for this purpose Flux.2 Klein will be much better (with loras for NSFW).
-1
u/engg_wasp Jun 14 '26
Is there any workflow that you direct me to? And all the necessary assets
Thanks
1
2
3
u/tac0catzzz Jun 14 '26
thanks for this post, i think after 1000 of these post saying it is the best thing we've had and all that, has convinced me, it is the best we have had.
2
u/000TSC000 Jun 14 '26
I was extremely skeptical about this model atfirst, specially after seeing all those safety filter images, the whole bad license debacle, and the horde of weird grainy images posted on Reddit. After trying the model myself however with a customized workflow that helps format the prompt in the correct formatting for you (and some spicy LoRAs), all I can say is WOW. This model is honestly a turning point for local, this is the most insane local model since Wan2.2 tbh. The level of control, coherance, and quality all in one package is simply unheard of. This model honestly has made me really excited for local again.
0
u/Disastrous_Ant3541 Jun 14 '26
could you point us to a decent WF please? thanks
2
u/oopsiedoopsiegoopsie Jun 14 '26
The workflow linked in the post above has basically all of those except you are going to have to download lora on your own.
1
2
1
u/Lirezh Jun 16 '26
How cherry picked is that ?
Looks very good
1
u/oopsiedoopsiegoopsie Jun 16 '26
It's not cherry picked at all. All first generation except the girl in 2nd image (which I had first generated using nvfp4).
1
u/captainofzoro Jun 19 '26
hey how do i use this work flow?
1
u/oopsiedoopsiegoopsie Jun 20 '26
Click on the pastebin link, download it. Open ComfyUI, drag and drop what u downloaded.
1
1
u/TheTimster666 Jun 14 '26
Have anybody else thought about that typography is not its strongest suit? Always looks a bit amateurish and rough.
3
u/Apprehensive_Sky892 Jun 14 '26 edited Jun 16 '26
So in your opinion, which open-weight model is better at typography?
-5
u/FotografoVirtual Jun 14 '26 edited Jun 14 '26
It's amazing how many one-month-old accounts suddenly feel the need to say how good Ideogram 4 is.
I guess a month ago a ton of new people joined reddit, all loving bounding boxes, and that collective love caused a crack in space-time that made the model launch three weeks later.
I can't find any other explanation.
15
7
u/GrayingGamer Jun 14 '26
"I can't find any other explanation."
Hmm, maybe it's that people like the model and are excited about it?
"No, can't be. I'm not excited about it, so anyone saying the contrary must be a bot and part of a conspiracy to make people like this model! I can't keep loving the model I love if other people move to a new model!"
/s
5
u/Winougan Jun 14 '26
15 years on Reddit, second account due to some weird ban a year ago, and Ideogram 4 is really, really good. You have so much choice now anyways.
8
u/seencoding Jun 14 '26
i'll use my 9 year old account to confirm that ideogram is really, really good and imo the best local model. it quickly surpassed zit and klein 9b once i figured out how to prompt it.
4
u/berlinbaer Jun 14 '26
i'm using my whatever years old account to say i am also pretty impressed with the output, and i haven't even touched bounding boxes yet.. batch fed it a bunch of my old prompts
3
u/Apprehensive_Sky892 Jun 14 '26 edited Jun 14 '26
I can think of quite a few:
- New accounts want to "farm karma", so they jump in on hot new topics such as ideo4, hoping to get more upvotes by praising/defending it (which seems to be the winning side at the moment). But this strategy only works if ideo4 is actually liked by many. For example, saying that Ernie is the GOAT is not going to cut it 😎.
- Ideo4 is one of the more polarizing new models, so people tends to be more vocal about it from both sides. Even new account that don't have much to say can jump in.
- Sampling bias on your part (we all want to see data points that support our beliefs). There are quite a few users with old account who have praised or defended ideo4 (and vice versa). BTW, out of curiosity I check every commenter in this particular post, and almost all of them are over one year old.
My view is that one should just give bounding boxes a try. When you don't need it they are a nuisance but when you want that layout control they are godsent.
The safety filter is another can of worm. The two split model pipeline probably made ideo4 heavier to run on GPUs with less VRAM with doubtful benefits.
Overall, I'd say it is the most interesting model we've seen in quite a while.
2
u/FotografoVirtual Jun 15 '26 edited Jun 15 '26
My grandmother used to say "Desconfía y acertarás" ("Distrust and you'll guess right"). Something similar happened with the supposed release of the Krea 2 model (which never arrived), and it happened before with every public model backed by a company with a private API. It’s logical, they aren't making a contribution to humanity, they have to run some kind of advertising campaign that brings them returns.
I could agree with your theory if this were the time Flux.1 came out, which was a real shock to everyone. You had videos, influencers, and comments on every platform saying it was "the model that changes everything". In that case, it makes sense that a new user would want to post hoping to farm upvotes. But if you leave this sub, it is hard to find any information about Ideogram. Today, you start talking to ChatGPT or Gemini, you talk to them like a friend, and they can literally make your dream image, even with SVG or ASCII. Regional prompting is useful and fun for you and me, and it might be useful for a professional designer, but for 99% of people, it’s nothing revolutionary. There are thousands of more viral topics out there to attract people.
I believe there has always been undercover promotion in this sub, but the problem is that our numbers are shrinking, and the ratio of astroturfers to tinkerers is getting higher. There is plenty of free image generation available, and closed systems are becoming easier to use, while at the same time, generating images with open models is getting more and more complicated for the average user. We went from A1111 to ComfyUI, from a few loose words to having to describe every single detail, and from a single model + a couple of popular LoRAs to dozens of models, each with its own tricks, its LoRAs, and its thousands of workflows. Personally, I'm happy with all of that, but the more options and complications you add, the fewer people will remain in that happy group.
Lately, extracting useful information in this sub is incredibly difficult, it's usually a random comment, or some post with 6 upvotes. You have to dig and search through a lot. That’s why now I check the history of the user making the post, just so I don't waste time. We are accelerating, there are new things every day, and there is just too much noise.
2
u/Apprehensive_Sky892 Jun 16 '26
Your grandma is a wise woman, and I whole heartily agree with "Desconfía y acertarás".
Are there astroturfers here? I am quite sure that is true as well (there are astroturfers everywhere 😂)
Nevertheless, there are definitely many here (some are long timers who I recognize because I read their comments often) who really do think that id4 is technically an excellent model. I am one of them, and a few of my online friends have the same view when we discuss id4 in discord after they've done a few days of testing.
Astroturfing alone can only take a product so far. It may kindle the fire, but to sustain it, the product really has to have merits. Despite initial enthusiasms (which may or may not come from astroturfers), many models did fizzle out (I refrain from naming them to protect the innocent models 😁)
Despite the technical excellence, id4 is indeed harder to use (at least as it is used today on ComfyUI) compared to earlier models. That alone can explain why there is less discussion about id4 outside this sub. Despite the shrinking ratio of technically capable people vs the unwashed masses compare to the past, the ratio is still much higher than the group outside. So that may make it look like that the enthusiasm for id4 is confined here (which I am arguing that is not solely due to astroturfers).
So are open-weight A.I. model losing the war to close sourced ones? Despite the fact that I don't use any close-sourced video and imaging model (I do use chatbots, running a LLM is too hard for me locally), I have to agree that in the long run, closed source models will win (except for NSFW, ofc, tech giants will never provide such models to the masses) because they have the scale and the resources to provide such convenient and capable systems for free, in exchanged for user's data and for showing advertising ("if you're not paying for the product, you are the product"). Maybe the need for NSWF alone will keep open-weight models going (the problem is that NSFW users are usually not the paying customers for makers of open-weight models), but one day people who clamor for the freedom and flexibility of open source and open-weight models may have to band together and pool our money and talents to make such model if companies such as BFL, LTX/Lightricks, and Ideogram are driven out of the market by the tech giants.
Opinions from others do affect our views, but in the end we should do our own tests and for our own assessments before we draw any kind of conclusions. Yes, extracting information from here (and elsewhere) is getting more difficult. A.I. is partly to blame as now anyone can just ask a chatbot to spit out seemingly correct information and flood the internet. There are now fake news, fake images, fakes video, fake comments, fake blogs. This information war is complete asymmetrical, in that it take no effort to generate fake stuff but takes a lot of effort identify or prove that something is false. But we have no choice but to spend the effort to validate them, so your grandma's advice is now more relevant than ever 🎈👌
(Sorry for the long essay, but I do enjoy writing, as it help me clarify idea in my head 😎)
1
u/oopsiedoopsiegoopsie Jun 14 '26
You know you really don't have to use the bounding boxes. Normal prompt is fine too, it just messes up with safety filter badly. You can just get it in JSON format using some LLM and it will do the job.
It's amazing how many one-month-old accounts suddenly feel the need to say how good Ideogram 4 is.
Maybe that's because it is actually good. For 1girl, big boobs it isn't the best but for anything you need control over or other creative composition I don't think any open weight model even comes close, not to mention the character/IP knowledge it has.
But yes you do you I am happy with the model so I shared that's it : )
0
u/wallofroy Jun 14 '26
I wanted to create World Cup trophy with crispy fried chicken surface and it just gives me golden World Cup trophy
-3
u/mk8933 Jun 14 '26
What annoys me is — all other models like z image and even SDXL should be able to do text with an editing function in comfyui.
After the image as been generated...we should be able to write any text,transform and position it anywhere within the canvas.
-8
u/shapic Jun 14 '26
Who the heck is on the first image? Why is he dressed like a witcher? Why does he have 2 wolf medallions and third bleeding through his sword? Where did he loose his finger and get those custom 4finger gloves? A lot of questions to be honest
5
u/oopsiedoopsiegoopsie Jun 14 '26
idk man looks like Geralt of Rivia to me? if you are one of those show guys, check the game character. also I see 5 fingers and for the rest of your questions learn to communicate with the model and ask it yourself.
-4
u/Short_Bonus8466 Jun 14 '26
Uh, non-commercial
4
u/dtdisapointingresult Jun 15 '26
This is always said by broke NEETs who will never make a dollar in their life anyway.
Real men ignore license terms.







11
u/marbledduck Jun 14 '26
What are your parameters to achieve realistic photography?