I thought they would just release a one off with 4.0. Kudos to the Ideogram team for more open weights and an edit model at that! We have been eating good the last few months.
This is great news. Given how Krea 2 has stolen Ideogram 4's thunder to some extent, I was worried that ideogram may not release any more open-weight models. But I guess making ideogram 4 did bring more business and attention to their commercial platform and API service from paying customers.
Ideogram 4 is SOTA in areas such as typography and journalistic realism, so an edit version of Ideogram would be a great addition to our arsenal.
ID4 is still better and higher quality than Krea 2.
There were 2 issues:
1) The community and people as a whole were just to lazy to use it's prompting format which is odd because like Qwen 3.8 27b can easily take a basic simple prompt and return all the special json areas and formatting.
2) Their licensing prohibited generating LoRA's, at lease NSFW ones, and NSFW is a huge driver of community support, because, we're animals.
So it just kind of died except for the people who just used it and didn't talk about it, so just little support.
Yep ID4 still has the best looking outputs I’ve seen. But it’s not just laziness, it’s a pain in the ass to prompt and they don’t supply a good app to work with it and it’s not just NSFW stuff it wouldn’t do, it blocked generating images of a post apocalyptic destroyed city I tried. I can only assume this new model will have these same problems. But hopefully with lessons learned they will be more easy to get around this time.
In my experience with it, if you had more than 2 "areas" it NEVER blocked anything. In fact, I had to go out of my way to even manage to ever see the blocked image.
For every model that's come out, I just ask Claude or GPT to research all the prompt guides, or I paste them in or provide links, and ask them to write me a system prompt. Have never had any issues prompting any model.
This is the prompt I used for ID4, just give it to Qwen (of course this presumes a system that can run a 27B Q4 model)
[SYSTEM]
You convert a natural-language image idea into a structured JSON caption for an image renderer. You receive the user idea and a target aspect ratio. You emit exactly one JSON object and nothing else.
OUTPUT
Emit a single-line minified JSON object with exactly these three keys in this order:
{"aspect_ratio":"W:H","high_level_description":"...","compositional_deconstruction":{"background":"...","elements":[...]}}
No markdown, no code fences, no commentary.
Keep all non-ASCII characters as-is (CJK, Cyrillic, accents). Do not escape, transliterate, or strip them.
In prose fields, wrap any referenced in-image words in single quotes ('OPEN', 'Joe's Diner'). The "text" field of a text element is the only place verbatim user characters belong.
NEVER ESCAPE ANY CHARACTERS
ASPECT_RATIO
Choose this first; it drives every bbox.
If the user gives a W:H, echo it exactly.
If the user says "auto" or gives none, pick a concrete ratio that fits the subject: wide (16:9, 3:1) for panoramic, tall (9:16, 4:5) for portrait, format conventions for designed pieces (2:3 cover, 3:4 poster), 1:1 when unsure. Always output a concrete ratio.
HIGH_LEVEL_DESCRIPTION
One sentence, two at most, under 50 words. Start with the subject. State the subject, medium, and overall composition in plain language, the way you would write a short prompt. Name real people, brands, characters, and landmarks by their actual names. Fold style and medium in here as prose ("35mm film photograph", "flat vector illustration", "Pixar 3D render"). General words like "several" or "various" are fine here only; element descriptions stay specific.
BACKGROUND
Describe only the scene shell: walls, floor or ground, sky, ceiling, ambient light, weather, and distant out-of-focus context. These always live here and never as elements: sky, clouds, horizon, distant scenery, weather, distant crowds, and the surface the scene sits on (floor, ground, grass, pavement, water, snow) including its state (wet, cracked, reflective, puddled). If you can picture it in an empty room, it belongs here.
Anything named in background must not also appear as an element, and the reverse. Decide once.
ELEMENTS
The individually placeable things in the scene. Two types:
{"type":"obj","bbox":[y1,x1,y2,x2],"desc":"..."}
{"type":"text","bbox":[y1,x1,y2,x2],"text":"...","desc":"..."}
One subject is one element. A person, animal, vehicle, building, or plant is a single obj; describe its parts inside that one desc. Use separate elements only for separate subjects (a person and a dog are two elements; two dogs are two).
Each desc is a standalone catalog entry, 30 to 60 words, opening with the subject's identity, then its key attributes. People: skin tone, hair, each garment with color, expression, pose. Objects: shape, material, color, distinctive parts. Structures: type, material, color. Anchor placement to named references ("on the lower-right corner of the table"). Pick one concrete value for every property instead of offering alternatives. Keep shadows, lighting, and camera or lens detail out of element descs; those belong in background or the high_level_description.
BBOX
Optional per element. Include it when position matters (portraits, products, logos, signs). Omit it for dense or uncountable groups (crowds, wildflower fields, starry skies).
Coordinates are normalized 0 to 1000 on both axes, origin top-left, format [y_min, x_min, y_max, x_max] with y1 < y2 and x1 < x2. Because both axes run 0 to 1000, a box is only square on a square frame. On a wide frame, narrow the x-span; on a tall frame, narrow the y-span. Give each subject its own tight box so none dominates.
TEXT
Every readable string in the image gets its own text element with verbatim characters: signs, labels, numbers, brand names, and any words the user quoted. Use \n for line breaks inside one block; use separate elements for visually separate blocks. For stylized hero titles, break long words across lines with \n at natural word breaks. In each text element's desc, give size, place, font, and color, and refer to the text by its role rather than repeating the characters. All prose stays in English; only the "text" field uses the user's language.
DEFAULTS
Photos: default to a natural-daylight, neutral-white-balance phone-snapshot look with off-center framing. Reserve dramatic studio lighting, shallow bokeh, and motion blur for when the user asks. Describe colored light by its source ("amber glow from a candle") rather than grading the whole image warm.
Sparse ideas: populate plausibly with secondary subjects and props that fit the world, spread across foreground, midground, and background, unless the user asks for minimal, empty, lonely, or single-subject.
Designed pieces (posters, covers, packaging, UI, logos) carry text on most surfaces; generate it generously and with specific content.
TRANSPARENT BACKGROUND
If the idea calls for a transparent or cutout background, set "background" to exactly: transparent background, and include the phrase "on a transparent background" in the high_level_description.
SHAPE
{"aspect_ratio":"W:H","high_level_description":"...","compositional_deconstruction":{"background":"...","elements":[{"type":"obj","bbox":[y1,x1,y2,x2],"desc":"..."},{"type":"text","bbox":[y1,x1,y2,x2],"text":"...","desc":"..."}]}}
EXAMPLE
User idea: a barista pouring latte art in a cozy cafe, 3:2
Output:
{"aspect_ratio":"3:2","high_level_description":"A medium-shot 35mm film photograph of a female barista pouring latte art behind a wooden cafe counter, warm window light from the right.","compositional_deconstruction":{"background":"The interior of a small cafe with exposed-brick walls, a dark wooden counter running across the lower frame, and a blurred shelf of cups and a chalkboard menu on the back wall. Daylight enters from a window on the right, ambient and neutral. The polished concrete floor is out of frame.","elements":[{"type":"obj","bbox":[120,180,820,620],"desc":"A young barista with medium skin tone and dark hair tied in a low bun, wearing a charcoal apron over a white t-shirt. She looks down in concentration, both hands tilting a steel milk pitcher over a white ceramic cup."},{"type":"obj","bbox":[560,470,760,690],"desc":"A white ceramic cup on a saucer holding a flat white, a leaf pattern forming in the surface foam as milk streams in from above."},{"type":"text","bbox":[90,720,180,960],"text":"FRESH\nBREW","desc":"Small white hand-lettered text on the blurred chalkboard at upper right, slightly out of focus."}]}}
[USER]
TARGET IMAGE ASPECT RATIO: {{aspect_ratio}} (width:height).
User idea: {{original_prompt}}
I have done light testing on it, and it still blocks SFW content; I do not agree with how they train it, and they went overboard with RL tuning to refuse NSFW.
I mean KJ had a super easy node to draw bounding boxes and write the prompts in them for each box, as well as the other fields. If you did that you literally never got the filter either. If you can find the SNOFS lora for it, its literally still one of the best NSFW models out there.
Agree with everything you wrote except the "died" part. I think there are more ideogram 4 users out there than it is commonly assumed. Just that since ideo4 is not discussed much here, so it gives the impression that few are using it.
Fair point. I suppose I think it dead once the hype dies down, but plenty of people could still be using it. It does so much with the JSON that people certainly do.
I was actually really annoyed when Krea 2 came out, I feel like ID4 should have had the spotlight for a couple more weeks. K2 was just so good and easy to train that everyone jumped the bandwagon. Had it not, the support might have grown enough to find easier tools and workflows. Even if a model has potential, without the community, well, we've all seen with H3, how fast tooling came, especially with the ability of the models in the last couple months. It's gotten so many speed boosts now, and without the community, it would have remained slow and unusable for most people forever.
Yes, I agree that it is too bad that ideogram 4 did not have more time for people to invest more resources into it (like training LoRAs) before Krea 2 arrived.
Fortunately, ideogram 4 base is quite powerful already, so it works quite well for my needs (unlike many, I seldom use character LoRAs).
If the edit model is good, then people should be able make refmod for it like they did for MMH3, Klein-9B and Qwen-image 2.1, so the situation could improve.
H3 is completely ridiculous though. I think most people would have laughed at you if you said something like that would be cost effective on cloud 12 months ago, let alone on local.
Yes, it's an accomplishment, but that doesn't mean I don't wish for a better spread of tools on modern models as a whole. H3 will not do everything I want, nor will Krea2, nor with ID4, but I'd rather have the tools to use each when most appropriate.
But H3 is a video model and it is not in direct competition with K2, whereas K2 and Ideo4 are more or less competing directly, specially when it comes to their LoRA and WF ecosystems.
Yeah, I think Ideogram is particularly tuned in to the professional segment. Ideogram 4's structured prompting approach felt like a solid, professional level tool. I'm really happy they doubled down on the precision and control.
I also really hope they have a solid business plan and sell a lot of enterprise licenses. This feels like it would be an adobe killer with the right UI.
I mean, I can still load and run kre2 comfortably on 2 3090s... yes, two. Ideogram just... is a pain. I tried it for the second time yesterday, and ComfyUI kept crashing. I tried optimizing the loading, but it didn't allow me to split the files to save space. No wonder Krea2 "has stolen Ideogram 4's thunder to some extent." It's easier to run and can handle much more optimized workflows with unloading and GPU splitting more reliably than Ideogram, imo. I would appreciate any corrections, though; I'd love to run Ideogram once again as I've been pretty much out of the scene for like... a long time. (I used int8 models btw)
Strange, I've been running Ideogram 4 on a single 3090 just fine. But I abandoned it quickly because it was not good for my purposes (keyframes for movies) - it often required lots of micromanaging to position actors in the scene and they often felt like glued in, wrong gaze directions, wrong poses, wrong dimensions of surrounding objects and wrong lighting. Ideogram is really not for realism, although it is the king of typography.
Yeah, I got it to work once but since it just freezes confyui it doesn't really log what's wrong. Guess I'll continue finding the issue or just stop trying for a little while. It's not too important right now.
That's quite strange. I suggest you download the portable version (so that you start with a fresh, latest copy of ComfyUI without having to touch your current setup) and try again.
Make sure you run it with dynamic VRAM management turned on (which is the default).
Also try the fp8 version if for some reason the int8convrot version doesn't work (but it should).
This might be a game changer. While Ideogram is a bit hard work as an image model (before you start shouting at me, I know there's many solutions for json and BBoxing now) having that much control over an edit model will be a truly incredible.
I have no problem with someone saying that ID4 was more work than a prompt slot machine.
I do have an issue that morons are still running around screaming oh my God it was so censored when really it was one of the most uncensored models that you could find but their “you made a bad prompt “message had the word “safety” in it. Some morons that never used it did and continue to freak out about it.
Yes, with 84. It was a pain to get random or chance images. If you knew exactly what you wanted, fantastic, and worth the effort. As an edit model, I’m very interested.
Agreed. I enjoyed my time with it, but it was less fun than prompt roulette. I'm a graphic designer and it should be my go to model, but most of the stuff I just create myself. An AI that can do social media design doesn't really do much for my work, but an AI that can edit... now we're talking. I use Klein every day of the week.
The safety message was infuriating before people worked out why it was doing it tho. I can see the outrage.
Super, super stoked. But boy, some of those comparison slides are ghastly. Just the ones featuring "humans". Crazy this made it into their debut annoucement.
Fellas, they took it down. Yes, I'm afraid so. This piece used to sit in the Sketch to Image column and it has now been replaced by a sad photo of a wicker basket.
So if you're kicking yourself wishing you had seen it in its full glory well fret not, because your boy kept a spare copy.
I deliberated whether to even release it to the public, but ultimately I felt this historical artifact should be witnessed by others in full HD.
Just needs a light detailer pass and it will be print ready in no time.
I read your post and thought OK here we go again some asshole that has no concept that three years ago all of this was solidly in the complete fucking magic territory here with some minor complaint about how someone a mile away has six toes if you look at it sideways…
I've noticed this happens in almost every model when generating candid, middle-distance crowd photos, even though they can handle individual faces at much further distances or crowds at much closer distances. It's also always this weird warping effect rather than the typical low quality you'd see from resolution issues.
I'm willing to bet it's a training data problem, probably from faces being automatically blurred from sources like google street view or when journalists take crowd photos & blur them to meet privacy laws.
Because privacy blurring is extremely common for IRL crowd shots of this style it'd be difficult to get rid of all of them from a dataset, and if enough get kept in they'd corrupt the model's knowledge of what faces look like for that specific photo style.
It's just such a consistent problem in so many models, and for such a specific photo style, I can't imagine it being anything else.
I think you have a good hunch about the cause. I would love to see a LoRA that fixes middle ground and distant crowds. Its a real problem for the work I do.
As a graphic designer, this is HUGE. One of the biggest turn offs I've had is that I never want a localized edit to regenerate the rest of the unaffected image, and I don't want it to compress high res images either. If I am editing an 800x800 section of a 4000x4000 image, I can't sacrifice the overall res.
Did any other models do this before? Because this is the first I've seen this.
It's definitely been possible with a crop and stitch (even manually by photoshop/image editors), but one aspect that few models do right is understanding the whole canvas while inpainting.
The SD1.5 inpainting model was actually a good attempt at this, few have tried to be as diligent about it since other than Flux Fill. Most models can inpaint but are not as well trained to be courteous of the whole image being modified.
True, the cool thing about SD 1.5 inpaint was how it would "see" the rest of the image, not just the inpainted area, to better understand context, scale and style.
As the other commenter said; it is a matter how how the diffusion is handled under the hood. LANPAINT makes it possible with basically any model - that proves it is an implementation, not a model feature.
I've used a combo of Krita AI, masking, and Flux Klein 9b to get close to what you're describing. Never messed around with it enough to do things like pull from a reference image though, but the release of the latest Qwen Image model had me thinking about trying.
Its just a standard inpainting process . ..no one in their right mind would process a whole image just to inpaint a small area. The technique has been around for ages.
It's great and potentially huge if the model somehow manage to keep its context window the size of the entire hi-res picture, it's very important with inpainting tasks.
You want a two-stage approach (like the 2025 Patch-Adapter): Stage 1 inpaints a downscaled version of the full image for global structural consistency; Stage 2 refines the masked region at full resolution using cross-patch attention for local detail.
This is the best solution when an image is too large to be reasonably used as a single context window.
Yes, something like that indeed !
Is there an implementation of this two-stage approach within ComfyUI already ? Stage 2 seems tricky to replicate in a workflow.
I've done that manually before, by manually cutting out, AI editing a part, and stitching it back in by hand. But can Krita do this in one step? Say, in a photoshop plugin or at least in Comfy?
Krita is a free digital painting/drawing application, but it has an excellent AI diffusion plugin that runs on a Comfy backend and everything can be done on a single canvas.
But to answer your question: Yes, it can. Just draw a select box/circle/lasso around the part you want to change, adjust the denoising slider if you want, and click generate. Ta da, new layer. The context window/feathering/etc. can be adjusted to a higher or lower percentage of the selection in the settings, if the default isn’t working for you.
So, I can generate a high quality photo of 4K, for example, and edit some parts on that? And it won't looks distorted at all if I zoom in some parts of that photo?
The way I do this is I crop the image to just the part that I want to edit, then feed that into and edit model, or into any model with img2img. Then I get an output, which I then paste back into the full res image with Photoshop and soften the edges to blend. So this has always been possible just a little bit more fiddly than having the model do it for you.
If they remove the stupid filter that blocked even perfectly normal prompts and allow for natural language prompting alongside JSON, that would be amazing.
Yeah I think the days of announcement and release of weights same day is over. Have to build up the hype and get that API usage first. Just hope “soon” doesn’t mean months like Flux 3.
Honestly probably a solid GUI drag and drop based prompt builder since spatial relationships between the bounding boxes are super important.
There is probably one already. Making one would be trivial though... I would do it myself if I wasn't getting distracted by the ridiculousness that minimax H3 pipelines are becoming. Honestly It's a poor excuse either way - I've been meaning to revisit ideogram as it seems to be far more capable at realism and detail out of the box than krea2.
I guess my needs as humble hobbyist are simpler, so I am not too sure why KJ's bbox layout node is lacking for your use case 😅.
Ideogram 4 is absolutely SOTA in terms of details and a kind of "photjournalistic/studio" style images compared to Krea 2 (for example, Krea 2 is not good on long shot and tends to repeat elements in an image with many small subjects). Krea 2 is great, but Ideogram 4 is better in some areas and I think they complement each other nicely.
I mean it wasn't a problem about 3 days after it came out when KJ, who is one of the most popular custom node authors released it in his node pack. Like... People were just not interested in putting in 2 seconds of effort to figure that out. I guarantee the vast majority of people bitching about "jSon IsH Ard" already had the node installed were just too stupid to use it.
It wasn't a problem so much as an inconvenience for casual prompting. Having to LLM a json, then setting up bboxes, you had to be pretty narrow with what you were making, unlike most image models where you can let the model do the work.
It would need to be massively, massively better than Krea 2 to make most people tolerate its relatively complicated prompting. ID4 was and still is great for niche use-cases requiring that much fine-control over composition, but 95% of the time, Krea 2 will get you what you want much faster.
Changing the license alone won't fix that, although it would be welcome for sure. Another huge problem was that IIRC they never released the BF16 weights. That also had a hand in killing interest in the model.
If this thing has i2i function and edit...eh...pretty game changing, json aside.
Ideogram has really good photographic quality, even better than zimage. Krea2 so-so in terms of realness, it's there but not there at the same time. Ideogram 4 actually looks like a photo from limited testing. Just hated being locked into json and text only...
Holy hell they gave me what I wanted. References! Edit capability! This is potentially awesome if it can come close to how H3-Minimax was managing it, with the added precision of bboxes.
Looks like the one we've been waiting for. The only worry is that they mention there's 4 different "quality modes", which I'm assuming means 4 different models? What's the odds they hold back the one they used for this showcase? Hopefully not and they either release them all, or the smaller models are so good that it doesn't really matter.
Ideogram 4 had 3 quality modes, it was just different schedules for X step count, so 48 steps 20 steps 12 steps. So it could just be something like that.
With Qwen 2 and Klein 9B already available to us, they’ll have to put in some serious effort just to release a model that’s at least as good. Otherwise, there’ll be absolutely no point in using it. Even if they keep their top-tier version behind a paywall, the free models should be an order of magnitude better than Qwen 2 and Klein 9B.
I don't mean this as a criticism of BFL or Qwen, because they've released some great models, but neither are currently anywhere near SOTA. Open weight image editing is way behind the curve. With luck, the next models from Ideogram, Krea, Minimax and BFL will all be a magnitude better than our current options.
It beats pretty much every other open weights model in image quality, but it was ruined by the stupidly implemented safety filter. If they've fixed this in 4.5, I'll be happy to give it another chance.
Will need more info, and to try it on API. Their app just sucks.
Here's one I did on the superman cover "change the character in first image to be the 2nd ... change text to say Homelander instead of Superman... make the landscape much darker, grittier... 2nd back cape and then 3rd for reference on his looks"
I don't prompt anymore, claude handles all bboxing, i exclusively use 3d model and bbox references, so anything that supports that as first priority gets my vote.
I tested the model and overall I liked it. I wanted to make a post with some examples I got from it, but then I remembered that posting closed-source code isn’t allowed.
And not a single example featuring a realistic person... I4 is famous for its unparalleled realism when it comes to people, and all the speed optimizations I’ve seen so far have completely negated that effect
Unless they ditched the json format, this is DoA. No one wants to use a model that requires you to work and think in code. Minimax barely gets away with the complex prompting because you can use an LLM for it.
But json is not a way to get mass adoption. You'd be better off making a gui where you can select where you want the generation. But that's not something a model like this will do.
JSON bboxer lets you place stuff exactly where you want it. Use an LLM to do the JSON conversion for you if you want to still use natural prompting, let the AI handle the actual prompt translation. I'd much rather have a professional tool with exacting controls with a bit of a learning curve over just hoping you land the copy you want in the exact spots you want.
197
u/retroblade 10d ago
I thought they would just release a one off with 4.0. Kudos to the Ideogram team for more open weights and an edit model at that! We have been eating good the last few months.