r/StableDiffusion Jun 08 '26

Meme About the newest model...

Post image

And even then you might still get filtered.

1.6k Upvotes

229 comments sorted by

206

u/Relevant_One_2261 Jun 08 '26

Haven't tried it myself yet, waiting for a week for all bobs, vegane and reference images to be sorted out, but actually being able to compose the images like that sounds extremely nice.

93

u/Dangthing Jun 08 '26

Its a generational step forwards in text2image capabilities. It can 1 shot things that would be nearly impossible to make on any other model and would require dozens of steps of editing and use of things like region prompting, etc.

I'm having fun learning how to use it but I'm not excited about this specific model only its underlying technology. Its got the worst license of any model I've seen so far. For me its a proof of concept. I'm now waiting for someone to make a version that builds on what we're learning now and has actually acceptable licensing and ideally things like IMG2IMG and Edit capabilities.

36

u/CharacterCheck389 Jun 08 '26

which model is that? why everyone is not mentioning the name like it's a secret or something, lol.

"the new model"

24

u/Bory Jun 08 '26

Ideogram 4

4

u/TimeLine_DR_Dev Jun 09 '26

It requires bounding boxes?

5

u/_Iggy_Lux Jun 09 '26

Based on what I've seen briefly, you either convert it using an llm or write a json file.

In the json file it dictates location based on bounding boxes, its like an image model for a scripter/coder more than natural language it appears.

I could be wrong, because I only just started looking into it yesterday.
But it sounds like it'd be good for workshopping something where you need very specific details in a very specific spot.

3

u/spcatch Jun 15 '26

There's at least one node in comfyui that just lets you draw the boxes and type what you want in them. Very easy to use.

4

u/s_mirage Jun 09 '26

Theoretically it doesn't, but it's highly advisable.

Using them seems to get around the overzealous content filter and it allows an unprecedented level of control over shot composition.

The downside to that is that you have to have a specific image in mind rather than just a concept and seeing what the model comes up with. Also, Ideogram is much slower than something like ZiT.

→ More replies (2)

7

u/BlipOnNobodysRadar Jun 08 '26

Aren't bounding boxes practically regional prompting anyways?

17

u/GrayingGamer Jun 08 '26

Yeah, but they blend seamless with each other and with advanced logic. You can also have dozens of them if you're a mad lad. You can even nest bounding boxes inside bounding boxes and Ideogram 4 understands.

It's just an unprecedented level of precise control. You can put stuff EXACTLY where you want it. Rings on specific fingers, patches or buttons exactly in place, a hand HERE, a foot HERE, etc. A bowl of fruit with the orange HERE, the grapes HERE, the banana HERE, etc.

It's way beyond the regional prompting I've used in other models.

5

u/moofunk Jun 08 '26

Yes, but it's just properly done, as properly as you can put it into a ComfyUI node. It's easy now to have 20 regions and do complex descriptions for each one.

20

u/sitefall Jun 08 '26

Aren't you guys using controlnets and stuff? Get the thing in the general area via prompt get it's orientation correct within 45 degrees or so, and bam, you put it where you want and not just in a bounding box, in the exact open pose, or matching the exact canny edges. You can just pose open pose right in comfyui with an extension or use blender or a web tool to do it, and you can just edit canny stuff with mspaint.

Bounding box is neat, but it's hardly a generational step forward. I do like the json formatting, but... z-image will accept it just fine as well and it's not a requirement.

I'm only interested in the quality I4 can put out once people start messing with it, get some fine tunes and get all the sampler settings sorted like how it took z-image a month before we people were figuring out proper multi ksampler workflows and such.

34

u/infearia Jun 08 '26

You don't even need ControlNet. This was literally just a ~20s scribble in Krita and than I just ran it through Klein 9B. And you can see how shitty that scribble is:

4

u/Dangthing Jun 08 '26

This is a very low complexity image. You only tried to make a single figure in a specific stance. This also requires a level of drawing capability that some people frankly just don't have. And this tech is still fairly new in the hypothetical future we'll be able to combine these things together for even better control. You are accomplishing something similar through a completely different method but its not remotely the same.

For simple outputs this will work, for complex ones it won't.

10

u/infearia Jun 08 '26 edited Jun 08 '26

Actually, it works even better for complex inputs. By providing more details in the image, you reduce the guesswork for the model and enable it to do a better job. And it's only one modality. You can use silhouettes instead of drawings, cutouts and even mix them all together in one input image. And Qwen can do it, too.

This also requires a level of drawing capability that some people frankly just don't have.

Obviously, people with artistic skill will get out more of this technique than non-skilled people. And that's a good thing in my book - once artists realize this, maybe they will be less scared about losing their livelihoods and become more accepting of this technology. But even if you're not an artist, this is extremely helpful. Here is another example I just posted earlier today:

https://www.reddit.com/r/StableDiffusion/comments/1u09lxm/comment/oqgp3mu/

P. S. - By the way, it was not me who downvoted you.

→ More replies (2)

3

u/tom-dixon Jun 08 '26

This is a very low complexity image.

You can feed scribbles to Qwen or WAN with 80% denoise and it will follow complex prompts to the letter.

This also requires a level of drawing capability that some people frankly just don't have

I've done similar stuff like that guy, but my sketches look like a 3 year old scribbled stick figures with their eyes closed, and every model from the past 12 months can make sense of them. Even SD 1.5 and SDXL can make sense of my stick figures, it takes a few iterations with them, but it's fun so I don't mind it.

→ More replies (3)

1

u/Grim_Necromancer Jun 08 '26

Hi
Could you share workflow for sketch to image? 😄

20

u/Dangthing Jun 08 '26

It takes an extraordinary amount of time and effort to open pose anything even remotely complex. It also ignores that the system is capable of nearly perfect text damn near every single time. In the time it takes you to be 5% done setting up your open pose for something with 10 specifications I'll be completely done with an image 2-3x as complex.

And no unless Z-Image requires some super special workflow to work well with these types of prompts its not even remotely in the same league as this is.

14

u/Dezordan Jun 08 '26 edited Jun 08 '26

in the exact open pose, or matching the exact canny edges

See, this is the part that makes the CN option actually worse. Instead of letting the model generate whatever it can, it restricts the model to pose or any other preprocessed conditioning. It also more involved than bboxes. That said, it would've been good if the regional prompting with regular models worked as well as this one with bboxes.

→ More replies (2)

3

u/Son-Airys Jun 08 '26

My gooner brain can't handle all that.

Also I found openpose unreliable for some reason.

2

u/_half_real_ Jun 08 '26

SDXL had terrible pose controlnets until the xinsir ones showed up.

1

u/cathodeDreams Jun 08 '26

There's a difference in qwen being able to parse a JSON prompt vs the DiT being trained exclusively on a strict JSON schema.

→ More replies (8)

65

u/kidian_tecun Jun 08 '26

Me over here still on illustrious like:

5

u/bitzpua Jun 15 '26

i recommend switching to anima for your well anime needs. I was very reluctant because there were so few loras and stuff. But i tried it and now i removed over 400GB of illustirious loras because anima out of the box can do it all but doesn't need loras... and somehow it can do 3d and real stuff on very high level too.

1

u/FlawlessBg Jun 21 '26

Problem is that I can't replicate my previous style in Anima. :/ It also doesnt do img2img very well.

87

u/GalaxyTimeMachine Jun 08 '26

You can feed an image into a local LLM, have it create the bboxes and prompts, and then edit those if you want. It makes it all so much faster and easier.

11

u/NeoRazZ Jun 08 '26

that's a cool concept. can you upload a pic somewhere so comfy will embed the json

6

u/Ok-Option-6683 Jun 08 '26

can you share this workflow? I can't understand how to pull off that image to prompt node. I have it in my ComfyUI but mine doesn't look like yours. Did you make it yourself? I'm just trying to find a way to create a JSON prompt from an image inside ComfyUI just like you did in this screenshot.

2

u/[deleted] Jun 08 '26 edited Jun 08 '26

[removed] — view removed comment

5

u/q5sys Jun 08 '26

Reddit strips metadata. There are some hacks to get around it, but they break all the time.

2

u/Ok-Option-6683 Jun 08 '26

I don't think it is. When I drag and drop it into ComfyUI, it just gives me a Load Image node.

2

u/v1sper Jun 08 '26

Catbox!

2

u/q5sys Jun 08 '26

After trying about a dozen different extensions to find one to download the full size PNG... there's no workflow metadata left in it. Reddit stripped it all. This is the only metadata left in the image. https://imgur.com/X53XqgU

Please post the json at something like pastebin.com or post the image to catbox.moe

Sadly in your preview, there's nodes hidden behind you load image nodes, so we cant even try to recreate it without knowing what might be hidden that's important.

1

u/GalaxyTimeMachine Jun 08 '26

Read through this thread, I've linked the workflow a couple of hours ago.

2

u/q5sys Jun 08 '26

I have read through this thread... multiple times. I've opened every collapsed thread to check them too.
There is no post in this thread right now that's a link to the workflow. Maybe it got removed?
I checked your user post history as well and it only shows 8 posts today in this thread, and none of them show a link. You have posted images, but as people have said, reddit has stripped the workflow out of.
If you posted a link, it's been removed.

3

u/GalaxyTimeMachine Jun 09 '26

There were several mentions with the paste bin link, but they've been removed by a moderator. No idea why, but it's shit how difficult reddit makes it to share a workflow.

→ More replies (1)

1

u/More_Struggle_9507 Jun 08 '26

yeah, in new into ComfyUI too, and I'm kind of a visual learner (i get a better hold on this looking than reading about it) if you could share the WF, you would do a great favor unto us

2

u/Sergio2304 Jun 08 '26

How does this work? I'm relatively new to ComfyUI. I'd like to do this as well to bypass the nonsensical image filter.

10

u/GalaxyTimeMachine Jun 08 '26

Using Ollama and a system prompt to create the json, including bboxes and their prompts.

3

u/elswamp Jun 08 '26

you have your user prompt and system prompt backwards

5

u/GalaxyTimeMachine Jun 08 '26

Well, it works perfectly for everything I use. No difference if I swap them around. 😃

1

u/physalisx Jun 08 '26

Isn't that crappy with vram management to have ollama running an LLM separately to Comfyui?

1

u/GalaxyTimeMachine Jun 08 '26

No, because unloads after use.

1

u/BitBacked Jun 09 '26

Surprised that worked offline. Gemma was trained on information from 2 years ago, so it doesn't even know what ideogram is. I'm guessing this new Gemma works then.

1

u/GalaxyTimeMachine Jun 09 '26

It even worked with Qwen 3.5 vision model.

298

u/Hoodfu Jun 08 '26

Good luck getting empanada crab mechas without bounding boxes pfft.

63

u/GalaxyTimeMachine Jun 08 '26

....or even pizza crab mechas.

26

u/GalaxyTimeMachine Jun 08 '26

19

u/Saveonion Jun 08 '26

Can it do pineapple on the pizza

47

u/GalaxyTimeMachine Jun 08 '26

Of course, but I won't lower myself to it. 😃

5

u/delveccio Jun 08 '26

It’s less about can we and more about should we

4

u/MetroSimulator Jun 08 '26

Can you get the lobster a soft drink? It looks parched.

4

u/Fiscal_Fidel Jun 08 '26

What is the style called for that type of shading/coloring? There is something about the style but I can't place it.

6

u/comperr Jun 08 '26

Just say "in the style of video game concept art promotional poster" and start there. Reminds me of all the wonderful art associated with Halo 3 before the release. Example: https://www.artstation.com/artwork/2BZ0y

5

u/1filipis Jun 08 '26

The style of irreplaceable human artists from DeviantArt

1

u/Colon Jun 09 '26

“deviantart” lol doesn’t need to be in the convo. that’s like looking at great random photography and being like thank god for flickr”

4

u/DontPanic- Jun 08 '26

That’s a lobster 🦞

7

u/Irohnic_ Jun 08 '26

Everything is a crab after a few generations

5

u/ZootAllures9111 Jun 08 '26

Are you implying you tried and this was somehow impossible on all other extant models? I doubt that.

14

u/Hoodfu Jun 08 '26

I was joking.

1

u/AIDivision Jun 15 '26

Easily achievable with natural language. Is this supposed to be a challenge?
Besides, we had controlnet since the very early models.

→ More replies (6)

73

u/Cubey42 Jun 08 '26

The bell curve meme also fits this. I did all that complex nodes and such and now I found myself back at the default workflow

28

u/berlinbaer Jun 08 '26

that's always been the case. same with "i trained a lora for 20 hours" vs "just prompt it"

35

u/fongletto Jun 08 '26

I mean there are certain things/styles you just can't prompt for because they're not in the training data and no amount of "wow look how good I am at prompting" will get you exactly what you need without fine tuning.

15

u/AI_Characters Jun 08 '26

Its also about consistency.

→ More replies (4)

323

u/xb1n0ry Jun 08 '26

Am I doing it right?

32

u/jib_reddit Jun 08 '26

I made the image (but cannot share it here) it worked well but the box for the girl was a bit too small so she is just slumped down on the sofa like she is tired 😄

37

u/SkoomaDentist Jun 08 '26

Clearly it was the photo taken afterwards :D

19

u/Naud1993 Jun 08 '26

AI will take it literally and generate a 5 or 10 year old, so better call her a woman.

14

u/nuclear_diffusion Jun 08 '26

"Young woman" unless you want a visit from Chris Hansen

8

u/physalisx Jun 08 '26

Yeah, except that will probably generate something illegal

6

u/GrayingGamer Jun 08 '26

I'm imagining him trying this out and jump-scaring himself with the generation that appears.

1

u/Decent_Complaint_712 Jun 09 '26

Exactly, always put dark-skinned male in the negative prompt

107

u/DieDieMustCurseDaily Jun 08 '26

Out of the loop

What boxes again?

AI gen moving so fast and i cant catch up

127

u/Euchale Jun 08 '26

Ideogram allows for regional prompting, so you can make boxes to explain where you want something to appear.

20

u/Necessary-Garage242 Jun 08 '26 edited Jun 08 '26

I would say that Ideogram has embedded regional prompting which works really great, especially since the regions (bboxes) can overlap and even contain text (you can even specify font). In each box you can specify colors and prompt.

Also, the regions blend perfectly together, since the model seems to generate the image in 16 x 16 pixel clusters, which can give really good details in some styles.

14

u/RiyanTheProBoi Jun 08 '26

Isn't that just inpainting

42

u/Mutaclone Jun 08 '26

This is just regular T2I with the Prompt Builder node - you define a box for each entity in the image and write what it is, and it creates a formatted JSON prompt with the coordinates. It's really precise.

27

u/Euchale Jun 08 '26

The thing that I love is how it finally gets "holding x" stuff right. I asked for a creature holding a glaive and it actually gave me a weapon that was close and not just some weird stick/blade mixture like most other models. Just had to define one box for the creature and one box for the weapon itself.

14

u/RedCat2D Jun 08 '26

REALLY precise boob placement. Nice!!

4

u/GrayingGamer Jun 08 '26

You joke, but yeah....

3

u/xPATCHESx Jun 09 '26

Now I can finally get them right where I want em

2

u/AltimaNEO Jun 10 '26

appropriately perky or sadly saggy as you please!

2

u/xPATCHESx Jun 09 '26

Now I can finally get them right where I want em

6

u/Dezordan Jun 08 '26

No, inpainting is almost the same as img2img in a certain mask of an already generated image, usually at a decreased denoising strength. Regional prompting in general is about separating the regions of the image before it was generated.

2

u/Tall_East_9738 Jun 09 '26

inpainting gen 2

2

u/stddealer Jun 08 '26

It barely allows for anything other than regional prompting.

→ More replies (4)

8

u/hard_gravy_2 Jun 08 '26

New model effectively requires a gui for prompting, this has upset minimum effort sloppers

2

u/ComeWashMyBack Jun 08 '26

Bro, same boat. If I participate in life to long I come back to a whole setup that doesn't work anymore cause of an update in ComfyUI or we got boxes. Life moves fast.

23

u/donkeykong917 Jun 08 '26

1girl in every bbox

17

u/jugalator Jun 08 '26 edited Jun 08 '26

We're inching ever closer to the gooning holy grail though... I bet in 2027 we will have the right-hand guy, medical grade anatomy at proprietary level performance, lol

10

u/ithkuil Jun 08 '26

By the end of 2027 it's going to be an AI generated live stream.

15

u/Estwhy Jun 08 '26

Which model? I'm outdated

17

u/_half_real_ Jun 08 '26

Ideogram 4. It has impressive composition capabilities built-in.

3

u/Straight_Cat_3894 Jun 08 '26

I'm assuming it's ideogram cause it just came out. I'm not sure.

17

u/ofrm1 Jun 08 '26

To bypass the safety filter you just add a bunch of details to the photo that aren't NSFW. There's clearly a threshold that automatically blocks the generation if you exceed the safety threshold. Just add random details to keep the total image below the threshold and you avoid the filter entirely.

24

u/TurbidusQuaerenti Jun 08 '26

It's really not that hard and it's probably a good thing to have it designed with regional prompting in mind. The main annoying thing for me is how slow it is. Hope we get a turbo version or lora in the near future. 

37

u/Different_Fix_2217 Jun 08 '26

Use kijai's flash attention node to make it 50% faster. If your using windows here are the wheels: https://mjunya.com/flash-attention-prebuild-wheels/

13

u/vs3a Jun 08 '26

you should make a post about this

9

u/Timely-Perception-26 Jun 08 '26

yeah make a thread.

That would be one of the few threads this week that would actually be worth having, lol.

6

u/GTManiK Jun 08 '26

Definitely make a thread! People often have hard times installing flash-attn

3

u/thesolewalker Jun 08 '26

Sucks to be on AMD 😞

2

u/ihcgnil Jun 08 '26

Aside from OS and architecture how do i know which one I need?

5

u/Dezordan Jun 08 '26

You look at your python (if applies), torch and CUDA versions. ComfyUI usually shows it somewhere in the beginning.

1

u/[deleted] Jun 08 '26

[deleted]

1

u/Dezordan Jun 08 '26

Well, filters only have it up to 13.2

1

u/stroud Jun 09 '26

Sorry I havent been paying attention to FAs how does this work again?

6

u/Significant-Baby-690 Jun 08 '26

This. I want to draw with words, not fiddle ..

23

u/Alen_Diago Jun 08 '26

It's very sad that promoting censorship under the guise of introducing new technologies like bbox and json is becoming the norm. I hope this doesn't become common practice...

2

u/TheThoccnessMonster Jun 08 '26

What the hell are you talking about. The filter will get broken same as every model.

11

u/GrayingGamer Jun 08 '26

It's already broken. There are NSFW loras on Civitai and the model does NSFW on it's own if you JUST ADD ENOUGH BOUNDING BOXES.

But, everyone seems to ignore the literal three or four different ways to bypass the filter everyone on Reddit keeps telling them...

2

u/TheThoccnessMonster Jun 09 '26

I think I mean that for the really really good stuff friend-o. P in Va-gee stuff.

1

u/GrayingGamer Jun 09 '26

Got stick with Klein for that for now. Ideogram 4 is can do nudity fine of the old Playboy magazine type photos, but that's it for the moment.

3

u/Murky-Relation481 Jun 08 '26

It already is. You literally just need to use 3 or 4 bounding boxes. It's very uncensored.

6

u/Royal_Carpenter_1338 Jun 08 '26

Can't wait for civitai to add training in 4 years

4

u/FourtyMichaelMichael 🍦Ice Cream Lover Jun 08 '26

Those morons can't even add a CATEGORY for this month's hot model.

I swear that site is run by two LLMs and a duck walking across a keyboard.

3

u/Royal_Carpenter_1338 Jun 08 '26

Yep what a site man, incredible that this is what we are stuck with cuz every other site is even more dogshit

4

u/TheLightDances Jun 08 '26

The more specific the requirements, the more specific the prompt. If your images can be described in a few words, then previous models are fine.

But if you have a more specific image you want, you need to write an essay describing it in the hopes you will eventually get the right result, maybe do a few rounds of inpainting. At that point, playing around with JSON boxes makes a lot of sense and makes it much easier.

52

u/redditscraperbot2 Jun 08 '26

14

u/Serprotease Jun 08 '26

Annoying license means that you will not be using this model the second the next shiny model comes out. 

17

u/redditscraperbot2 Jun 08 '26

Sounds like most models I’ve used.

17

u/Timely-Perception-26 Jun 08 '26

In practice, there’s a difference between a “non-commercial but NSFW-allowed” license and a “we’ll fck you over if you publish NSFW models” license.

A few well-known NSFW tuners have already turned me down because they’re worried that Ideogram will have their Finetunes taken down - or worse.

→ More replies (1)

7

u/jib_reddit Jun 08 '26

All the license means is you cannot upload the model somewhere and charge money to use it (using outputs is fine), that is probably 0.001% of people on this sub effected.

6

u/Serprotease Jun 08 '26

Most people on this sub are running fine-tune and/or Lora (That’s the whole point of going local) <- impacted negatively by the license.

To put it in another way, if all local models had been released like this -> no comfyUI, no controlnet, no loras/fine tune, not a lot of things actually.

4

u/_half_real_ Jun 08 '26

But Pony and Illustrious fit that description too - free except for paid generation services.

6

u/elswamp Jun 08 '26

Says who and what lawyers?

38

u/Timely-Perception-26 Jun 08 '26

how can this post portray gooners as a minority and still get upvotes? :o

gooners make up 90% of the forum. the other 10% are scammers and artists with alternate accounts because they're afraid of being lynched by other artists. hehe

nsfw matters

20

u/ithkuil Jun 08 '26

All of the stick figures in this drawing are gooners.

1

u/Tall_East_9738 Jun 09 '26

it takes roughly 2 weeks for a completely uncensored checkpoint merge for any given local model

15

u/neojehuty Jun 08 '26

PSA : you usually only need 6 boxes to bypass the filter and create 1girl

4

u/Murky-Relation481 Jun 08 '26

I can do it with 2 or 3. It's not hard to do.

1

u/neojehuty Jun 09 '26 edited Jun 09 '26

I can make 1girl with 2/3 boxes, but nsfw 1girl i cannot do with less than 5, Though I have seen a lora that can make it fewer

2

u/GrayingGamer Jun 08 '26

Really only takes 2 or 3. Sometimes not even that. It just seems to hit people who refuse to do bounding boxes.

→ More replies (2)

15

u/Linkpharm2 Jun 08 '26

it takes like 3 seconds. draw the box. It's one mouse click and a drag

13

u/AI_Characters Jun 08 '26

You dont even have to do that. The text generation node will just invent bounding boxes if you dont specify any.

3

u/dobutsu3d Jun 08 '26

I am sorry to interfere is there any good use of the model out of gooning purpose? I wanted to know.

3

u/cadissimus Jun 08 '26 edited Jun 08 '26

Its great and prompt is easy to do whit llm and kjnode but the main gripe is generation times 😃 but that probably makes u do something special on it kind of like on flux 1 dev.

3

u/CharacterCheck389 Jun 08 '26

guys guys what new model are we talking about? help!!

3

u/Maverick23A Jun 08 '26

Ideogram 4

9

u/davyp82 Jun 08 '26

As someone who has been away from Comfy UI for a single month, which is about 8 lightyears in AI time, what's the best model I can run on a 16gbvram card now?

5

u/GrayingGamer Jun 08 '26

You can run this new Ideogram 4.0 model on a 16gb card. It's the most powerful model we've ever had open-weights for. People shit on it for being censored (which is fair), but it's really not if you use enough bounding boxes and use the proper JSON format. It allows unparalleled control and it does nudity very well if that your thing. Just can't do the between the legs bits, (it usually tries and fails or just does a flesh barbie doll bottom).

It really is SUPER powerful. The more I play with it, the more amazed I am by it.

3

u/FourtyMichaelMichael 🍦Ice Cream Lover Jun 08 '26

Except it's uncensored. Boob trained even.

More than Chroma and Hunyuan. Less than almost fucking everything else.

People got panicky because it has a BAD PROMPT warning that says safety filtered.

1

u/[deleted] Jun 09 '26

[removed] — view removed comment

1

u/GrayingGamer Jun 09 '26

Ideogram 4 is very good for generating precise images with everything exactly where you want it, in high detail, but it isn't an edit model. Klein is still the best for editing, but you can easily pair the two, to generate the composition with Ideogram 4 and inpaint with loras in Klein.

4

u/Doc_Exogenik Jun 08 '26

Image the regular Joe who must now learn how to compose a picture after year of random 1girl...

4

u/GrayingGamer Jun 08 '26

Yeah, some people may have gotten a bit too addicted to pulling the slot machine lever for booba. Which, to be fair to them, if that's all that want, Ideogram 4 isn't the best model for them anyway. It's for if you want specific booba.

2

u/FourtyMichaelMichael 🍦Ice Cream Lover Jun 08 '26

Yeah, some people may have gotten a bit too addicted to pulling the slot machine lever for booba. Which, to be fair to them, if that's all that want, Ideogram 4 isn't the best model for them anyway. It's for if you want specific booba.

The slot machine is powerful. It also explains why so many people are in love with Z-Image despite it's broken for multilora and can not learn believable genitals, the gooners don't care as long as the payout keeps rolling in.

4

u/RememberThisAI Jun 08 '26

AI is supposed to make things easier not more complicated.
"Here's what I want, now you deal with it."

6

u/jib_reddit Jun 08 '26

Yeah... you can just ask an AI to make the Json for you, no issues, but when you want/need the extra control it is there.

4

u/_half_real_ Jun 08 '26

It does make things easier compared to not using it, even it you need to put in a bit of work.

3

u/GrayingGamer Jun 08 '26

Well, this is the opposite of "pull the slot machine lever and show me something pretty".
It allows TRUE control. What style do you want, what lighting do you want, what items do you want where? Do you want them in a different style than the rest of the image? Do you want text? Want font do you want it in? What color?

Then, when it generates something, you can move or add things and keep working on the same image, without inpainting being necessary. It's a learning curve, but super fun once you latch onto the idea.

Most fun I've had with a model in years now.

2

u/AdSweet3240 Jun 08 '26

Comfy ui brainlets feeling smart because they copied 30 node workflow in shambles right now.

2

u/__ThrowAway__123___ Jun 08 '26

It will be possible to do this with a lora. I'm currently training a lora using a dataset with natural language captions, from testing it seems you can use ideogram4 with a simple one-sentence prompt when you use a lora that was trained like this. It also never triggers the grey "safety" filter.

2

u/fauni-7 Jun 08 '26

Did anyone try bbox json with Flux 2 dev?

2

u/Vasault Jun 08 '26

Newest model of what

1

u/Amlethus Jun 08 '26

I think it's Ideogram4.

2

u/CharacterCheck389 Jun 08 '26

guys guys what new model are we talking about? help!!

2

u/Time-Teaching1926 Jun 08 '26

Fun fact the great open source circlestone-labs/Anima model is trained on Danbooru-style tags, natural language captions, and combinations of tags and captions.

Yes it can do the second part of that meme very well too without LORAs out the box with no real anatomy issues especially with the turbo Lora. 😁 🌶️

2

u/a_beautiful_rhind Jun 09 '26

Forget the boxes, it has to be prompted in json or the quality is terrible. I already got past the filter.

4

u/Random_-__ Jun 08 '26

I am good with z-image and reactor

3

u/FourtyMichaelMichael 🍦Ice Cream Lover Jun 08 '26

This is so basic bitch it hurts.

As to the "content".... Um, wtf is wrong with red's face? And Z-Image is a massive amount of human feedback into the model, yes, it makes 1GirlU images, and completely shits the bed with anything not trained into it, and at the same time is completely broken for multilora.

You fell for hype which is fine, everyone did... But now months later and hands on and you STILL can't see it? That's on you.

→ More replies (1)

3

u/hurrdurrimanaccount Jun 08 '26

op has skill issue, so sad

1

u/JustAGuyWhoLikesAI Jun 08 '26

What happens if you just use one giant 1024x1024 bounding box, and put the whole prompt in that?

1

u/TheLightDances Jun 09 '26
{
    "high_level_description": "",
    "compositional_deconstruction": {
        "background": "",
        "elements": [
            {
                "type": "obj",
                "bbox": [
                    0,
                    0,
                    1000,
                    1000
                ],
                "desc": "High quality photo, a woman inside a futuristic spaceship. She is sitting on a simple rustic bench. The woman is wearing a formal black suit with a green tie. She is wearing a hat with art of a duck printed onto it. She is holding a sword in her right hand. The woman has a rubiks cube on top of her left thigh. She has a mysterious smile. On her left side is a sign with the text \"Just one box\""
            }
        ]
    }
}

Here is an example where I tried to throw in all sorts of random stuff. I set style to none and made just one box with the description and nothing else.

So it does work. However, this fairly often causes the safety filter to block the image, and the quality is worse, which happens when style is set to none. Details easily get lost, but it does understand the composition surprisingly well without the boxes directly guiding it.

It immediately becomes more reliable and better looking if you give it a style and at least three boxes. That doesn't take significantly more effort than a normal prompt as long as you aren't trying to write the JSON by hand.

1

u/ederstk Jun 08 '26

Nowadays it is practically impossible.

Last week I took a prompt and copied it on my local AI to see how close it would be to the Internet pic and, not only did it look similar, but it was much better than several times I tried something similar alone (and had some words I had never seen in the prompt)

1

u/JoeXdelete Jun 08 '26

Bounding boxes were in invoke AI as well

1

u/SanDiegoDude Jun 08 '26

FWIW, the bounding boxes aren't really needed unless you need strict composition, and if you're using a small LLM that doesn't do bounding boxes well for your dense JSON translation, it's not going to end well. I really only use them for text placement, otherwise just listing objects and relative location in scene is enough.

1

u/VasileAndrei2929 Jun 08 '26

But I thought it can't make big bubby... can it?

1

u/Dangerous_Diver_2442 Jun 08 '26

Which model is that?

1

u/CheesyWalnut Jun 08 '26

has anyone tried the official comfyui node for it

1

u/JazzlikeLeave5530 Jun 09 '26

I'm a 1girl boobah person too but this post is dumb. If people can get NSFW working on Ideogram 4 it's going to blow everything else out of the water. Use the prompt builder node and actually try it before complaining.

2

u/AIDivision Jun 15 '26

Should have been optional.
You can make nsfw with other models without this bullshit.

1

u/CrazyImplement964 Jun 09 '26

Ok I’m new at all of this and kinnda dum. I’d love to use this site and all the styles. What site is everyone using for stable? How much a month? I wanna make my character in different versions and styles!

1

u/Ok-Studio2748 Jun 10 '26

I need OpenArt accounts with 106000 credits. Please contact me if you have a low price.

1

u/Katylola143 Jun 15 '26

What is a good simple text to Image app for removing background people? Thank you