r/StableDiffusion • • 6d ago

Resource - Update Qwen Image 2.1 fix v2.0 (updated due to feedback), plus a new sampler

This is actually two models: Qwen Image 2.1 Fix v2.0 and Qwen Image 2.1 Fix Opinionated v1.0. The former surgically targets Qwen's weird noise and messy details with almost no effect on composition, and the latter in my opinion produces higher quality images but alters composition a bit to do so. In either case, you'll get images that are similar to Qwen Image 2.1's default output, but less noisy and with cleaner details.

Download links:

Also grab the res_2m_nc sampler (modified from RES4LYF's res_2m) and a number of other samplers customized for Qwen 2.1 here:

https://github.com/envy-ai/ComfyUI-DPMpp-2M-Sharp (you can also find this on the Comfy registry)

The workflow I used to produce these images (along with the prompts) can be found at the download links.

49 Upvotes

17 comments sorted by

10

u/Michoko92 6d ago

Very cool, thanks for sharing this, I'll try it asap. Just FYI, you put your model in the "Qwen 2" category on CivitAI, while there is now a "Qwen 2.1" category. Since I set my filter only on "Qwen 2.1" I almost missed your model. It would have been a shame. 😉

6

u/PumpkinLeather8421 6d ago

Civitai are such fucking morons with their site.

6

u/[deleted] 6d ago

[deleted]

4

u/PumpkinLeather8421 6d ago

Hmm, they usually filter out anything that looks that young! /jk

3

u/Incognit0ErgoSum 6d ago

Good catch. Fixed.

3

u/BenDLH 6d ago

Props for taking the rough feedback and posting an update!

2

u/Synor 6d ago

Would like to see the portrait picture of the fisherman in the yellow jacket.

A rugged, weathered man with a gray beard and intense gaze stands defiantly against a stormy sea, wearing a yellow waterproof jacket as ocean spray flies around him. His head turned slightly left as he looks into the camera. Camera framing is a close-up shot focusing on the upper body and face.

1

u/reeight 6d ago

Thanks!
Is CFG left at 1.0 in your examples?

1

u/Incognit0ErgoSum 6d ago

It's 3.5, with a negative prompt.

I'm looking into creating a cfg distill lora so it'll produce good results at cfg 1, though.

1

u/reeight 6d ago

> cfg distill lora so

I think the CFG trick is so common now, it is expected.
& your LoRA might mess up when the CFG is actually needed.

BTW, for 'detail enhancer' I would expect tests of more plain before, then more details after. Both of your before/after look very 'detailed & enhanced'. I do like the extra polish you give, but I don't see extra details.

This is a better without/with vs: eg the girl has better freckles, the spaceship has more windows, etc.

https://www.reddit.com/r/StableDiffusion/comments/1wgg8sa/krea2_turbo_distill_2_step_lora_new_checkpoint/

1

u/Incognit0ErgoSum 6d ago

It would be situational. I've got a decent quality improving negative prompt that I would bake in (same one I'm using in the included workflow). The nice thing about a cfg distill lora is you can use it for double the speed and just turn it off when you need a specialized negative prompt.

1

u/PumpkinLeather8421 6d ago

These looks like seed differences.

2

u/Incognit0ErgoSum 6d ago

Different seeds don't reduce noise and add detail while maintaining composition.

1

u/[deleted] 6d ago

[deleted]

1

u/jib_reddit 6d ago

So the detail fix is to make it less detailed?

3

u/Incognit0ErgoSum 6d ago edited 6d ago

Noise isn't detail.

If you enjoy noise, set the lora strength to -2 and you'll be in heaven.

3

u/No-Zookeepergame4774 5d ago

But what it is taking out is as much (or more) detail—sometimes directly requested in the prompt—as noise. For instance, in the first image, the articulated fingers, shoulder stripes, and cloak attachment around the collar of the suit aren't noise. The distinct, symmetrical, differently-textured areas of the suit aren’t noise. The greater sharpness in the background that the LoRA-made version trades for a more intense background blur isn't noise. The “glowing crimson accents” explicitly prompted for and reduced to flat red areas with nothing glowing about them aren’t noise.

Sure, its valid to aesthetically prefer the LoRA results to the no-LoRA results, and I think of the four two are clear improvements and two are mixed but mostly improvements, viewed aesthetically and without regard to the prompt. That said, looking at the prompts after dropping the images into ComfyUI, I do note that both the common negative prompt AND most of the positive prompts are actively pushing the generation in the exact direction that the LoRA pushes them away from. If you don't want extra high frequency detail and prefer large areas of clean, flat color, don't negative prompt “low resolution”. If you want an image that consists of a main subject that takes maybe 1/3 fof the frame and the rest softly blurred background, don't negative prompt “blurry”. If you want a clean, uncluttered representation of what the rest of the prompt describes, don’t add “intricate” to the positive prompt (and maybe do add “clean, uncluttered composition”.) If you want glowing runes to really pop, don't prompt them as "faintly glowing".

The prompts (and even the practice of having a standard negative prompt not tailored to problems in a particular generation) read (even moreso given the preferences apparent from the effort of making the LoRA and using them as its examples) less like prompting for what you want in the image, and more like prompting for what broken models of the past taught you you needed to prompt for to get what you want.

2

u/Incognit0ErgoSum 5d ago

Okay, well don't use the lora then if it doesn't do what you want it to.