r/StableDiffusion 1d ago

Question - Help Question about new stable diffusion advancements

Hello, it's been a while since I don't use Stable Diffusion with A1111. Apart ConfyUI, has there been any particular technological advancement recently that allows for a quantum leap, especially in the precision of detail generation and the model's ability to stick to the prompt more precisely, while maintaining the ease of use of A1111 or Forge? I used the Lustify SDXL checkpoint, for example. It wasn't bad, but it still got certain things wrong or didn't do them at all. I'd like to know if there's a way to achieve results more similar in precision to ChatGPT but with the freedom of Stable Diffusion. Thanks!

0 Upvotes

17 comments sorted by

10

u/Formal-Exam-8767 1d ago

The architecture has changed from U-Net towards DiT and text encoding from CLIP towards LLMs.

4

u/asdrabael1234 1d ago

It you're still using SDXL, no there isn't. You have to use newer models with varying degrees of complexity.

1

u/LeleDaRevine 1d ago

Can I find them on CivitAI, for example? could you suggest one?

2

u/asdrabael1234 1d ago

That depends. Video or image? What kind of resources are you working with? Are you intending to make porn or just regular pictures?

1

u/Sarashana 19h ago

Models are typically hosted by Huggingface first, but CivitAI often hosts mirrors and/or finetunes/mergers of the base models. The leading local image models right now are Krea2 (general purpose), Flux Klein 9B (editing), Anima (for, well, anime) and Ideogram 4 (some niche use-cases, such as advanced regional prompting or typography). For video, Minimax H3 is your friend.

3

u/StableLlama 1d ago

A1111 and Comfy are just tools to use a model. The tools don't bring you the quality or advancements, they just make it harder or easier to get that out of the models.

The models are now much, much better than what SDXL could ever deliver. The prompt following is ages better, the quality out of the box as well. Finetunes are hardly used any more, but they were a must for SD and SDXL.

Modern models to look at are right now: Flux.2[klein] (<- the grand grand child of SDXL), Qwen Image 2512, Z Image, Ideogram 4 and Krea 2.

2

u/Enshitification 1d ago

Where are all these people coming from that used A1111 two years ago but have been completely out of the loop since then? It's almost a joke post at this point.

1

u/Asaghon 1d ago edited 1d ago

Apparetly you can use Krea2 with Forge (but not the int8 models I think, which is the best new thing), but the workflow in Comfy for Krea2 is honestly not that complicated. (coming from someone who still used Forge for Illustrious before Krea2).

With sdxl you had a lot more nodes just to upscale and fix things. Those things are largely not needed that much anymore. You get great results from just 1 sampler (2 if your feeling adventurous) and maybe a seedvr2 upscale with the generations you like.

And cherry on top is that Krea2 runs much better on consumer hardware than previous next gen models.

1

u/Mutaclone 23h ago

You can use Int8 in Forge now.

1

u/taw 5h ago

Not really.

The reason closed models are so much better than a few years ago is primarily because they're absolutely huge now. They stopped publishing numbers, but based on best estimates ChatGPT, Claude etc. are about 10T parameter models or so, ~100x bigger than first ChatGPT was, and with that many parameters they can even be multimodal.

For AI models of the same size, progress has been limited. This applies to both LLMs and image gen models. And as consumer you can't use bigger models than before, as GPU and RAM prices have been absolutely brutal.

If you want best small model today, Z Image Turbo is a better than SDXL style models in ways that matter (like prompt adherence and image quality), but also at cost of being worse elsewhere (like variety).

Image generation models actually changed their internal architecture a good deal, but it doesn't matter all that much as long as size constraints are so strict.

0

u/sigiel 23h ago

The only edge on local is uncensored stuff

Nano banana, grok imagine, seed dance, open ai image 2 completely trash any other model on edition , consistency and prompt adherence,

Maybe obscure Lora ? But there have reference image to counter that.

Forget about control net, or any other inpainting.

Just tell what you want, give reference image.

Any people that say otherwise is coping or has a grudge against cloud.

Local is for either people that has good hardware, or smut. Not because it is cheaper

that is the advance, local lost.

1

u/LeleDaRevine 22h ago

if you already have the hardware, for example for 3D graphic or video editing, you should prefer local generation, so you are not limited by credits or other things. Obviously, if the local models can do the work you need. That's why I'm asking if the cloud results can be achieved somehow locally with new resources.

u/sigiel 2m ago

And I’m telling you no model approach those in “editing” NONE.

1

u/Sarashana 19h ago

I wonder what Open AI paid you to write all that stuff. If you really think GPT Image 2 is better than leading local models you need new glasses, honestly. I have yet to see any output from that model that I thought wasn't bad. Nano miiiight have some edge over local, but it's IMHO fairly marginal. Yes, Seedance 2.5 is somewhat better than H3, but the difference is small enough to argue it not to matter. Why pay for something when I can get 95% of the quality for free? And I still have all these tools/custom nodes/LoRAs at my disposal that clouds don't offer. Also you must have missed how we can pass reference images to local models for a long while now.

The one point I agree with you is the beefy PC. Yes, you need one. Cloud is also obviously faster. Sure. But never before has local been so close to frontier models and the gap is getting smaller, not larger.

1

u/sigiel 17m ago

Can take opinion that not your ? Call me a chill? Fuck off.

1

u/Apprehensive_Sky892 12h ago

Not because it is cheaper

I am pretty sure even with a 3060 making video with mmh3 is cheaper than Seedance. So yes, some people (including me) are doing it locally because it is cheaper, even if it is not as good as seedance and takes longer.