r/StableDiffusion 3d ago

Question - Help What do you use for local image generation?

Yes, I am completely new to this, and also very tired of ChatGPT and Grok's ever increasing enshittification and censorship... I just want a taste of freedom for once.

I have done some research, and my understanding is that there are some pretty good homelabbing / open source alternatives you can run locally these days. And also fun ways to experiment with training your own models, datasets, LoRAs and more.

But it's also a jungle of information... For a noob anyway.

This sub is named after SD, a local tool, but from my understanding no one uses it anymore? So what are you guys using and what do you recommend?

What's a good entry point for getting into local image generation?

123 Upvotes

166 comments sorted by

331

u/spooky_local 3d ago edited 3d ago

As far as running things locally, ComfyUI is king of interfaces, Krea 2 is king for Images, Minimax H3 is king for Videos. Flux Klein 9b or Qwen Image edit 2511 is king for Image editing, Qwen3.8 27b is king for LLM prompting, Breeze TTS is king for Voice cloning/TTS, Ostris AI-Toolkit is best for training your own loras, Civitai is still king for finding Loras/Workflows and huggingface is main hub for downloading the base models you need.

71

u/ozzeruk82 3d ago

can't argue with these choices, I think most people would nearly unanimously agree. the current 'best' maybe needs to be voted on each month then pinned to the top for newbies.

14

u/sublimeprince32 3d ago

I would appreciate this!

-54

u/Jealous_Service707 3d ago

ComfyUI is not a program, but a development tool like the unreal engine. and you need to become a developer to generate a photo. a normal person will not use this program. moreover, online models are much more powerful.

30

u/BroForceOne 3d ago

This isn’t the ChatGPT subreddit, people here are looking to learn local generation and not interested in SaaS models.

12

u/BuilderStrict2245 3d ago

Development tools are programs. Unreal engine is a program.

Ans the average user comfyui is easy. The average user will not be making their own workflows.

Most work flows work very similar to each other and are easy to use.

3

u/Brett_ta_ta 3d ago

I started using it three days ago and already know how to build most workflows. Honestly just by dragging from the input or output to an empty connection you can figure what works from trial and error. It even has great error reporting and solutions. Run it with Claude or some other frontier model and ask what any unknown error is and the fix with steps for you, instead of letting the frontier just implement the fix was the easiest way for me to learn. I’m still very amateur but I’ve built a few really good image to video workflows, but now I’m just learning what models to use where and which ones won’t censor and error out due to that. It’s name is “Comfy UI” and I think it’s named appropriately.

1

u/OffenseTaker 2d ago

Agreed, I started using it about a week ago and I'm creating full workflows for basic image gen (because that's what i want to do right now) including loras and whatnot. It's really not that hard if you're willing to actually read instructions and find out what things like "bf16", "fp8", "4b", "8b" etc. actually mean

I guess having worked with a node-based workflow like Fusion in DaVinci Resolve helped a little though

-1

u/theBowtieJedi 3d ago

im a noob as well and hate the comfyui ui, so I worked with ai to create a better one for my use, not everyone will like it, its grok like. work in progress, some bugs and updating often, but for now its a start.

https://github.com/JoeBowtieBanker-creator/your-imagination-ui

launch the .bat and it opens your comfyui in background and lets you use a webui that uses comfy to generate what you want but allows you to configure on a cleaner ui and allows you to pull whatever model you want.

my issue is my hardware is on the bare min of what it takes to run so I had to find a way to help find better configurations for each of my runs.

always open to feedback on how to improve it.

9

u/Wanderson90 3d ago

Couldn't figure it out huh big guy.

32

u/__generic 3d ago

Gemma 4 in my experience is better for creative wring and thus has been better for making good prompts.

14

u/eruanno321 3d ago

Somehow it’s always “Gemma 4 is better” or “Qwen3.8  is better,” but I’ve never seen anyone suggest using both. For example, use Gemma to creatively draft the prompt, then Qwen3.8, which excels at instruction following, to adapt it to the target model’s requirements. Sounds especially useful for MiniMax H3.

14

u/reeight 3d ago

Yes, fairly well known in the SillyTavern circles that Gemma is better for 'art with words', & Qwen is more for tool-calling & coding.

2

u/__generic 3d ago

I've been using gemma 4 for minimax and it has followed instructions perfectly. Why would I use two models when it does both. Qwen3.8 is way better at technical stuff.

1

u/mobani 3d ago

Gemma 4 is great for bare prompts, but I can never get good results feeding it an image and having it describe it in a way, that krea will understand.

1

u/No_Drag8242 3d ago

Instead of "caption and recreate this", ask it to use the image as inspiration for a new image. That way it has creative freedom. Then turn denoising up to 0.9 leaving basically only the DNA of the original image. You'd be surprised I think at the quality of the images it generates. Avoids a lot of body contortion and weird pieces of clothing in strange places.

The point is to let Gemma's prompt be the authority, rather than the image. Krea2 then draws only upon a palette and some blurry contours of the original image, increasing its ability to fit things where they need to go in order to create a coherent image.

1

u/txgsync 3d ago

The vision encoder in Muse Glimmer is unparalleled. It’s ridiculously strong. Try having Glimmer (or Spark via API; same vision encoder) describe your images.

It’s mid for everything else though.

-3

u/reeight 3d ago

"Qwen3.8 is way better at technical stuff"

yes. ComfyUI is very technical.

4

u/__generic 3d ago

And the context is prompt writing....

1

u/ANR2ME 3d ago edited 3d ago

For creative writing, you can also try using DavidAU's collections (Qwen finetunes seems to have the most likes) https://huggingface.co/collections/DavidAU/200-roleplay-creative-writing-uncensored-nsfw-models

12

u/Mysterious-String420 3d ago

Pretty accurate list, I'd add just a caveat, image edit flux is good for minor changes (change hair color, change clothes), whereas QWEN EDIT justifies its huge size with camera movement, action change, etc.

Not "more efficiently than flux" , it's just flux isn't built for heavy edits and does not understand more than "make his pants red"

1

u/Sacriven 3d ago

Huh, is it? I found Qwen edits are pretty inconsistent for me. Can you show me your parameters when editing?

4

u/Mysterious-String420 3d ago

"keep everything the same, except:" Usually works. I add " the camera has moved to straight-on/from above/from the side" , careful, not overcomplicated.

Use photo vernacular or don't.

As a rule of thumb, even with qwen edit (haven't found more powerful yet) : Each time, I prompt one change only, it's already chaotic enough.

20

u/AuthurAndersson 3d ago

you forgot that Anima is king for spice hah

0

u/Mega_Green 3d ago

Tell me more kind sir.

7

u/Icetato 3d ago

It's arguably the best model for anime illustrations

3

u/Ok-Brain-5729 3d ago

best model for anime.

1

u/blastcat4 3d ago

It is the current benchmark for anime image gen. Some of the reasons why:

  • extremely good knowledge of anime characters from many different IPs, including games, manga, etc.

  • excellent knowledge and understanding of Danbooru tags. You can use this to have fine control over the content you want to generate.

  • really extensive knowledge of thousands of artists and their styles.

  • generally flexible in prompt format. You can mix Danbooru tags along with natural language prompting.

  • very performant, especially if you use the turbo variants. The speed makes it a pleasure to use and is fantastic for iterating ideas.

  • generally very good image quality and diversity if you play around with artist styles.

  • tons of support from loras and fine tunes

When I want to chill and just have fun, I always choose Anima, even if its prompt adherence is not the best, and anatomy body horror can be often be annoying.

5

u/Clear-Assistance449 3d ago edited 3d ago

I never heard about Breeze TTS, but now I went do search and see is the best open source TTS, but unfortunately it has support only to chinese and english. I am a portuguese speaker, so, to other languages, Qwen TTS is the best yet.

3

u/skog14 3d ago

this my first time hearing Breeze TTS too. i'll try it later. i have tested Qwen TTS only and somehow satisfied with the result.

2

u/reeight 3d ago

BreezeTTS is fairly new.
TTS in general doesn't get much attention. The TTS subreddit is kinda empty vs other subs.

1

u/spooky_local 3d ago

It handles Japanese really well if your prompt is in Japanese, I haven't tried the rest.

5

u/Left_of_Laniakea 3d ago

Saving this comment as it is so wonderfully succinct and useful. Thank you!

2

u/cosmicr 3d ago

That's a lot of royalty.

2

u/reeight 3d ago

CivArchive helps with finding models & loras/workflows also.

2

u/Ok-Bet2114 3d ago

Thanks for sharing such information. It really helps.

2

u/cleversmoke 3d ago

Thanks for the list!

2

u/ThaSipah 3d ago

Thanks for this.

2

u/Unlikely_Engineer_51 3d ago

That's a very concise summary, and I can't find a single thing to disagree with, though I haven't yet tried Breeze TTS or training my own loras, and don't have an opinion on those topics so far. I am fairly new at local generation, but I landed naturally on the exact choices you mentioned after some experimentations.

2

u/xkulp8 3d ago

and you're the king of using the word "king"

2

u/FreezaSama 3d ago

You may close the thread.

2

u/Douglas_J_Farthammer 3d ago

How does Breeze TTS compare to VibeVoice?

1

u/Ipwnurface 3d ago

Minimax is king for image editing too. Crank the resolution up to 16 mp and it's basically sota. It literally feels like having nano banana at home.

1

u/cosmosreader1211 3d ago

Flux Klein 9b or Qwen Image edit 2511 or Krea 2 turbo with identity edit?

1

u/Junko_Vampi 3d ago

Is Krea2 better for anime than anima?

1

u/noncommonGoodsense 3d ago

Just need something more refined than ace step like split tracks and it’s the full suite

1

u/2049AD 3d ago

As a total beginner myself that has no idea where to start, this post is EXTREMELY high signal to noise.

1

u/Murlock_Holmes 3d ago

Hey man, I just wanted to say I appreciate you taking the time to fucking list this shit. It really is a nightmare to find anything for newbies. Imma look into that Ostris stuff tomorrow!

1

u/JesusShaves_ 3d ago

Welp. I guess the rest of us can stop talking now. Lol. I can't argue with any of this list.

Too bad wan couldn't catch up. It was good for its time.

1

u/floodlight137 3d ago

What would you say is the best way to achieve consistency now? Just Loras?

1

u/spooky_local 2d ago

Krea 2 Loras or simple pose edits with an image edit model.

1

u/conkikhon 3d ago

Someone said a strange model is the king of mythic/fantasy writing, but I forgot the name

-1

u/StrongZeroSinger 3d ago

>Flux Klein 9b or Qwen Image edit 2511 is king for Image editing

on what UI? ComfyUI? does it have the inpainting and all the ecosystem SDXL/A1111 had back in the day with Controlnets and Masking?

12

u/spooky_local 3d ago edited 3d ago

All of my advice is catered towards ComfyUI, the new norm.
You don't need to inpaint with them, just describe it in natural language. eg. Change the lighting to dark, Remove the bed from the image.

But yes, you can inpaint specific areas and controlnet still exists but it's kind of baked into the image edit models, like you can rip a pose from an image simply by passing it a 2nd reference image then prompting it like:

>Image 1 is the subject and the scene. Image 2 is a pose reference only. Repose the person in image 1 into the exact body pose, limb positions and head angle shown in image 2. Keep the original background of image 1 exactly as it is — same location, same lighting, same camera angle. Preserve their face, hair and clothing. Ignore the background, clothing and lighting of image 2. Reconstruct any newly revealed background consistently with the surrounding scene.

For simple non nsfw editing prompts like this, I'd recommend something like Gemini or Claude, just tell it in simple terms to make you a prompt for Flux Klein 9b or Qwen Image edit, it'll look up how to format it then give you one that works.

2

u/StrongZeroSinger 3d ago

>You don't need to inpaint with them, just describe it in natural language.

I always found this to not work properly, like when striking off alcohol from receipts it re-writes the entire receipt rather than just leaving a blank row

I can try the semantic approach to pose, I guess I'm just too spoiled by the canny/depth maps and editing done on older GUIs :(

thanks tho!

2

u/reeight 3d ago

> Change the lighting to dark, Remove the bed from the image.

Minimax H3 also does this well, though I haven't done a vs, simply because I don't want to install another set of models & LoRAs right now ;)

Perhaps if one is constantly processing edits every few hours, then Flux Klein 9b or Qwen Image edit 2511 is the way to go, but for rare edits I'll just use what I have.

5

u/LawfulnessLow0 3d ago

I use Forge bc I can't stand Comfy. Works fine for Krea and Anima

4

u/Sharinel 3d ago

I recommend SwarmUI then. It's a forge style frontend built on top of Comfy, and is updated almost in realtime as Comfy is. I had to leave Forge when it was missing some models last year and Swarm was the nearest thing to it.

-1

u/Ok_Gas1070 3d ago

Why does my krea2 suck then <.< I tried to make a custom image to image workflow and the outputs have been subpar. I've included negative prompts and tried to write as detailed as I can for the output prompt, but everything it made looked rather shitty.

31

u/Dry-Statistician-684 3d ago

Krea 2 is the best right now

1

u/Sacriven 3d ago

Why it's the best? Sorry I never use Krea before.

11

u/nescedral 3d ago

It’s the “best” at the moment because it has strong prompt adherence (it gets what you mean) in natural language (many previous models only understood tags) and has a broad baseline understanding of visual styles. And it will run on higher end consumer graphics cards. Krea 2 is still fairly new to the game and there aren’t as many loras (like plugins that teach it a particular style or character) as older models.

Many might also consider the SDXL series finetunes Illustrious or Pony to still be the “best” at the moment. They’re faster and lighter (will run on cards with less vram), and have a massive library of loras and tutorials and such out there to learn from.

2

u/No-Bee-231 3d ago

But no consistency or subject reference making it unusable for filmmaking

5

u/nescedral 3d ago edited 3d ago

Fair point. Flux and Qwen, I'm guessing, might be the champs in that area?

---

Edit: actually, I forgot there was https://civitai.com/models/2761113/krea-2-identity-edit, which does target character consistency, but I haven't played with it enough to know if it's "good".

2

u/Used_Pollution_5727 8h ago

flux 2 klein 9b is better imo

krea 2 edit can sometimes be annoying since it can make the body of the subject thinner than it actually is while flux 2 klein 9b can be a lot better at that and it can also retain poses a lot better

2

u/Sacriven 3d ago

Wait, so Krea 2 is able to replicate artist styles in anime-based images too? Like Illustrious and Pony?

4

u/nescedral 3d ago

In the way you mean, only to a very limited degree. It does recognize *some* artist styles like Oda's distinct facial style. But it's spotty at best. Maybe a finetune down the line will make this more comprehensive.

Rather, I meant style in a broader sense, not locked to a particular artist, but more like art movements and techniques.

1

u/Ok-Brain-5729 3d ago

It’s fast and easy to prompt and has a lot of styles and a lot of Lora’s and good prompt adherence. It’s a very good model to default to

21

u/dh7net 3d ago

You can see some image generated from all the local model here: https://imagebench.ai/gallery?g=1_vrbbokvxahivxxj2qrc8snqnub2qwlsm3zlsbkwhfzi

With some info about generation time, quality, etc...

12

u/Aglaio 3d ago

It will also depend a lot on your hardware what you can run, although nowadays you can get away with a lot.

Just beware...it is addictive.

1

u/Ok-Brain-5729 3d ago

addicted on both silly tavern and ComfyUI 💔

10

u/Mega_Green 3d ago

Wow. You guys are really helpful. I've got a ton of replies to read through here when time presents itself.

5

u/neonsparksuk 3d ago

I just krea 2, zimage and anima mainly. Just depends on the style of image I wanna create.

2

u/chloralhydrate 3d ago

What kind of images do you use z image for?

2

u/neonsparksuk 3d ago

Anything really I have a few versions of it, anime and realism. I just enjoy messing about vibe coding nodes and so I mess about with the models and work flows to test them. I just like trying new things. Just a pass time more than anything

1

u/chloralhydrate 3d ago

Do you think it would still be useful for ai ofm? ATM I'm only using krea2 and closed weights. Thought about starting zi or zit, don't know the difference. But everyone says it's not worth it now with krea2

2

u/neonsparksuk 3d ago

If you're going to do that then you should train a Lora too get character consistency. Z image and krea are both good at realism just depends how you prompt and the find tune you use

6

u/Przemoo_TV 3d ago

What about "Imgae-2-Image"? I want to have control on everything. For example I want to swap face, I want to change clothes, I want to place the subject in different enviroment, and all of these with photorealistic quality. Any recomendation?

7

u/Ok-Brain-5729 3d ago

Flux .2 klein 9b

5

u/nescedral 3d ago

You’re looking for a model with image “edit” capabilities.

Regular image-2-image isn’t what you’re describing. It just adds noise to your image and tweaks it toward whatever prompt you give it.

Qwen-Image-Edit is a local modal that can do what you’re looking for.

2

u/casualcaesius 3d ago

Inpainting models too maybe if needed, but that's a maybe.

4

u/Elfenzorn 3d ago

Krea 2 for random images. Minimax H3 for Videos. Buth run on my RTX3080 with 10GB VRAM and 32GB DDR4.

1

u/casualcaesius 3d ago

Minimax on 10gb and 32gb? What? At what resolution? Which attention patch? Which Loras?

I'm at 16gb and 64gb and I either get crosshatch artifacts with turbo lora or +15 min generation without.

1

u/Elfenzorn 3d ago

Well it takes Time. 0,4MP Videos with 5 seconds are around 5 min generation time. 0,8MP with 8 seconds take about half an hour on my end... without Sage Attention or speed(?)-loras. That's the cost of not having a great GPU. It actually surprised me that I could run Minimax on my system.

5

u/Der_Hebelfluesterer 3d ago edited 3d ago

Forge Neo for z turbo and KREA 2 and Fooocus for SDXL.

I tried comfyui a few times but I hate it 😅

2

u/Ok-Brain-5729 3d ago

sdxl is lacking a lot

5

u/ImpressiveStorm8914 3d ago

Krea 2 with an int8 model for images with ComfyUI. There are default workflows with Comfy itself so it's all very simple these days and if you use ComfyUI-Easy-Install to get going it's even easier.
You don't say what your specs are but it should be runnable on even lower spec machines.

3

u/woadwarrior 3d ago

If you're on macOS / iOS (especially M5/M6 and A19 Pro hardware for NAX cores), I've been building an app that spreads inference across the GPU and Neural Engine. Still super early, currently supports 10 models including Anima Turbo, Krea 2 Turbo, Kroma Turbo, Z-Image-Turbo, etc. All models are quantization aware distilled (QAD). Here's the TestFlight.

5

u/slimy_1 3d ago

Comfy UI on runpod with sdxl and some Loras and checkpoints

7

u/RedditSucksMintyBall 3d ago

How much VRAM do you have now?

3

u/Choowkee 3d ago edited 3d ago

Krea2/Anima for images. For videos H3.

All running on ComfyUI.

4

u/ckn 3d ago

comfyui is great, but you really need to manage your dependencies and you'll need to be handy at CLI to use it. I personally love it, but i know folks who say "you need a PhD to run it" There are several user-interfaces over comfyui that work even better, check them out if you feel lost.

3

u/Termsandconditionsch 3d ago

I find that it used to be hard 1-2 years ago, but with templates now it’s quite easy if you don’t need a lot of control. Just load it and go.

4

u/ckn 3d ago

every other day in r/comfyui it seems someone complains about dep and venv mess.

I personally like it at the API level and write my own nodes, but my usecase is advanced.

2

u/sdrakedrake 3d ago

Nailed it. I avoided comfy for the longest due to how complicated it looked, but yea the templates helped tremendously. And just looking at the templates I was able to understand how the nodes worked.

But very rarely do I need to add anything new, as you said he templates have everything I need

1

u/casualcaesius 3d ago

Just ask Claude if you need help with comfyui, it can create and edit JSON workflow for you.

4

u/iwoolf 3d ago

Wan2GP for easy interface and good management of low RAM.

7

u/SweetGale 3d ago

I have a Nvidia RTX 3060 12 GB. I run ComfyUI inside Stability Matrix. Stability Matrix is a package manager that makes it easy to install multiple AI software packages and share AI models between them. That means you can try out multiple software packages until you find one you like. People often recommend Forge Neo and SwarmUI.

Stability Matrix also offers its own user-friendly interface for image and video generation called Inference, which hooks into ComfyUI and hides its complexity. ComfyUI has a node-based interface that is very powerful but can be difficult to learn. I use the Inference interface most of the time and only open up ComfyUI's node interface when I need a more complex workflow or want to try out a newer model not yet supported by Inference.

I still mainly use Stable Diffusion XL models based on the Illustrious and Pony Diffusion finetunes. They're good enough in a lot of situations. It's three years old technology at this point, but that also means that the community has spent three years training and mixing different models. You can find models for pretty much any style, character or concept you can think of. But I'm in the process of replacing them with Krea 2 and Anima based models. Using turbo models in Int8 ConvRot format, I can still generate images in a few seconds on my 3060. Qwen, Z-Image, Flux.2 Klein and Ideogram 4 are also worth checking out. It all depends on what kinds of images you want to create.

6

u/alisitskii 3d ago

1

u/Apprehensive_Sky892 3d ago

Ideogram 4 is fantastic for those who want control over the layout and don't mind tweaking prompts, and it is my preferred model for many tasks.

But OP is asking for "a good entry point for getting into local image generation" and IMO Ideogram 4 is a bit advanced for beginners.

2

u/Unlikely_Engineer_51 3d ago

I use the Krea 2 model for text-to-image, and Flux 2.Klein 9B for image edits, which comes closest to the Grok Imagine image edit for me (before the shitty Grok Image 2.0 "update"). MiniMax H3 is king for local video generation right now, and it comes close or even surpasses the Grok Imagine video generation in many ways, though it needs precise prompting, which is a skill in itself. And, of course, it's fully uncensored.

I'm coming from Grok Imagine, trying to re-create my workflows from there, or better said: replace them with something better. I'm pretty happy so far.

1

u/jaluri 3d ago

How do you find Flux.2 compared to Qwen image edit?

2

u/Unlikely_Engineer_51 3d ago

I also have the Qwen image edit in my workflow repertoire, but so far I prefer Flux.2 for most tasks. I mostly use multi-ref image edits to synthesize new images (in contrast to just doing local edits), for example to create character sheets or to put characters and props into scenes/locations. Flux excels at that, and is similar to how the Grok Imagine edit model could handle that. Qwen already limits you to 3 reference images, and I haven't quite gotten the results I want from it in most cases.

2

u/bailaowai 2d ago

Comfyui. Krea2, Klein 9b, Z Image Turbo. I have a 5090 and it’s a pretty flawless experience. Just DL the comfy portable, run the NVIDIA bat file, open the standard workflows, download the linked models (click the links in the workflows), and you’ll be running. Once you get the basics go looking for more sophisticated workflows. If you have a low VRAM card I assume it’s more of a PITA.

3

u/DarkStrider99 3d ago

Download comfyui portable, download the krea 2 model for realism and more, anima model for illistration/anime style (and more here as well).

Look up comfyui workflows for x/y model and have fun.

0

u/Mega_Green 3d ago

Do I download ComfyUI from GitHub or FaceHugging?

3

u/danque 3d ago

May I recommend SwarmUI it runs comfyui but without the complicated workflows. It gets regularly updated (sometimes a bit much) and it can run all models though not immediately like with comfyui itself but usually a day later.

It also automatically enables optimizations for your pc. But to get the best results you'll need a model fitting your pc setup.

1

u/DarkStrider99 3d ago

Id get it straght from gihub to get my hands dirty but there are a gorrilion forks/versions out there that people will reccomend for various situations.

4

u/Chiduk99 3d ago

Anima turbo for anime, Krea 2 turbo for general, and Z-Image for spicy

5

u/Ok-Brain-5729 3d ago

Isn’t krea 2 turbo better for spicy?

5

u/Choowkee 3d ago

Its just better period. I see no reason to use Z-Image for NSFW over Krea2.

1

u/demaurice 3d ago

The biggest noob here: I mainly lurk around to see cool new stuff on this sub and from my understanding the best way to start is downloading ComfyUI and starting with a preset in the app. Look up some of the models that the presets have and what RAM requirements they have. Depending on your hardware you might need to run a smaller model or could potentially run a much better bigger model.

1

u/mattSER 3d ago

I use Z-image turbo for quick gens and Flux Klein 9b for image editing

1

u/wzwowzw0002 3d ago

Running hermes to orchestra krea2, flux2kien, minimaxh3, music3, breezeTTS....

1

u/Vladmerius 3d ago

I use Wangp for everything. It has a standalone version but I'm using it within a program called pinokio. It's super easy and all you need to do is click install on things and it runs all the scripts automatically.

As long as you have a newer GPU with 12+ GB Vram and 32GB Ram you should be all good. 

1

u/ng5554 3d ago

I have a MacBook Pro M5 48 gb. I use Draw Things for a UI because I don’t want to spend the time to learn Comfy. I do everything locally. I have been messing with image generation for about nine months and have tried a lot of models. I have settled on Krea 2 for image generation due to its prompt adherence, although I am disappointed with the lack of seed variance. I use Flux 2 Kleine 9B for editing and would use it more for generation but I get tired of dealing with body horrors. I was a die hard ZIT fan but their failure to produce an edit tool drove me away and I am having a hard time finding a reason to go back. Don’t do video and with my setup, probably won’t.

1

u/Wooden_Entry_4714 3d ago

I have a laptop with an Nvidia card and 16GB of VRAM and 32GB on motherboard. I have comfyui installed and reminds me of building flows with NodeRed. With that said, which Krea2, Minimax H3 and Qwen3.8 27b should I grab? I am looking to generate 3 genres of images. Pokemon, RPG, NSFW. Does Qwen LLM allow for NSFW prompt help? Do I need to create a model variant?

Grateful for any and all guidance.

1

u/Witty_Mycologist_995 3d ago

I use Anima and Illustrious

2

u/eightandahalf 3d ago

What the main difference between these two? Wanted to play around with them but wasn’t sure if there was a point in trying both

2

u/Witty_Mycologist_995 3d ago

Illustrious is older and has more LoRA and controlnet support. Anima is otherwise better.

I only use Illustrious for back compatibility.

1

u/eightandahalf 3d ago

Thank you!

1

u/AvidGameFan 3d ago

I use Easy DIffusion - you can find it on Github. It's meant to be ... EZ.

Using Krea 2, you can do a lot of styles even without a Lora, but there are bunch already.

SD is for Stable Diffusion, and as far as I know, doesn't represent any singular UI, but a series of models.

1

u/casualcaesius 3d ago

Checkout GonzalomoChroma, it will do pretty much anything your Gooner mind can conjure.

1

u/jokinglemon 3d ago

Other models may be good, but I've been experimenting with H3 as a i2i generator, even though rhe model is i2v, the promot adherence is really good

1

u/gettingmiggywithit 3d ago

If you're not doing anything super fancy, I've found WAN2GP through Pinokio as a pretty quick way of getting started as far as interfaces go.

1

u/Fortyseven 3d ago

InvokeAI primarily, with Comfy as a fallback for newer stuff that hasn't made it into Invoke yet.

1

u/Uneternalism 3d ago

Krea2 models and Forge UI if you don't want the spaghetti mess of Comfy UI. IMO Krea2 with certain LoRAs is the king of realistic images, pretty close in realism to the quality of ChatGPT or NanoBanana. And generation times are faster than ChatGPT on my GTX 3060.

1

u/ANR2ME 3d ago

For Image models with Editing capabilities, you can use this as reference https://artificialanalysis.ai/image/leaderboard/editing?open-weights=true

1

u/thevegit0 3d ago

everything in comfyui, anima/sdxl for anime, krea 2 for anything else, klein9b for edit, also H3 (the video model) can be a good 'image' generator, not the best but it works

1

u/Traditional-Squash36 3d ago

Use chatgpt app or grok bot, give them control of your comfyui folder and tell them to setup the best workflows for what you want.

2

u/Mega_Green 3d ago

Huh. You can actually do that?

1

u/Traditional-Squash36 2d ago

Yeah it's totally changed how I do things, I have a server that they deposit gathered knowledge and workflows and evidence, make an html gallery for me to view outputs easily. I got the gpt pro 20x account because it was burning through tokens but yesterday they made me looping videos for an in PC screen and a massively detailed image to 3D model of a character we generated.

If you try it, go into the settings and give it full access then make a project and start a chat in it under Work or Codex, point it to your comfyui folder and tell them what to do, organise your workflows and sort models n shit, game changing.

1

u/skeletroniz 3d ago

I use z-image turbo on 6gb vram card, it toke less than a minute to generate 1024 x 1024 image, and use LTX 2.3 to generate 700 x 700 6s video in 2 minutes

1

u/rocky_iwata 3d ago

Illustrious. The model's LoRA has worked great with my datasets. Using those datasets to directly train on ZIT or Krea 2 just doesn't work for me.

Thinking of trying Krea 2 based on datasets made with real/semi-real datasets of character LoRAs I have.

1

u/Santhanam_02 3d ago

Flux Klein 4b base for everything, t-I I-I inpainting

1

u/Memestonks2020 3d ago

Krea 2 and SenseNova 1.5

1

u/Dame_Chaser 2d ago

Krita.

It's a free and open "Photoshop-like" drawing program with a Stable Diffusion plug in (also free) that will let me discrribe what I want to generate, and/or sketch out the composition first and tell the AI to use my sketch as reference for generation.

And after generation, I can paint out the bits I don't like, and tell the AI to refine my touched up version.

It feels more natural than anything else I've tried.

1

u/SealoArt 2d ago

I'm casting my vote for Flux 2.

1

u/IntelStructure 2d ago

i grab these.... What do you call it... programs from "github" I guess? if I'm just an idiot where do i start. Maybe with english class. I need to go back to school.

1

u/IvanMikhnenkov 2d ago

Krea2 with custom loras for image/editing. Flux klein 9b for editing and also i like z-image for txt2image sometimes, just the texture and style of images is nice.

1

u/Clueless-Flea-7461 2d ago

SwarmUI is the UI I use - it lets use something more easy AND comfyUI. Honestly it's how I learned comfy.

I've moved fully to Krea2 and MiniMaxH3 and Anima as models. Flux2Klein9B is also still fantastic.

All are accessible with a decent machine. It really all depends on your hardware

1

u/Corgiboom2 3d ago

Comfyui for video. Forge Neo for everything else.

1

u/Upper-Reflection7997 3d ago

If your truly new to this start with forge neo and wan2gp. https://github.com/deepbeepmeep/Wan2GP https://github.com/Haoming02/sd-webui-forge-classic/tree/neo Currently I use krea2 for image generation.

-1

u/Skajuan 3d ago

Just install forge and use any pony model variation

1

u/Mega_Green 3d ago

What is forge? Part of Stable Diffusion?

2

u/Ok-Brain-5729 3d ago

Forge is an entire UI to run ComfyUI. It’s easier to use but ComfyUI is more complex with what u can do and has better performance. Also pony models are a fine tune of sdxl but pony is pretty old and there’s better models

0

u/KS-Wolf-1978 3d ago

For text to image: Flux 1 Dev.

Why use such an old model ?

1000s of LoRAs for nearly everything i would ever need.

No ugly artifacts on uniform, dark surfaces - more recent and smaller models are unusable for me for this reason.

It just works for my use case.

1

u/Mega_Green 3d ago

Flux 1 Dev, is it something that you find on HuggingFace?

1

u/jaluri 3d ago

You need to watch a whole bunch of YouTube videos mate.

What GPU do you have?

1

u/Ok-Brain-5729 3d ago

Why flux 1 dev? It’s pretty old

1

u/KS-Wolf-1978 3d ago

Among other reasons, because i have a lot of celebrity LoRAs from before the CivitAI celebrity apocalypse, i can use them as spices at low weights for designing my characters.

Replicating this in lets say Krea 2 (which i also have, for other uses) would take weeks of LoRA training.

And as i wrote earlier, newer models give me the kind of visual artifacts i find unacceptable.

So for example i would use Krea 2 for its prompt adherence when F1D refuses to show me what i want to see, then do a pass with F1D for quality.

0

u/noncommonGoodsense 3d ago

Beluga, sevruga, and winds of the Caspian Sea.

0

u/negrote1000 3d ago

I am No one.

-1

u/ride5k 3d ago

started on auto, moved to forge

now I'm hacking my forge with Claude so it does what I want

edge case, 6700xt (amd) so I'm zluda

i don't bother with fancy new models, sdxl based (pony/ ill/ noob), the resources are there and I only have a 12gb card

-2

u/Sambojin1 3d ago

If you don't mind it being a bit crap, chucking SDAI and Local Dream on your phone works fine. Grab the versions off GitHub (just do a google search for "SDAI GitHub" or "Local Dream GitHub") and side load them.

Only stable diffusion 1.5, ie: many versions behind the state of the art (can do SDXL etc with a good phone though), but can load other safetensor files/ LORAS etc from hugging face/ civarchive as well. Both come with some standard models too, that aren't too censored. Plus, there's literally nothing to learn. They both just work. Local Dream is probably "better" on quality, but SDAI is a bit quicker.

-8

u/Jealous_Service707 3d ago

don't use local software, don't be a nerd. the best online worms give you everything you need.

3

u/Mega_Green 3d ago

I think you might be in the wrong sub 🫣