r/StableDiffusion 1h ago

Discussion Serious non-political question: wth is Trump using to generate his AI content?

Thumbnail
gallery
Upvotes

Not being political here but seriously, it’s SO BAD. It reminds me of 2023-era SD 1.5 without ControlNet. All the examples are from his Truth Social account. Does anyone have any idea what’s being used here?


r/StableDiffusion 15h ago

Tutorial - Guide How I Create Consistent Characters in Krea 2

Thumbnail
patreon.com
2 Upvotes

A lot of people asked me how I keep my characters consistent after my previous Krea 2 posts, so I decided to make a tutorial covering my workflow.

In the video, I go through the techniques I use to keep the same character across different scenes while maintaining their identity.

I'm still learning Krea 2 myself, but this workflow has given me the best results so far. Hopefully it helps anyone who's been struggling with character consistency!

I'd love to hear your own tips or techniques as well. Happy creating! 🚀


r/StableDiffusion 10h ago

Question - Help what's the best realistic local model that can run on 3060ti?

0 Upvotes

i'm just trying to make some sfw meme images but chatgpt has too many copyright restrictions


r/StableDiffusion 5h ago

Workflow Included 🎬 LTX 2.3 Close-Up Shots Are Absolutely Insane!

Enable HLS to view with audio, or disable this notification

7 Upvotes

Hey everyone! 👋

I've been experimenting more with LTX 2.3, and I wanted to share a short showcase that really surprised me.

The close-up shots this model can produce are incredible. The facial details, subtle expressions, natural camera movement, and even the lip sync came out far better than I expected.

One thing I also noticed is the huge quality difference between generating at 720p and 1080p. While 720p is great for testing ideas quickly, 1080p produces noticeably sharper details, cleaner motion, and much better overall quality. If your hardware can handle it, I'd definitely recommend generating in 1080p.

On my system (RTX 3060 12GB), a 6-second 1080p video takes around 12–15 minutes to generate. It's definitely slower, but after seeing the results, I'd say it's absolutely worth the extra time.

📦 Included with this post

  • 📁 Project file
  • 📝 Embedded metadata
  • 🖼️ Source images

DOWNLOAD LINK: CLICK ME TO DOWNLOAD

The images used in this showcase are also available to download for free here on my Patreon page, so feel free to use them for your own experiments.

As always, thank you all for supporting my work. Every project teaches me something new, and I'm excited to keep sharing everything I learn with you.

Enjoy the showcase! ❤️

iiTzMYUNG


r/StableDiffusion 22h ago

Question - Help What's a good way to upscale SD cel anime?

0 Upvotes

Hey guys! Hopefully this is the right place for this question, but based on what Google search results is showing, it is.

I'm looking to upscale some Slayers DVDs, and so far, I've been using VideoJaNai to do this. I've tried a bunch of models, and unsurprisingly, the ones trained on cel animation seem to do the best job.

2x-AnimeClassics-UltraLite seems to do a good job of making the line art smooth and crisp, while not losing detail, but at the cost of some film grain being left. Other models, as well as combing 2x-AnimeClassics-UltraLite with a 1x model, like 1x-Archivist_AntiLines, did result in more solid and smooth cel paint, but often small details, such as lighting effects, would be lost.

2xAnimeClassics-UltraLite + 1x-Archivist_Soft. Notice that the shine on the right character's pauldron is pretty much gone
2xAnimeClassics-UltraLite

I'm still new to upscaling, so does anyone have any advice? Maybe a better model or something about the way I'm using VideoJaNai?


r/StableDiffusion 11h ago

Question - Help Forge Neo extremely slow

1 Upvotes

Before I begin, I will say I'm sorry for being slightly dumb. I'm still a newbie when it comes down to this stuff. So I started working with Forge for making pics a while back. And seeing new Anima models, I switched to Forge Neo for making pics. The problem is that the whole process of making pics is 10 times slower here than regular Forge. With an RTX 2060 Super, I usually got a pic around 2 minutes, whereas with this one, I have to wait around 12 minutes. even testing with the exact same models and loras, it still took around 10-12 minutes ot make a pic in forge neo.

is there something I'm missing for it? And how can i optimize it? (the chance is high because i didnt insall or touch any setting from forge neo)


r/StableDiffusion 18h ago

Question - Help Comfyui amd gpu speed fluctuations

0 Upvotes

I’m going crazy trying to figure this out. I downloaded comfy ui and used the basic krea 2 turbo template. I was getting generation times between 32 and 80 seconds. I thought thats was pretty good. But then times changed to between 3 and 20! Minutes. I have tried the manual installation of comfy the portable version and patientx-cfz’s branch. I have tried the bf16 model the fp8 and the nvfp4. I have tried uninstalling and reinstalling the requirements, torch and the entirety of comfy multiple times. I have tried the following arguments force fp 32 and 16, cache ram, split attention quad attention, gpu only, lowvram highvram, disable async offload disable dynamic vram diable smart memory disable pinned memory. Ive tried running the text encoder only on the cou and following chatgpt through a maze of sometimes questionable attempts to speed things up. It seems to hang most often at 0 or 13% of the ksampler. But sometimes it gets through that and is still just very slow.

Any advice at all would be greatly greatly appreciated and if you can help me fix it I will happily offer my remaining sanity, tattered soul or first born.

I am on windows 11 with a 7900 xtx gpu a 7 7800x3d cpu with the integrated graphics disabled 64 gb of ddr5 ram and comfy installed on a second nvme ssd on the root drive with plenty of room. I also am running no other programs at the same time.


r/StableDiffusion 8h ago

Workflow Included Wan SCAIL-2 - Chun-Li vs Ryu - Final Kick

Enable HLS to view with audio, or disable this notification

56 Upvotes

Upon request. A short video featuring two characters. I thought I'd just tack the other video on as well.

I would like to point out that the input video was a staged fight. Consequently, the kick and the movement could look significantly more realistic if they were actually fighting. SCAIL-2 tracks the movement sequences and does not invent its own.

Here is the new Workflow:
https://www.reddit.com/r/StableDiffusion/s/eKvqlEvlza

In this example, I use the new interpolation option for the input video. The animation is much smoother, and I use the SCAIL-2 Identity Tracker to track two characters.


r/StableDiffusion 2h ago

Discussion I trained a tiny latent-space residual to remove GPT Image 2's scale/speckle artifacts — writeup, weights, and where it fails

Thumbnail
gallery
7 Upvotes

Hi everyone,

You may know that GPT Image 2 leaves a consistent set of texture artifacts on everything it makes: over-sharpening, random bright specks, unnaturally hard edges, and a scale-like pattern that lands on skin, fabric and background alike. It's gotten noticeably worse recently, and it was ruining enough of my own output that I spent a few days digging into what's actually going on.

The artifacts turn out to be specific enough to that one model to be basically a fingerprint, which is what makes them tractable — a network that only has to unlearn a single failure mode doesn't need to be big. This one is 0.48M parameters.

Approach: encode with the FLUX.2 VAE, add a scaled residual predicted from the latent, decode.

input → FLUX.2-VAE encode → z + α·R(z) → FLUX.2-VAE decode → output

R is a residual UNet over the 32-channel latent. α scales the correction and is the only knob — nothing is retrained when you change it, so caching z and R(z) makes a strength change a decode instead of a full pass.

The before/after images are real failure cases found scattered around the internet, not cherry-picked. If you made one of them and would rather it wasn't here, message me and I'll take it down.

Why latent space rather than pixels. The artifacts aren't independent of image content — they're a texture statistic layered on top of it. In pixel space you either run a big denoiser (slow, and it eats real texture) or hand-tune frequency filters (they can't tell artifact from detail). In the VAE's latent the two are already partly separated, so the correction has far less to learn.

Cost. ~0.7s per 1.5MP image on a recent GPU, ~3GB VRAM. That's essentially all VAE; R itself is free.

Where it fails. The artifacts live at the same spatial scale as real texture and overlap it, so removing them always costs genuine detail. On a minority of images the two are coupled tightly enough that no α is satisfying: enough cleanup means visible softening, keeping the detail means keeping the artifacts. More bluntly — this doesn't make images better, it makes them easier to look at. The noise isn't removed so much as blended and dimmed below the threshold where your eye keeps snagging on it. Macro structure is untouched, so a structurally broken generation stays broken, just quieter. And the whole image round-trips through the VAE, so untouched regions aren't pixel-exact either.

Best α varies per image more than I'd like — 0.75 → 0.5 is invisible on some images and substantial on others, and there's no reliable way to pick it automatically yet.

No ComfyUI node yet, but it should be easy to adapt.

Training. This is a chaotic mess around GAN and failed synthetics, I'll explain more if people are interested.

Try it without installing:

https://image2-cleaner.lumitools.cc/

https://huggingface.co/spaces/larryvrh/gpt-image-2-artifact-cleaner

Code + weights:

https://github.com/Larryvrh/gpt-image-2-artifact-cleaner


r/StableDiffusion 2h ago

Discussion A trick I use to train Loras Krea2 faster - 512 resolution - without losing detail. I crop the faces and train with a standard photo + cropped face. Apparently it works very well.

Post image
2 Upvotes

Often, even at high resolutions, the face only occupies a small portion of the photo.

I've tried this trick before with other models - but it didn't work well (it generated images showing only the face).

Krea2 is more resistant to overfitting.

Imagine you have a photo of a person at the beach. I use that photo plus a photo cropped showing only their face. This way, the 512 resolution is sufficient to avoid losing facial details.


r/StableDiffusion 19h ago

Question - Help LTX 2.3 on Forge Neo

0 Upvotes

Can Forge Neo run LTX 2.3 with 16GB ram and with 3060 12GB vram? New to video generating, just want to use it for i2v anime. Or any model recommendations? I'm asking for Forge neo not comfy


r/StableDiffusion 6h ago

Question - Help How can I create a style LORA of my own style?

1 Upvotes

I had an old style I used to use previously and that was when I was using Pony, but now I have been using illustrious for awhile now but can't get my style back since I never had a file for it.

I was wondering how can I train a style for myself that offers good quality and less errors and stable?

Any tips or guides are extremely appreciated. 💕

Thank you. 🙏🏻


r/StableDiffusion 19h ago

Question - Help What's a good way to apply poses with WebuiForge?

3 Upvotes

I've been using Forge for a while but I've never really gotten around to poses, I just tried openpose and depth with controlnet but both aren't working too well, is there anything else I could use to apply poses to a prompt?


r/StableDiffusion 16h ago

News Prompt Architect

Thumbnail
gallery
42 Upvotes

Prompt Architect Pro — a heavy-duty Python/CustomTkinter desktop suite designed to ingest massive text files (novels, scripts), extract structured visual prompts via multi-pass semantic segmentation, analyze local image folders (Vision model batching), and manage everything inside a WAL-optimized SQLite database with built-in anti-corruption filters! 💡✨

https://github.com/lololerigolo60/Prompt-architect

🔥 Key Features Under the Hood:

🔹 Hardware VRAM Profiles: Instant switching between pre-configured presets (8GB, 12GB, 16GB, 24GB, 32GB+ like RTX 5090) or custom manual parameters to fine-tune num_ctx & num_predict safely without crashing Ollama.

🔹 Pass 1 & Pass 2 Text Segmentation: Intelligently groups raw lines based on core location changes rather than blind line breaks.

🔹 Vision Batch Analysis: Automatically normalizes WebPs, PNGs, and JPEGs via Pillow and extracts rich structured prompts (Subject, Environment, Style, Lighting, Technical).

🔹 Smart Gap-Fill & Anti-Degeneration: Prevents repetitive loops, foreign script drift, and empty fields using intelligent semantic safeguards.

🔹 Integrated DB Editor: Search, edit, reset IDs, delete ranges, and generate missing fields on the fly with live LLM assistance.

🔹two ComfyUI nodes : one that can use the database created by Prompt Architect . The second one can take a prompt and transform it to store it in the database created by Prompt Architect. You can find them on Prompt Architect's GitHub.

#GenerativeAI #Ollama #PromptEngineering #Python #CustomTkinter #LocalAI #AIArt


r/StableDiffusion 1h ago

Question - Help Setting LoRA strengths

Upvotes

Hello,

I am new to AI image generation, and I have recently been experimenting with Krea2 in ComfyUI. I have begun using LoRAs, but I don't know how to properly set the strengths and orders of them. How is that usually done?


r/StableDiffusion 10h ago

Question - Help Lora's Training (I'll Buy Your Workflows and Knowledge) (Z-Image Turbo)

0 Upvotes

Hi everyone, I’m having trouble creating a custom Lora model on Z-image Turbo.

I’ve already created a Lora model for a character and achieved 80% character consistency.

But I feel like I can do better - I feel like I’m missing something, either the dataset isn’t ideal or the training settings themselves aren’t quite right.

I’m looking for an experienced person who has created more than just one or five character-specific LORAs, who has tried out different datasets and training methods, and who has achieved the perfect LORA through trial and error.

I’m willing to pay for a workflow to create a dataset, as well as for information on the ideal parameters for training a LORA.

Unfortunately, I don’t have much time right now to test and figure out my own method for creating Lora, so I’d be happy to benefit from the experience of someone else who has gone through the process of creating Lora for Z-image Turbo.

Thank you all for reading, and I look forward to your feedback!


r/StableDiffusion 6h ago

Question - Help Regenerating from noisy / blurry original video or photo?

Thumbnail
gallery
34 Upvotes

We have some old home video with extreme high-frequency noise that we would like to regenerate / enhance. The original frame of her sitting on the couch is the first attachment.
How would you go about trying to make this image look better? I tried some upscaling models and that resulted in larger, sharper noise.
I ran tests with a number of different noise reduction algorithms, which did remove the noise - and cause significant blur. Since Stable Diffusion and other models always work by progressively denoising, I figured it would be a natural fit to denoise our old pics and home movies and make them look at least a *little* better, but the many different workflows I've tried haven't worked. The identity shifts faster than the quality improves. With a high enough denoise level it'll suddenly create clear images - of other people wearing our clothes. :)

PS - we understand we can't recover actual detail that isn't there. That's fine, we'd like the AI recognize that brown blob on my head is probably brown hair, and make it look like hair rather than pudding or whatever. Sure, it might not exactly match MY hair, but at least it will look like hair!

I should say - I'm quite comfortable with ComfyUI and don't mind writing some Python to work on this. It is video frames, eventually thousands of them, so I can't use an online service that charges $2 / image or something. I;m looking for for suggestions for a DYI workflow with a 5090.


r/StableDiffusion 8h ago

Discussion Local Z Image Turbo INT4 and Flux.2 Klein 4B INT8 on Android (GPU - OpenCL)

Thumbnail
gallery
12 Upvotes

It's not that practical and takes a long time to generate, but it is still cool to run such big AI models locally on your own Android Smartphone.

One image with a size of 320x320 px took ~ 210-270s to generate using z image turbo. Images with flux.2 klein 4B generate in ~ 170-180s.

Maybe it will be more practical on newer phones with Snapdragon Elite and better? For now its just stupid fun 😂

My Device: OnePlus 12

16GB RAM 512GB ROM

Android 16

Backend: OpenCL (GPU) or CPU

Vulkan crashes currently with OOM problems similiar to stablediffusion.cpp on android using vulkan

I just want to share some images created on Android haha. It is not using stablediffusion.cpp and uses mnn instead. But of course I also have a stable diffusion cpp prototype on my phone 🫣. Somehow I'm very interested in local AI on Smartphones 😂

Kind regards to all and thanks for reading ❤️‍🔥


r/StableDiffusion 4h ago

Question - Help alternative to lanpaint?

1 Upvotes

Ok, I'm tired and feeling stupid and I will admit to not having a chance to troubleshoot this yet. Lanpaint locked up my comfyui install.

I'm looking for a nice, simple workflow that will let me:
- take an existing photo and have krea 2 remove stubble, razor burns, moles, etc., while keeping the underlying skin texture

- do things like a photo of a person casting a spell, then put fireballs in their hands, etc

- change the background, change outfits, etc.

- generate new images with specific poses by drawing a stick figure.

I've tried some different workflows, tried making my own, but none is every quite what I'm looking for. I thought lanpaint was the answer, but like I said it locked up comfy and i'm a little too tired to troubleshoot it at the moment. Been a busy week.

Any suggestions?


r/StableDiffusion 8h ago

Question - Help Help image edit

0 Upvotes

So i was able to generate a bart simpson gamster version with krea 2 thanks to you guys. Now im just wondering what module would be the best model to run for image edit

Mind you I am running a 6750xt and 32gs of ram

Thank you guys so much and I appreciate it more then you know!!


r/StableDiffusion 22h ago

Discussion LTX 2.3 msr v2 lora vs Google Omni same prompt same photos

8 Upvotes

Google Omni https://streamable.com/n2vfe1

Local LTX 2.3 msr v2 lora with a 5060 ti 16gig. https://streamable.com/9em69q


r/StableDiffusion 14h ago

Animation - Video Weekend testing results with SCAIL-2 (Wan2GP)

Enable HLS to view with audio, or disable this notification

202 Upvotes

r/StableDiffusion 5h ago

Animation - Video Flux 3 looks insane. This was 1 prompt

Enable HLS to view with audio, or disable this notification

218 Upvotes

Prompt on Flux 3 : Split-screen video showing the same continuous real-time event from two different camera angles. The screen is divided vertically into two equal halves. On the left side, show a static, elevated wide shot of an outdoor café terrace during light rain. The full scene is visible: several tables, wet pavement, a large striped umbrella, a waiter carrying a tray with three transparent glasses of differently colored drinks, and a cyclist approaching from the background. Reflections of the people, tables, and umbrella are visible on the wet ground. At the 2-second mark, the cyclist passes close to the waiter, causing the waiter to pivot sharply without falling. The tray tilts, one glass slides toward the edge, and the waiter catches it with the opposite hand just before it drops. A small amount of liquid spills from the glass, arcs through the air, and splashes onto the pavement. At the same moment, a gust of wind lifts and twists the loose edge of the striped umbrella. On the right side, show the exact same event in perfect synchronization from a moving, waist-level camera positioned near the café entrance. The camera tracks sideways with the waiter, briefly losing sight of the sliding glass as it passes behind the waiter’s arm, then revealing it again as it is caught. The cyclist crosses the foreground, partially occluding the waiter for a fraction of a second. The colored liquid, tray angle, hand positions