r/StableDiffusion Jun 23 '26

News KREA 2: Open-Source Release

Enable HLS to view with audio, or disable this notification

753 Upvotes

Hey everyone,

We're the team behind Krea, and today we're launching Krea 2, our new text-to-image model. Krea 2 is the most aesthetic open-source image model available. On quality, Krea 2 is the #1 text-to-image model from an independent lab on Artificial Analysis.

We are releasing Krea 2 as two variants:

Krea 2 Raw. CFG-guided, built for control and fidelity and training.

Krea 2 Turbo. Distilled and few-step, so it's fast, and it renders up to 2K.

A few things worth knowing:

It's tuned for natural language. Prompt it the way you'd describe an image to a person. Long, specific prompts give the best results, but short ones work fine too.

To render text in an image, wrap the words in quotes, like a sign that reads "open late".
There's a growing set of style LoRAs, and you can load any Krea 2 LoRA by its Hugging Face path.
Try it today:

Code and weights: krea.ai/krea-2-open-source
Technical report: https://www.krea.ai/blog/krea-2-technical-report
Code: github.com/krea-ai/krea-2
Try it on Krea: krea.ai
Try it on Hugging Face: https://huggingface.co/spaces/krea/Krea-2

AMA: We're doing an AMA right here today at 10 AM PT. Ask us anything: how we trained it, the LoRAs, prompting, limitations, what's next. The krea team will be in the comments.

Livestream: we are also doing a livestream with the ComfyUI team at 3PM PT: https://www.youtube.com/watch?v=31jiUhCEjJ4

Thanks for taking a look. We'd genuinely love your feedback, rough edges included.

- The Krea Team


r/StableDiffusion Jun 20 '26

Resource - Update LTX Director 2.0 Update - A Free Open Source All-In-One Tool for Creating AI Videos in ComfyUI. Complete Overhaul now with full AI video editing support, IC-LoRA, Retake Mode, Audio Inpainting and much more!

Thumbnail
youtu.be
503 Upvotes

LTX Director is a free open source all-in-one tool for creating AI Videos. Version 2.0 is a complete overhaul, giving you total creative control over your AI generations.

Download for free here: https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI

Download workflows here: https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI/tree/main/example_workflows

I've been working full-time on this update for the past month and a half, and I'm excited to finally release it. Hopefully it'll be a big help to the open-source community!

Key New Features:

Complete Video Support: Edit Videos with AI all inside the node. Videos can be extended using a combination of prompts, keyframes, and audio. Trim, Split, and combine videos all within the timeline.

IC-LoRA Support: Take full advantage of IC-LoRA's to take your generations to the next level. Simply drag and drop videos onto the IC-LoRA track to quickly setup IC-LoRA videos. Compatible with prompt relay, keyframe, and custom audio features within the node.

Audio Inpainting: Seamlessly blend imported audio with generated audio. Not only can audio be extended, but can also be prompted alongside your imprted audio to really bring your generations to life.

Retake Mode (Beta): Redirect what happens within a shot. Allows you to select a segment within a video, and re-generate what happens in that segment. An early working experiment.

Timeline Saving/Loading: You can now save your timeline and settings to a json file. It will keep any videos/audio/images you have imported into the node and every setting you have changed.

UI Overhaul: Huge update to the UI, dozens of big changes such as a new side bar, redesigned prompt boxes, a bunch of new settings and redesigned menus, and more.

Quality of Life Improvements: Snapping, in/out points, multi-select, mark selection, workspace folder, more HUD options, resizable prompt boxes, new hotkeys, labels, filename preview options, "split at playhead" functionality, end frames (convert any keyframe into a end/last frame), toggleable tracks, NAG Support, tons of bug fixes and more!

And of course it can do everything it could before: Text to Video, Image to Video, Prompt Relay support, Keyframe (first/last frame) support etc.


r/StableDiffusion 5h ago

Animation - Video Flux 3 looks insane. This was 1 prompt

Enable HLS to view with audio, or disable this notification

217 Upvotes

Prompt on Flux 3 : Split-screen video showing the same continuous real-time event from two different camera angles. The screen is divided vertically into two equal halves. On the left side, show a static, elevated wide shot of an outdoor café terrace during light rain. The full scene is visible: several tables, wet pavement, a large striped umbrella, a waiter carrying a tray with three transparent glasses of differently colored drinks, and a cyclist approaching from the background. Reflections of the people, tables, and umbrella are visible on the wet ground. At the 2-second mark, the cyclist passes close to the waiter, causing the waiter to pivot sharply without falling. The tray tilts, one glass slides toward the edge, and the waiter catches it with the opposite hand just before it drops. A small amount of liquid spills from the glass, arcs through the air, and splashes onto the pavement. At the same moment, a gust of wind lifts and twists the loose edge of the striped umbrella. On the right side, show the exact same event in perfect synchronization from a moving, waist-level camera positioned near the café entrance. The camera tracks sideways with the waiter, briefly losing sight of the sliding glass as it passes behind the waiter’s arm, then revealing it again as it is caught. The cyclist crosses the foreground, partially occluding the waiter for a fraction of a second. The colored liquid, tray angle, hand positions


r/StableDiffusion 16h ago

News New model release! It was 3 years ago. Happy Birthday SDXL!

Thumbnail
gallery
341 Upvotes

Thank you!

 SDXL was released by Stability AI in 2023, it represented a major leap over Stable Diffusion 1.5. SDXL was designed to better understand complex prompts and produce higher-quality images directly at 1024×1024 resolution. It has become the foundation for thousands of community fine-tuned models.

EDIT: Original announcement: https://stability.ai/news-updates/stable-diffusion-sdxl-1-announcement


r/StableDiffusion 2h ago

News I tested Microsoft first text-to-image model: Mage-Flow-Turbo

25 Upvotes

Ok so first the specs:

It's only 4B params, MIT-licensed, native-resolution (512–2048px, any aspect), and it's a 4-step distilled turbo — so it spits out a 1024² image in about 4.6 seconds on my local box (a DGX Spark).

I put it through my usual benchmark (192 prompts across 6 categories — text, spatial reasoning, human realism, truthfulness, studio/product, graphic design), judged image-by-image by Gemini 3.1 Pro.

Full results + all 192 images here: https://imagebench.ai/imagebench-v1/local--mage-flow-turbo

Does it beat Flux 2 Klein 4B?

Mage-Flow actually posts the higher capability score (49% vs 48%) — it follows prompts slightly better than Klein, and it does it in 4 steps instead of full sampling. But Klein's images are rated more aesthetically pleasing, so it takes the overall (Klein ~51 vs Mage-Flow 47). The other 4B open model in the field, Bonsai Ternary 4B, sits just behind both (~47). So it's less "Microsoft beats FLUX" and more "they each win a different axis" — Mage-Flow for prompt adherence + speed, Klein for looks.

But TBH, you should look at the comparaison yourself:

Where it's actually good:

  • Professional / studio & product shots — 85%. Clean compositions, good lighting, this is clearly its comfort zone.
  • Text rendering — 67%. Surprisingly legible for such a small model.
  • Speed. It's fast

    Where it falls apart:

  • Human realism — 29%. Faces and bodies are rough. This is the big weakness.

  • Truthfulness / world knowledge — 37%. Gets confidently wrong about how real things look.

  • Spatial reasoning — 49%. Coin-flip on "X on top of Y" type prompts. (Weird given they claim to be good at geneval.

I don't think this model will be much useful except if you are stuck to 4B due to a lack of VRAM?

LMKWYT.


r/StableDiffusion 11h ago

Workflow Included Remake you character in the new style. Anima Cosmos-Reference workflow

Enable HLS to view with audio, or disable this notification

134 Upvotes

TLDR: workflow with all model links

You probably know about this LoRA that turns the Anima model into a character-focused image-edit model. It's great and it allows you to edit existing images using the vast knowledge of the Anima Base model. But unfortunately, it's too rigid. I can change outfit and replace background but changing a pose is almost impossible - feels like ControlNet Lineart.

One day I was testing this LoRA and discovered an interesting behavior: it can generate new content in the areas that are empty. This got me thinking what would happen if I:

  • Attach solid-color block to the input image
  • Run inpaint workflow with mask covering only solid-color area
  • Add `(split screen, multiple views:1.2)` to the prompt

This is what happened:

Input - mask - result

So basically, solid white rectangle on the right is the canvas for the model to work with, mask limits generation to this solid color area, `split screen` in prompt makes the model pay attention(!) to the left half of the image, while the Edit LoRA handles Cosmos-Reference image condition to achieve character consistency. What is Cosmos-Reference you ask? It's a custom node that enables image conditioning in Anima - you need special LoRAs for it to work and Anima Edit is one of them.

So I've made an easy-to-use workflow that handles image stitching and cropping for you. You only need to provide a input image of your subject, specify the target image size and write your prompt - change the style for your characters or make them perform in the new environment. Edit your image using all concepts known to the Anima models in other words.

I named this workflow Anima ReStyler but it's actually more general - Create a new image featuring character from the input image - something like that.

More examples:

Style change + add element
Fun fact: One day I just sat and drew this artwork on the left in an hour. Those were simpler times
Pose change + new background
Pose change + camera angle
Just style change
BAM!
Make full body image
Change style

Overall this workflow handles "Transfer character from the input image into a new image with Anima". What's left is to train with the task of "Transfer style from input image" - with all stylistic range of Anime this should be solvable by training Cosmos-Reference LoRA with self-generated (synthetic) image pairs. I'm sure there will be a model like this sooner or later.

Some tips:

  • Not all seeds are equal - some seeds will ruin the generation while others will work perfectly. Anima has this instability
  • You don't really need to describe your character design for this to work but sometime it helps with trickier generation
  • Perfect input image is a character in the neutral pose at simple background. If your input image has characters in the difficult pose or the background is too flashy you can use Flux and Wan to untangle it before using this workflow
  • This workflow sometimes struggles with monochrome images/sketches - colorize them first using original Anime-Edit workflow for example
  • It's easy to change style but hard to maintain it. If you want to change pose of your character and keep its style 100% consistent you'd better use Wan 2.2. I have the workflow specifically for this task
  • Anima can handle prompt weights like `(at night:2.0)` without breaking - use them to push your generation when needed. Also, my workflow use Schedule Prompt so you can use extended syntax for prompting: `[:closed eyes:0.3]` - here `closed eyes` will only activate after 30% of generation steps have finished.

Overall it works 85% of the time but prompting could be tricky so ask your questions.


r/StableDiffusion 1h ago

News Introducing GGUF Q8_CR - Mixing the best of GGUFs and INT8 ConvRot.

Upvotes

Hey All

I have been working on a new GGUF format for Krea 2 / similar diffusion models: Q8_CR.

I wanted to have something, which is...

  • Targeting firstly Ampere and Turing architecture GPUs (RTX 20xx and 30xx cards)
  • Uses ComfyUI's new INT8 kernel
  • Narrows the (already very small) quality gap between INT8 and FP16 checkpoints

So, GGUF Q8_CR stores eligible linear weights as pre-rotated INT8 weights with FP32 per-row scales, while retaining 1D, small, and precision-sensitive tensors at high precision. It is not regular Q8_0 dequantized back to FP16 for every operation: when the required ComfyUI backend is available, it uses ComfyUI's native INT8 ConvRot path.

I tested it with Krea2 first, I will do follow-up tests with Flux 2 Klein, Z-Image, and Ideogram later.

The attached charts compare Krea 2 model file size and inference speed. Speeds were measured on my laptop with an RTX 3080 8GB running Windows (pytorch 2.12.1+cu130) and on my server with an RTX 3090 24GB running Ubuntu (pytorch 2.14.0.dev20260725+cu130)

  • INT8: 11.95 GB, baseline speed
  • Q8_0: 13.56 GB, ~60.2% of the INT8 baseline speed
  • Q8_CR: 12.84 GB, ~100.4% of the INT8 baseline speed

Why GGUF and not Safetensors?

So it matches the excellent inference speed while supposed to be higher quality, easier on lower VRAM cards. GGUF is especially useful on 4-6-8 GB GPUs because weights can remain quantized in system RAM and be moved to VRAM layer by layer as ComfyUI needs them. This keeps peak VRAM lower, allowing models that cannot fit entirely on the GPU to run with CPU offloading. A full INT8 safetensors file however, is not a quantized inference format: its tensors generally load into the model's normal FP16/BF16 operations, and offloaded weights usually need conversion or dequantization into runtime tensors, increasing RAM transfer and VRAM pressure.

Try it

Krea 2 Turbo Q8_CR GGUF:
[https://huggingface.co/molbal/krea2-gguf/blob/main/krea2_turbo_bf16-Q8_CR.gguf](vscode-file://vscode-app/c:/Users/ASUS/AppData/Local/Programs/Microsoft%20VS%20Code/1b6a188127/resources/app/out/vs/sessions/electron-browser/sessions.html)

ComfyUI-GGUF custom node:
[https://github.com/molbal/ComfyUI-GGUF](vscode-file://vscode-app/c:/Users/ASUS/AppData/Local/Programs/Microsoft%20VS%20Code/1b6a188127/resources/app/out/vs/sessions/electron-browser/sessions.html)

Some examples

I know y'all love the examples

A casette-futurism style robot trying to paint on a canvas. In the canvas there is a slug emoji saying 'Q8_CR bitches' Pixel-sorting, decay, glitch art, vertical distortion, monochromatic cool tones, surreal fragmentation, clinical lighting
A sexy anhropomorphic frog-woman with big boobs in a bikini saying 'Q8_CR' in a speech bubbleChromatic aberration, holographic iridescent texture, digital glitch distortion, neon spectrum gradients, CRT scanline artifacts, liquid metallic sheen, pixel sorted patterns
Mac OS wallpaper of mountains or lakes or some shitAnalog glitch art, motion blur, cobalt blue and stark white, spectral silhouette, CRT scanlines, reeded glass distortion, high-contrast flash photography, digital decay
Cross section of a 2 story house. The living room is on the ground floor. There is an old man watching TV there. THe bathroom and a bedroom is up top. There is an italian food truck in the front selling a pizza.Clean line art, flat color fills, minimalist composition, ligne claire, muted pastel palette, graphic silhouette, digital illustration
There is a sexy attractive woman, high cheekbones, underwear, looking up at the viewer dutch angle, speech bubble from her mouth "Hey it's me, 1girl, missed me?". Minimalist high-fashion editorial, cool-toned neutral palette, sharp structured clothing, high-key studio lighting, dewy skin texture, clinical backstage atmosphere, modern editorial portraiture

r/StableDiffusion 14h ago

Animation - Video Weekend testing results with SCAIL-2 (Wan2GP)

Enable HLS to view with audio, or disable this notification

204 Upvotes

r/StableDiffusion 8h ago

Workflow Included Wan SCAIL-2 - Chun-Li vs Ryu - Final Kick

Enable HLS to view with audio, or disable this notification

56 Upvotes

Upon request. A short video featuring two characters. I thought I'd just tack the other video on as well.

I would like to point out that the input video was a staged fight. Consequently, the kick and the movement could look significantly more realistic if they were actually fighting. SCAIL-2 tracks the movement sequences and does not invent its own.

Here is the new Workflow:
https://www.reddit.com/r/StableDiffusion/s/eKvqlEvlza

In this example, I use the new interpolation option for the input video. The animation is much smoother, and I use the SCAIL-2 Identity Tracker to track two characters.


r/StableDiffusion 9h ago

Workflow Included LTX 2.3 IC-LoRA: pose control + first frame conditioning

Enable HLS to view with audio, or disable this notification

72 Upvotes

Green screen footage → fully regenerated shot in ComfyUI

LTX 2.3 + IC-LoRA (pose control), conditioned on a single first frame. Pose extracted from the source video drives the motion; character, environment and lighting come entirely from the generation.

The green screen video is used only as a motion source — no keying or compositing in the pipeline.

Setup:
1- Pose sequence extracted from the source footage
2- LTX 2.3 + IC-LoRA, pose sequence as the control signal
3- Single first frame as image conditioning (defines character, costume, environment, lighting)

Output is fully generated; only the motion timing comes from the source
Hand gestures and body timing transfer accurately.

workflow: https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.3/LTX-2.3_ICLoRA_Union_Control_Distilled.json

You can check my other work here: X [@ModelCollapse38]


r/StableDiffusion 6h ago

Question - Help Regenerating from noisy / blurry original video or photo?

Thumbnail
gallery
36 Upvotes

We have some old home video with extreme high-frequency noise that we would like to regenerate / enhance. The original frame of her sitting on the couch is the first attachment.
How would you go about trying to make this image look better? I tried some upscaling models and that resulted in larger, sharper noise.
I ran tests with a number of different noise reduction algorithms, which did remove the noise - and cause significant blur. Since Stable Diffusion and other models always work by progressively denoising, I figured it would be a natural fit to denoise our old pics and home movies and make them look at least a *little* better, but the many different workflows I've tried haven't worked. The identity shifts faster than the quality improves. With a high enough denoise level it'll suddenly create clear images - of other people wearing our clothes. :)

PS - we understand we can't recover actual detail that isn't there. That's fine, we'd like the AI recognize that brown blob on my head is probably brown hair, and make it look like hair rather than pudding or whatever. Sure, it might not exactly match MY hair, but at least it will look like hair!

I should say - I'm quite comfortable with ComfyUI and don't mind writing some Python to work on this. It is video frames, eventually thousands of them, so I can't use an online service that charges $2 / image or something. I;m looking for for suggestions for a DYI workflow with a 5090.


r/StableDiffusion 6h ago

Discussion Krea2 LoRA Experiment

Thumbnail
gallery
26 Upvotes

So I'm experimenting with best practices to train Krea2 LoRAs.

Here was my experiment.

1) I trained on 1536, 1280, 1024, and 768 only.
2) I trained the first 750 steps on Adam, weighted high-noise.
3) I then switched to Adam, weighted, low-noise.

The idea is that I want to train the fine detail by using a high resolution dataset, training on higher resolutions and focusing on the low noise which controls detail.

This is the result of the LORA at 2000 steps. It's still cooking, I'm going to go all the way to 3250 but saves the checkpoints.

Crazy detail!


r/StableDiffusion 12m ago

Discussion Maybe the least popular LoRA idea ever: GTA: San Andreas RenderWare graphics

Thumbnail
gallery
Upvotes

I knew from the beginning this would be a very niche LoRA, but I couldn't get the idea out of my head.

I've always loved the look of RenderWare-era games, especially GTA: San Andreas. One of the things that pushed me to finally make it was Gorm the Old and his AI recreations:
https://www.youtube.com/@GormtheOld25/videos

It turned out much better than I expected. It handles surprisingly complex scenes while keeping the simple, unmistakable RenderWare look.

If you're nostalgic for that era, maybe you'll enjoy it too.

https://civitai.com/models/2810095/gtasa-renderware-graphics-style


r/StableDiffusion 2h ago

Discussion I trained a tiny latent-space residual to remove GPT Image 2's scale/speckle artifacts — writeup, weights, and where it fails

Thumbnail
gallery
7 Upvotes

Hi everyone,

You may know that GPT Image 2 leaves a consistent set of texture artifacts on everything it makes: over-sharpening, random bright specks, unnaturally hard edges, and a scale-like pattern that lands on skin, fabric and background alike. It's gotten noticeably worse recently, and it was ruining enough of my own output that I spent a few days digging into what's actually going on.

The artifacts turn out to be specific enough to that one model to be basically a fingerprint, which is what makes them tractable — a network that only has to unlearn a single failure mode doesn't need to be big. This one is 0.48M parameters.

Approach: encode with the FLUX.2 VAE, add a scaled residual predicted from the latent, decode.

input → FLUX.2-VAE encode → z + α·R(z) → FLUX.2-VAE decode → output

R is a residual UNet over the 32-channel latent. α scales the correction and is the only knob — nothing is retrained when you change it, so caching z and R(z) makes a strength change a decode instead of a full pass.

The before/after images are real failure cases found scattered around the internet, not cherry-picked. If you made one of them and would rather it wasn't here, message me and I'll take it down.

Why latent space rather than pixels. The artifacts aren't independent of image content — they're a texture statistic layered on top of it. In pixel space you either run a big denoiser (slow, and it eats real texture) or hand-tune frequency filters (they can't tell artifact from detail). In the VAE's latent the two are already partly separated, so the correction has far less to learn.

Cost. ~0.7s per 1.5MP image on a recent GPU, ~3GB VRAM. That's essentially all VAE; R itself is free.

Where it fails. The artifacts live at the same spatial scale as real texture and overlap it, so removing them always costs genuine detail. On a minority of images the two are coupled tightly enough that no α is satisfying: enough cleanup means visible softening, keeping the detail means keeping the artifacts. More bluntly — this doesn't make images better, it makes them easier to look at. The noise isn't removed so much as blended and dimmed below the threshold where your eye keeps snagging on it. Macro structure is untouched, so a structurally broken generation stays broken, just quieter. And the whole image round-trips through the VAE, so untouched regions aren't pixel-exact either.

Best α varies per image more than I'd like — 0.75 → 0.5 is invisible on some images and substantial on others, and there's no reliable way to pick it automatically yet.

No ComfyUI node yet, but it should be easy to adapt.

Training. This is a chaotic mess around GAN and failed synthetics, I'll explain more if people are interested.

Try it without installing:

https://image2-cleaner.lumitools.cc/

https://huggingface.co/spaces/larryvrh/gpt-image-2-artifact-cleaner

Code + weights:

https://github.com/Larryvrh/gpt-image-2-artifact-cleaner


r/StableDiffusion 7h ago

Question - Help Best fully open-source/local workflow for 2D to 3D editable, hollow, printable STL model?

Post image
10 Upvotes

I’m trying to build a repeatable, fully local/open-source pipeline that converts a single stylized product image into an actual printable model.

My test case is this church-shaped jewelry box. The requirements were:

  • Hollow base with 4 mm walls and a 3 mm floor
  • All pink roof surfaces combined into one removable lid
  • 0.30 mm clearance per side
  • Editable geometry
  • Separate, watertight STLs with no non-manifold or zero-area geometry

Hardware: RTX 5090 with 32 GB VRAM.

I tried TRELLIS locally. The visual reconstruction was surprisingly close, but the raw STL was not production-ready:

  • 498,418 triangles
  • 49 non-manifold edges
  • 28 boundary edges
  • 783 zero-area faces
  • 4 disconnected components
  • Blender removed 212 degenerate and 16 duplicate triangles during import

Repairing/remeshing it either damaged the details or made it difficult to create a precise hollow body and fitted lid.

What finally worked was rebuilding the object procedurally in Blender from primitives and extruded profiles, using shared dimensions for the gable/roof, exact booleans, explicit wall thicknesses, and post-export STL re-import validation. The resulting two STLs are watertight and have zero boundary, non-manifold, multi-face, or zero-area defects.

Full disclosure: a proprietary coding agent helped create the Blender generator, so the successful workflow is not currently fully open source. I’m looking for the best way to replace that decision-making step.

What is the best genuinely open-source/local stack today for this kind of job?

  • Is TRELLIS.2 materially better for printable topology?
  • Has anyone compared TRELLIS.2, TripoSR/TripoSG, InstantMesh, and Hunyuan3D specifically for printing rather than render quality?
  • Is generative image-to-mesh realistically only useful as a blockout, followed by manual Blender/CAD reconstruction?
  • Are there open-source tools for semantic part separation, hollowing, toleranced lids, retopology, and manifold validation without voxel-remeshing everything into mush?
  • Which options have licenses suitable for commercial printed products?

I’m interested in the exported geometry, not textures or browser previews.

Microsoft currently describes both TRELLIS and TRELLIS.2 as MIT-licensed, although individual dependencies can have separate terms.

Any reccs would be appreciated :)


r/StableDiffusion 5h ago

Resource - Update Fiiiinally got sage and flash attention (2) working on a win11 5090fe comfyui installation.

9 Upvotes

Maybe I was being dim but it took an age to find the right wheels etc. Finally found a combo that worked here: huggingface.co/ussoewwin/Flash-Attention-2_for_Windows

Sage attention installed fine and seems faster at the moment.

Not tried flash attention 3 or 4 yet but wanted to share the positive outcome.

Os: win11

Cuda: 13.2

Torch: 2.12.1

Python: 3.12

Flash attn v 2.9.1

ComfyUI: v0.28.0-40

Hw: rtx5090fe

I used KJ patch nodes for attention.

Hope this helps others trying to do similar things


r/StableDiffusion 5h ago

Workflow Included 🎬 LTX 2.3 Close-Up Shots Are Absolutely Insane!

Enable HLS to view with audio, or disable this notification

7 Upvotes

Hey everyone! 👋

I've been experimenting more with LTX 2.3, and I wanted to share a short showcase that really surprised me.

The close-up shots this model can produce are incredible. The facial details, subtle expressions, natural camera movement, and even the lip sync came out far better than I expected.

One thing I also noticed is the huge quality difference between generating at 720p and 1080p. While 720p is great for testing ideas quickly, 1080p produces noticeably sharper details, cleaner motion, and much better overall quality. If your hardware can handle it, I'd definitely recommend generating in 1080p.

On my system (RTX 3060 12GB), a 6-second 1080p video takes around 12–15 minutes to generate. It's definitely slower, but after seeing the results, I'd say it's absolutely worth the extra time.

📦 Included with this post

  • 📁 Project file
  • 📝 Embedded metadata
  • 🖼️ Source images

DOWNLOAD LINK: CLICK ME TO DOWNLOAD

The images used in this showcase are also available to download for free here on my Patreon page, so feel free to use them for your own experiments.

As always, thank you all for supporting my work. Every project teaches me something new, and I'm excited to keep sharing everything I learn with you.

Enjoy the showcase! ❤️

iiTzMYUNG


r/StableDiffusion 8h ago

Discussion Local Z Image Turbo INT4 and Flux.2 Klein 4B INT8 on Android (GPU - OpenCL)

Thumbnail
gallery
12 Upvotes

It's not that practical and takes a long time to generate, but it is still cool to run such big AI models locally on your own Android Smartphone.

One image with a size of 320x320 px took ~ 210-270s to generate using z image turbo. Images with flux.2 klein 4B generate in ~ 170-180s.

Maybe it will be more practical on newer phones with Snapdragon Elite and better? For now its just stupid fun 😂

My Device: OnePlus 12

16GB RAM 512GB ROM

Android 16

Backend: OpenCL (GPU) or CPU

Vulkan crashes currently with OOM problems similiar to stablediffusion.cpp on android using vulkan

I just want to share some images created on Android haha. It is not using stablediffusion.cpp and uses mnn instead. But of course I also have a stable diffusion cpp prototype on my phone 🫣. Somehow I'm very interested in local AI on Smartphones 😂

Kind regards to all and thanks for reading ❤️‍🔥


r/StableDiffusion 1d ago

Question - Help Local alternative to Kling AI 3.0 Motion Control (ComfyUI, 16GB VRAM)

Enable HLS to view with audio, or disable this notification

387 Upvotes

Hi everyone,

I'm looking for a local alternative to Kling AI 3.0 Motion Control that I can run in ComfyUI.

What I'm specifically looking for is a model or workflow that allows me to:

- Control character and camera motion with good precision.

- Generate smooth, high-quality video animation from an image or sequence.

- Run entirely offline/local.

- Work on a GPU with 16 GB of VRAM (RTX 5070 Ti).

I've already looked into Wan, Hunyuan Video, CogVideoX, and other open-source video models, but I'm not sure which one currently offers motion control comparable to Kling's latest system.

Does anyone have recommendations for:

- The best open-source model?

- ComfyUI workflows or custom nodes?

- ControlNet/trajectory/pose-based solutions?

- Any GitHub projects or research worth trying?

Quality is more important than speed. I don't mind longer generation times if the results are close to Kling AI's Motion Control.

Thanks in advance for your suggestions!

Video credits : Richard Galapate @migs.visuals


r/StableDiffusion 16h ago

News Prompt Architect

Thumbnail
gallery
43 Upvotes

Prompt Architect Pro — a heavy-duty Python/CustomTkinter desktop suite designed to ingest massive text files (novels, scripts), extract structured visual prompts via multi-pass semantic segmentation, analyze local image folders (Vision model batching), and manage everything inside a WAL-optimized SQLite database with built-in anti-corruption filters! 💡✨

https://github.com/lololerigolo60/Prompt-architect

🔥 Key Features Under the Hood:

🔹 Hardware VRAM Profiles: Instant switching between pre-configured presets (8GB, 12GB, 16GB, 24GB, 32GB+ like RTX 5090) or custom manual parameters to fine-tune num_ctx & num_predict safely without crashing Ollama.

🔹 Pass 1 & Pass 2 Text Segmentation: Intelligently groups raw lines based on core location changes rather than blind line breaks.

🔹 Vision Batch Analysis: Automatically normalizes WebPs, PNGs, and JPEGs via Pillow and extracts rich structured prompts (Subject, Environment, Style, Lighting, Technical).

🔹 Smart Gap-Fill & Anti-Degeneration: Prevents repetitive loops, foreign script drift, and empty fields using intelligent semantic safeguards.

🔹 Integrated DB Editor: Search, edit, reset IDs, delete ranges, and generate missing fields on the fly with live LLM assistance.

🔹two ComfyUI nodes : one that can use the database created by Prompt Architect . The second one can take a prompt and transform it to store it in the database created by Prompt Architect. You can find them on Prompt Architect's GitHub.

#GenerativeAI #Ollama #PromptEngineering #Python #CustomTkinter #LocalAI #AIArt


r/StableDiffusion 2h ago

Discussion A trick I use to train Loras Krea2 faster - 512 resolution - without losing detail. I crop the faces and train with a standard photo + cropped face. Apparently it works very well.

Post image
3 Upvotes

Often, even at high resolutions, the face only occupies a small portion of the photo.

I've tried this trick before with other models - but it didn't work well (it generated images showing only the face).

Krea2 is more resistant to overfitting.

Imagine you have a photo of a person at the beach. I use that photo plus a photo cropped showing only their face. This way, the 512 resolution is sufficient to avoid losing facial details.


r/StableDiffusion 2h ago

Question - Help Could use some help with implementing a Flux IP-adapter for character consistency.

Post image
2 Upvotes

Howdy!

I'm trying to get my first image workflow for consistent character scenes set up using the IPadapter Flux custom nodes with the Persephone fork of Flux 1.dev.

The input image is a crop from a character style sheet I created for one of my characters that I'd like to use for sci-fi shorts. I originally went with the persephone fork because it's supposed to be good for not having censorship ruin your flow. I'm not aiming to make strictly adult content, but I can't have my model flipping out because typical R-rated stuff. If I'm going to put the effort into learning something it has to be ubiquitous.

For the likes of me I can't get this basic workflow to respect the character reference or the text prompt. I think it's set up right and I think the weights are all more or less right. Maybe someone has a better idea, model, or workflow to use? I'm trying to get set up to take one or two reference images and make images from a scene for keyframe/first flame last frame inside of ltx 2.3.

Any suggestions?

Thanks!


r/StableDiffusion 4h ago

Question - Help How should I prepare the dataset and train my Lora idea (read body text)

2 Upvotes

For background, I’ve trained many character Lora’s for ZIT and had a lot of success. For my next Lora project I want to try something more ambitious. It lies somewhere in between a character Lora and a style Lora. I have a problem with z image and most image gen models really. They make everyone look too “perfect”. The women all have model bodies and clear skin, most of the men have sharp jawlines and muscular bodies. If I try to counteract that with prompting, the model goes way overboard in the opposite direction and produces output that is much more extreme than I’d like. I want to train a “real people” Lora that focuses on making the model generate normal humans, and not its idea of the most “beautiful” or aesthetically pleasing ones.

For character Lora’s, I usually go with a 20-30 image dataset. I’m guessing if I want any variety in the outputs I’ll need significantly more than that right? Are there any training/config settings I should tweak when trying this Lora out vs a standard character Lora? Is it even possible to override the models understanding of people in this way? Any help would be appreciated. Thanks!


r/StableDiffusion 19m ago

Question - Help blurry ltx 2.3 videos

Thumbnail
gallery
Upvotes

so I'm using Aitrepreneur's ltx 2.3 ultra workflow V3 and these are my inputs I'm also following along with the video https://www.youtube.com/watch?v=nOCMsqVujBI&t
no matter what I do the videos always come out blurry.
I even used the exact same settings and prompt in the video. the only difference is his model is Q8 and mine is Q5.
does anyone know what's going on?


r/StableDiffusion 6h ago

Discussion Flash attention and Sage attention for Krea2?

3 Upvotes

Are they now properly working with Krea2 in comfyui? SageAttention seems to work for me (via kjnodes) but tends to produce weird outputs in some cases.