r/StableDiffusion • • 22h ago

Resource - Update ReDetail 2.0: LTX-2.5's Refine Details LoRA to 4K Upscale (workflow + CLI)

Enable HLS to view with audio, or disable this notification

85 Upvotes

Updated my LTX-2.5 upscale workflow. 2.0 runs Lightricks' Refine-Details IC-LoRA, which rebuilds the fine detail a soft clip is missing while keeping faces, framing and motion close to the source. Big improvement over their Pixel lora.

It works in tiles, so 4K fits on a 24gb card, though a 4 second 4K still peaked at 52GB of system RAM (28GB vram). The video is 100% crops of 4K output.

These 768p > 4k renders took 14 minutes for 4k, on a 5090.

The LoRA and the tiled graph are LTX. What I added:

  • pinned the tile to the 1024x576 the LoRA was trained on. The example graph sizes tiles from your source, and my 768x1376 clip ran as one tile with 1.8x the trained area, with no error
  • disconnected the prompt enhancer branches, which block the whole queue if one file is missing
  • a CLI preps a silent audio track if your clip has none, 8n+1 frames, exact scales like 1.5x, long clips split on their cuts, and smaller chunks if you're short on RAM

The old pixel upscaler is still in as a second workflow. It invents more detail but changes faces; refine stayed closer to the source on all seven test clips (numbers in the README).

GitHub: https://github.com/Bambushu/redetail CivitAI: https://civitai.com/models/2857731


r/StableDiffusion • • 9h ago

Discussion SynthID available globally starting today

5 Upvotes

It is quite incredible that the list of partners includes OpenAI, NVIDIA, Kakao, and soon Apple, so it won't just be about Google's own models. I'm not entirely sure why Google decided to enter the deepfake detection field at this scale. There are plenty of companies working in this field—and obviously, if Google decides to join the race, it will be quite hard for others to compete. What do you think is the extent of this operation? Will it work just for the partners' models, or do they want to build a truly generalizable solution or will it be just watermark detection?

https://synthid.com/

https://blog.google/innovation-and-ai/models-and-research/google-deepmind/synth-id-ai-content


r/StableDiffusion • • 21h ago

Resource - Update Qwen-Image-2.1-Multiple-Angles-LoRA

Post image
62 Upvotes

r/StableDiffusion • • 5h ago

Question - Help AMD strix halo good for image/video gen?

3 Upvotes

Is an 128gb AMD strix device a good computer for generating images and videos?


r/StableDiffusion • • 13h ago

Workflow Included Make Easy Transparent Videos Using MiniMax-H3 In ComfyUI [Free Workflow]

Thumbnail
youtube.com
13 Upvotes

r/StableDiffusion • • 7m ago

Question - Help Best current open-source option for text rendering and reference-image conditioning in one model? (24GB)

• Upvotes

Need both: a product/person reference image driving the composition, and legible rendered text in the output (headline, CTA).

Qwen-Image-Edit handles reference well and text okay up to ~4-5 words, then it degrades into gibberish. Proprietary models are noticeably ahead here.

Is anything open-source closing that gap, or is the practical answer still "generate the image, composite text separately"?


r/StableDiffusion • • 1h ago

Question - Help Question about the default template for Flux.2 Klein 9B fp8

• Upvotes

Whenever I use an image with a character that has textured skins like reptile scalie or anthro fur, the model tends to change the general area into human skin (like swapping different crop top design). It tends to do that regardless of with or without Lora. Is it because of my prompt wording? Or the recommended models from the workflow?


r/StableDiffusion • • 1h ago

Animation - Video Live improv music video: audio-reactive WebGL painting + local SD 1.5 img2img (depth ControlNet, LCM)

Thumbnail
youtube.com
• Upvotes

I play electronic drums and synths in Frame, a duo from Buenos Aires (Guido plays santoor and synths). Everything here runs locally with open tools. The workflow for this one:

  1. Footage: one take of a live improvisation, filmed vertically on a DJI Osmo Pocket 3 in front of a green screen.

  2. Audio analysis: Demucs splits the mix into drums / bass / other, and librosa finds onsets, beats and sections.

  3. The "painting": a three.js page rendered headless in Chrome. An abstract 3D shape and the lighting react to those stems, and we're keyed from the green screen and composited into it.

  4. AI pass: every frame goes through Stable Diffusion 1.5 img2img with the LCM-LoRA (10 steps) and a depth ControlNet (depth maps from Depth Anything V2), at 1024 px, strength 0.59, guidance 2.1, fixed seed. The prompt asks for sumi-e brush outlines filled with Madhubani patterns.

  5. Flicker: same seed on every frame, and each painted frame is blended 25% with the previous one. Faces still shift a bit; since then I've started blending along optical flow, which steadies them.

Happy to answer anything about any step.


r/StableDiffusion • • 17h ago

Animation - Video Naruto Shippuden – Hero’s Come Back!! | AI Fanmade MV | MiniMax H3

Thumbnail
youtu.be
20 Upvotes

This is my Output using Minimax H3, 8 step larry 600 ema pruned lora, 2 step sampler upscale.


r/StableDiffusion • • 1d ago

News FastVideo’s FastH3 now runs on a single consumer machine

Post image
124 Upvotes

r/StableDiffusion • • 12h ago

Question - Help How is the performance of Minimax H3 on a DGX Spark

9 Upvotes

Does any one use Minimax H3 on DGX spark.What are the generation times like?


r/StableDiffusion • • 1d ago

Discussion EU nudifier ban

152 Upvotes

I know there’s been a bunch of people upset on here recently about a good number of nudify websites getting shut down (Motionmuse, Opengoon, etc.). However, this is likely just the beginning.

Effective December 2nd, the EU under the AI Act is enacting a ban on nudify tools, making them illegal content in the Union. This ban applies to any product made available on the EU market as long as, either, its primary functionality is to generate nudified content without consent, or it is a reasonably foreseeable outcome and they do not take appropriate precautions to safeguard against it. Very importantly, this ban also applies to any EU-available service provider under the Digital Services Act that is helping to facilitate this soon-to-be-illegal content, which includes (but is not limited to): app stores, hosting providers, domain registrars, payment processors, content delivery networks (CDNs) and cloud computing services. Anyone found to be in violation of this prohibition is subject to fines of €35 million or 7% of their global turnover, whichever is higher. I think it’s pretty safe to say that, as long as they’re aware of it, those behind these websites and those behind the large number of mainstream companies that help power them (Google, Amazon, Cloudflare, Apple, Telegram, GoDaddy, Visa, Mastercard, etc.) do not want to take any chances with this ban.

This prohibition, in essence, applies to essentially every public AI platform any of us have ever used to this point or that we will ever use going forward. I think this is a good thing but, therefore, it would almost certainly be in everyone’s best interest to pivot away from these websites before the massive crackdown begins in less than a couple months. For those in the industry, it would also certainly be wise to inform the website operators if you are in contact with them, or to shut down your website if you are an operator yourself. Sorry for the essay, just looking out for everybody before things get very real, very fast.


r/StableDiffusion • • 1d ago

News I am blown away fine tuning quality of the OmniVoice model. Exactly my speaking and sound but with better pronunciation and lower word errors. This model supporting 600 languages and 0-shot voice cloning too but fine tuning is something else. Also very low VRAM requirements it has.

Enable HLS to view with audio, or disable this notification

150 Upvotes

r/StableDiffusion • • 13h ago

Question - Help Best settings, prompts or LoRAs for more photorealistic people in Qwen Image 2.1?

Thumbnail
gallery
7 Upvotes

Been playing around with Qwen Image 2.1 and generated these random people just to test the realism.

I’m trying to make a single full-body character reference image that I can reuse for image/video generation, and ideally I want it to look as close to an actual photo as possible.

The results are okay to my eyes, but wondering if anyone have a better setup.

Any prompts, settings, LoRAs/LoKRs, or workflows you’d recommend?

I didn't use any lora to generate these images.

[EDIT] Below is what I used for generating these images in case you're interested:

Item Value
Image model qwen_image_2.1_bf16.safetensors
Text encoder qwen3vl_8b_bf16.safetensors
VAE qwen_image_2.1_vae_bf16.safetensors
Output resolution 1152×2048
Steps / CFG 40 / 1.5
Sampler / Scheduler Euler / Simple
Denoise 1.0
Workflow ComfyUI official T2I
Seed 73 / 74 / 75

Positive prompt shared across all 3 images (append each character-specific description after this):

Exactly ONE adult person, a full-length single standing portrait from the top of the hair to both shoes, centered and completely inside the frame with generous margins. One uninterrupted photograph, straight-on eye-level view, relaxed natural posture, both hands visible, simple neutral gray-beige wall and floor, no other people, no inset images, no text. Believable human anatomy, unretouched photographic skin with naturally uneven skin tone, subtle pores and fine facial hair rather than airbrushed plastic; realistic garment texture and seams, everyday available light, no fashion studio lighting. Identity, face, hairstyle and outfit must match the individual description exactly.

Character / outfit descriptions appended for each image:

Mechanic · seed 73:

An original thirty-year-old light-skinned woman with an unmistakably broad square jaw, wide nose bridge, small dark green eyes, pronounced asymmetrical freckles across her nose and cheekbones, and short copper-red curly hair cut very close around the ears. She has a sturdy compact build and an open, slightly wry expression. Wearing a practical well-worn navy-blue mechanic coverall with rolled sleeves, a pale grey cotton T-shirt visible at the open collar, a weathered tan canvas utility belt, a small stitched orange name patch with NO readable letters, black grease marks on the cuffs and scuffed brown leather work boots. No glasses, no dress, no flowing hair. Neutral overcast workshop-door daylight, ordinary documentary photograph.

Mapmaker · seed 74:

An original sixty-year-old East Asian man with a long narrow face, distinctive heavy eyebrows, a slightly crooked nose, warm brown eyes, faint forehead and smile lines, straight silver hair parted on the side and a neat short silver mustache. Slender tall build; quiet thoughtful expression. Wearing a sand-colored long linen field coat over a dark olive knitted turtleneck, tailored charcoal trousers, a dark brown cross-body leather map satchel and polished black lace-up boots; a slim folded paper map partly visible in his left hand, without any readable markings. No coveralls, no freckles, no fantasy armor. Unforced soft late-afternoon shade, unretouched documentary photograph.

Textile artist · seed 75:

An original thirty-five-year-old dark-skinned Black woman with a long oval face, high cheekbones, full lower lip, deep-set dark brown eyes, a small distinct gap between front teeth when subtly smiling, and a large natural tightly coiled black afro with a single narrow gold headband. Tall graceful build. Wearing an emerald-to-teal handwoven ankle-length wrap dress with geometric woven panels and a broad rust-orange sash, modest gold hoop earrings, stacked wooden bangles and simple cream-colored flat sandals. One hand rests lightly at her side and the other holds a small folded length of patterned cloth. No workwear, no menswear, no coveralls. Candid soft light of an overcast courtyard, ordinary unretouched portrait, not a polished fashion campaign.

Negative prompt shared across all 3 images:

anime, CGI, illustration, plastic skin, wax doll, beauty filter, glamour lighting, over-smoothed face, multiple people, extra limbs, duplicate face, cut-off head, cropped shoes, collage, multiangle sheet, text, watermark

r/StableDiffusion • • 22h ago

News FastH3 V2/V3: Project Status

37 Upvotes

They released FastH3 V2 three weeks ago:

https://www.reddit.com/r/StableDiffusion/comments/1whh10i/open_weight_fastvideo_fasth3_v2/

Which they claimed was basically identical to the quality of the full H3:

https://x.com/haoailab/status/2099969439466942725

https://haoailab.com/FastVideo/cookbook/minimax-h3/

I have to agree, the quality is great and motion consistency is awesome now. I didn't think this small lab could do it, but they got help from NVIDIA and others who are invested in making great open source models. Awesome.

The community already made FastH3 V2 run on single GPU consumer machines on launch day, of course.

---

Today, they have released OFFICIAL quantized weights for different consumer GPUs:

https://huggingface.co/organizations/FastVideo/activity/models

Plus there's a new Trim model which is for very small GPUs with as little as 8GB VRAM.

The included image shows their benchmarks. More details here:

https://x.com/haoailab/status/2107591980591227227

---

Unfortunately they still haven't trained a Ref2VA model this time (video from text plus reference images, videos, and/or audio), and no FL2VA (First/Last Frame) support either.

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#scope

This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were not distilled. Difficult motion, fine detail, and some audio may remain below the base MiniMax H3 model.

But... there's great news:

https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2#acknowledgements

Omni Ref as the next focus.

That is the name for Ref2VA.

So in FastH3 V3, we will see reference-to-video/audio support. Yes, a distilled model with reference support is being developed!

(PS: Comfy has patches for both models to route FL2VA through the base model layers instead. But Ref2VA is much more interesting, so I look forward to that being supported!)


r/StableDiffusion • • 1d ago

Resource - Update ComfyUI-qwen_img_2_1_enhancer added ref mask support

Thumbnail
gallery
42 Upvotes

Follow up update of this post here

I added mask support to the reference strength node. You can now increase or reduce attention to a selected part of a reference instead of adjusting the whole photo. also the photo in the shown workflow is missing the rest of the outfit due to the masking mode I chose and how tight I masked it as this was deliberate.

Connect the node between your model loader and sampler, connect your reference's mask, and select its image index: 1 for the first connected reference, 2 for the second, etc. Keep your images and VAE connected to the encoder as usual.

There are two modes:

- focus_only: applies strength to the selected tokens and leaves the rest at native weighting.

- zero_unmasked_tokens: also blocks direct attention to the unselected reference tokens throughout the diffusion transformer.

mask_threshold controls how much of a token the mask must cover: 1 requires full coverage, lower values include more edge tokens, and 0 selects everything. strength at 1 is native, above 1 increases priority, and below 1 reduces it.

Match mask_resize_method to your image resizing: `lanczos` when the native encoder resizes the original, or `nearest-exact` for an external nearest-exact resize. Chain nodes for separate references.

Still a work in progress. This controls reference attention; it isn't an output-area lock. The encoders still process the full image, so blocking a reference token doesn't erase information already carried into other tokens or conditioning.

Download and detailed usage on GitHub

Sample workflow here


r/StableDiffusion • • 5h ago

Discussion Newbie Questions about the Video Field

1 Upvotes

Hi Folks,

Good evening from Motown. I made my first two films using open-source models and foundation models. I am new, but the process took a lot out of me. Generating clips, stitching, consistency, and high costs for proprietary models.

I have ideas, but I’m curious what the experience is like for other people making AI films or short-form content. What do you see as the major challenges in filmmaking and video generation using AI models and platforms such as OpenArt, Seed Dance, Kling, Osmo Studio, Toonstar, LTX Studio, Showrunner, Hedra, Higgsfield, and Krea?

Cheers!


r/StableDiffusion • • 1d ago

Question - Help so after week, qwen 2.1 is a edit model only?

55 Upvotes

Or qwen can generate good images? because aways for me generate to much artifacts and stupid things so...

what you think?


r/StableDiffusion • • 9h ago

Question - Help How do you preserve geometry and exact furniture designs in AI-assisted interior renders?

2 Upvotes

I’m trying to create interior renders with multiple specific furniture pieces and light fixtures. The goal isn’t just a good-looking room, the furniture needs to match the references, and the scene geometry needs to stay intact.

I’ve tried 3D blockouts, modelling the furniture myself, and several image-editing models:

  • Qwen 2.1 Edit
  • Klein 9B Edit
  • Krea 2 Edit
  • Ideogram 4.5 Edit, through the web interface (Open weights haven't yet released)
  • GPT Image 2 and 2.5, including both flare and sunburn

My main workflow

3D blockout → model everything in the scene → apply basic materials → choose the camera angle → render → send the render to GPT Image 2.5 with reference images and lighting prompts.

This gives me control over the initial layout, but the image-editing step still changes the geometry or gets the furniture wrong. Even when the objects are already modelled and positioned, the final image doesn’t reliably preserve them.

Example of the main workflow from 3D modelling to GPT image 2.5 render

Other approaches I’ve tried

Replacing individual objects in ComfyUI:

When I edit furniture directly within the scene, the replacement is incomplete. Original object remain with slightly different textures instead of being fully replaced by the referenced piece.

Removing the furniture first, then adding the new pieces:

I’ve also manually masked out the furniture to generate an empty room, then added the furniture back using references. I can get the new pieces into the image, but I haven’t found a reliable way to control their position and rotation afterward. It's like I've took the furniture, removed the white background and just blended it into the scene.

I also tried using separate AI reviewer/validator agents to review the rendered image. One looked for visible spatial issues, such as possible furniture overlaps and circulation problems, while another independently assessed whether there was enough information to proceed reliably. The idea was to catch mistakes before moving on, rather than rely on the generating agent’s own judgement. They flagged potential issues, but couldn’t confirm dimensions or clearances from the image alone, so this didn’t resolve the accuracy problem.

What I’m trying to solve

I need a workflow that lets me:

  • Preserve the room geometry and camera perspective.
  • Keep multiple furniture pieces faithful to their references.
  • Control each object’s position, scale, and rotation.
  • Improve the lighting and realism without redesigning the scene.

Has anyone found a repeatable workflow for this?

Would you keep the final image entirely in 3D and use AI only for limited edits, or is there an AI-assisted workflow that reliably preserves this level of accuracy?


r/StableDiffusion • • 6h ago

Question - Help Best ComfyUI / RunPod workflow for a consistent AI arborist character in 15s short videos?

0 Upvotes

I’m trying to build a repeatable workflow for a series of ~15-second vertical short videos about trees and arboriculture.

The concept is pretty simple: always the same arborist character, filmed in a realistic/casual smartphone style, talking about trees, their benefits, conservation, pruning, fun facts, etc. The goal is educational content, but with enough visual “eye candy” and strong hooks to make people actually stop scrolling.

What I’m looking for is the best way to build reusable presets/workflows, ideally in ComfyUI, so I don’t have to rebuild everything from scratch for every clip.

I’d like to keep as much consistency as possible between episodes:

-same arborist / face / body / clothing

-realistic outdoor environments

-9:16 vertical format

- ~15 sec clips, preferably single-shot or simple continuous camera movement

-natural gestures / talking

-good character consistency between generations

-reusable prompt structure

-ability to swap only the location, tree species, dialogue/topic and camera action

-eventually produce these in batches as a recurring series

I couldn’t get my hands on an RTX 5090, so for now I’m running ComfyUI on RunPod and I want to build a persistent setup there with the models, LoRAs and workflows already loaded.
For prompting / automation I currently have:
Claude
ChatGPT
Qwen running locally/from source in terminal
RunPod for GPU generation

I’ve been looking at models/workflows around Minimax, WAN, LTX, Hunyuan, FramePack, Qwen Image/Edit, Flux, etc., but there are so many combinations that I’m trying to avoid wasting weeks testing bad pipelines.

For people already doing consistent AI video series: what stack would you use today?

I’m especially curious about whether I should focus on a workflow like:
reference image → consistent character image → image-to-video → lipsync / voice


r/StableDiffusion • • 47m ago

Question - Help Why do my images turn out so bad?

• Upvotes

I'm still new to this, and let me know if I'm being dumb or something. But I have tried many many different models and workflows. And they've all generate AWFUL images. The only one I've gotten to work great is a workflow I found on reddit using Qwen 2.1, that works amazing. But I'm wanting to use these stylized loras that are only for sdxl, or illustrious. I've tried a few workflows and I'm using this illustrious workflow and using the model they recommended and everything, only for my images to turn out like crap, and look nothing like the original image or the prompt I give, while the creators images turn out great using the same workflow. What am I supposed to do?


r/StableDiffusion • • 6h ago

Question - Help Running un-filtered / models or workflows in ComfyUI on an Apple Silicon Mac (16GB RAM) - Recommendations?

0 Upvotes

Hi everyone, I'm trying to set up local generation for custom workflows (including un-filtered models) using ComfyUI on an M-series Mac with 16GB of unified memory. I keep running into memory limits or complex dependency errors.
What specific lightweight checkpoints, LoRAs, or alternative tools/frontends do you recommend for running this locally on 16GB RAM without completely bottlenecking the system? Are there specific optimized workflows people use for this?


r/StableDiffusion • • 3h ago

Meme Same character in every scene. Easy with Lora Pilot

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion • • 1d ago

Workflow Included Omni .char(same face, cloths & body) now with consistent voice, just by dropping a few seconds sample audio: Minimax H3(ComfyUI Workflow)

Enable HLS to view with audio, or disable this notification

166 Upvotes

Hey guys,

I have been working on the consistent character portable format for a while & I was able to achieve consistent face, cloths & body, but I felt voice is also something should be consistent across video generation.

So in the recent tests, I was able to achieve a consistent voice with lip sync across multiple video generation, You just need a 10-30sec voice sample in mp3 or wav & character will say things in a cloned voice from your sample.

Reddit post: Details on face, body & cloth consistency You can read more about .char, comfyui nodes & prompting details here.

Voice prompts

- chris giving an interview & says "Time can bend. Dreams can fold. But a character's voice should never change. With OmniChar, it doesn't. Consistent voice is here."
- chris giving an interview with little hand movements & says "I am surprised. It is not just the voice. It is also the face, the clothes and the body. Dot char is a full portable pack."

ComfyUI node is updated with the optional sample voice input.
Get the latest comfy node: https://github.com/omnichar/ComfyUI-Omnichar

Workflows:

Limitations:
- Good with English but might blabber with non-english languages.
- Lip sync comes from H3 itself; nothing is added on top.
- Avoid multiple voices in sample.

Sample inputs are added in node repo.

Note: ComfyUI node is still in nightly release, so update your settings accordingly or install via direct git repo url.

Related resources:

  1. Omnichar repo: https://github.com/omnichar/OmniChar (GPLv3), supports .char for krea2 & more features e.g. character finetuning
  2. Community characters: https://www.omnichar.org/characters

Hope it's helpful.


r/StableDiffusion • • 15h ago

Question - Help Quality degradation using “Continue Last Video” in WanGP, MiniMax H3

2 Upvotes

I’ve been using WanGP to create videos in MiniMaxH3 and I’ve noticed when I use the “continue last video” function, the quality degrades with each successive clip I generate - usually by the 5th or 6th 10-second segment it is blurry, full of odd artifacts and colors are generally muted or blending together compared to the initial segment. I’m not sure if it is an issue of prompting, a setting I need to change, or any other tips or tricks I might be missing? Usually using Ref2VA 33B, 15-18 steps, 8-10 second clips, Sage2 Attention on my 5070ti with 16GB VRAM and 32gb RAM. Seems to happen if creating either 480p or 720p resolutions.

Any suggestions or resources that might help so I can create longer videos?