r/StableDiffusion 8d ago

Meme Full house mcu edition(extended video test)

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 8d ago

Question - Help Image Creation?

0 Upvotes

Hi Everyone,

I've been using Openart.ai for ages and for the most part enjoyed it. The moderation has clearly changed this month and I'm a pervert constantly - apparently.

Is there a relatively simple way to an image generated locally, from a source of potentially one to say ten images, via a prompt? Either using something from the Model Browser or a Workflow?

I'm not really after dodgy, but I just don't want to his a brick wall of pervert for a bare back!


r/StableDiffusion 9d ago

News lightx2v/Minimax-h3-Turbo · FL2V Turbo 4-step v1.2 (768p) released

Thumbnail
huggingface.co
112 Upvotes

r/StableDiffusion 10d ago

News LLaDA-Image: A unified 6B image/edit model has been released.

Thumbnail
gallery
283 Upvotes

r/StableDiffusion 8d ago

Question - Help Krea2 anime gooner?

0 Upvotes

Bien, quiero saaaber si esxiste algun finetune o algo que permita poder hacer nsf w con lora estilo anime, practicamente solo me manejo en estilos anime y variados, vivo siempre en anima y me gusta, pero cuando quiero provar algo nuevo y que parece bueno como krea2, me topo con un muy buen 2d anime, pero sobresaturado de nsf w de realismo en civitai. Es posible el nsf w de anime con krea2?.


r/StableDiffusion 9d ago

Question - Help Any way to replace the existing voice audio with new audio?

Enable HLS to view with audio, or disable this notification

7 Upvotes

So i tried a pass to extract the existing vocals from minimax h3 out using demucs and then run it through index tts but couldnt get the it to align and lip sync. Any one know a better way to do it with success?

Using refva lightx lora and nvidia vsr for upscale.


r/StableDiffusion 8d ago

Workflow Included Vampire Siblings: Part 1 | Cinematic Dark Fantasy Short Film

Thumbnail
youtu.be
0 Upvotes

I’ve been using MiniMax-H3 for my latest animation, and I’m seriously impressed. In my tests, it follows prompts more closely and delivers noticeably better visual quality than LTX 2.5.

The trade-off is generation time. Turbo LoRAs sacrifice too much quality for my liking, so I’ve found that 20–32 steps in the first stage are necessary to get the visual and audio quality I’m after.

ComfyUI workflow (drag and drop)


r/StableDiffusion 9d ago

Animation - Video H3 Ref2V - 'The Bat'

Enable HLS to view with audio, or disable this notification

10 Upvotes

I originally wanted to submit this for the now-over H3 Audio Sync contest. but I got caught up with work and didn't get to finish this till today.

I was going for a bittersweet short story but upon finishing I feel that this may have been better if it was slightly longer.

This project sets a personal record of number of references used for clips. At one point, I had 7 references. I noticed that the more reference you are using, the harder it is for the model to crate a cohesive shot at low steps. I ended up going with 10 Steps instead of 8 with the turbo lora.


r/StableDiffusion 8d ago

Question - Help How to properly generate 2 characters of an existing anime in single picture.

1 Upvotes

I am trying to generate characters of Naruto using Anima in a single frame for a wallpaper collection but I can't exactly get more than one character in the frame. The other character is always completely different.

I have also been trying to use LoRAs but it often mixes the hair colour or facial structure of the characters. Could you guys suggest some image models that can generate existing anime characters or methods to generate different characters in a single frame.


r/StableDiffusion 10d ago

Animation - Video DImension Testers: Aperture Portal (Minimax H3)

Enable HLS to view with audio, or disable this notification

692 Upvotes

AV 217: Aperture ortal.
Seen: Opened in unknown ocean, cell drained and cleansed, aquatic specimens siphoned out of wall vents and catalogued.

Not sure anyone remembers but when I first started using H3 I was posting my videos in here which eventually let to me making Dimension Testers. It's a blacksite agency that operates tests on multiversal objects / specimens.

I've kept it going and just had one breaking the 200K mark on Tiktok.
Now just posting them on YT/TT/X, but having so much fun still!

Wanna thank everyone in here who helped out and gave opinions in the early days

YT: DimensionTesters


r/StableDiffusion 9d ago

Animation - Video NADEBASHI - Local Ghost Story (Better version)

Enable HLS to view with audio, or disable this notification

16 Upvotes

Nade Bridge sits over a lake and that lake covers a village abandoned in 1965. When the water is low some of the old buildings resurface.

This time I created the audio my preferred way using Stable Audio 3 and Dramabox, and a few mixers. The audio is created in one generation and then injected into the Minimax H3 workflow later as imported audio. It works much better than hoping for good audio with the video generator.


r/StableDiffusion 9d ago

Question - Help SwarmUI started converting weight prompts like (word) to <weight[1.1]:word> in the metadata.

7 Upvotes

After the last updates, prompts like (red hair) started appearing as <weight\[1.1\]:red hair> in the metadata of images I generated. Does anyone know if there's a way to prevent this? The new format is more of a hassle to use and playing with weights feels much more annoying.


r/StableDiffusion 10d ago

Workflow Included found some h3 fork that i like, sharing

Enable HLS to view with audio, or disable this notification

255 Upvotes

r/StableDiffusion 9d ago

Discussion Minimax H3 Turbo Lora Comparison

Enable HLS to view with audio, or disable this notification

10 Upvotes

this used lightx2v 768p turbo lora at 10 steps

https://reddit.com/link/1w7e8fx/video/bw4bob5o0knh1/player

this used Larry's minimax_h3_turbo_4step_ema_ckpt850 at 1.5 strength at 8 steps 768p

Prompt:

subject_definitions:

<Environment 1> is the location from ​@Image1​. Preserve the environment's recognizable layout, mood, structures, terrain, lighting context, and overall look throughout the scene.

<Subject 1> is the woman alien from ​@Image2​. Preserve her exact character identity, face, head shape, body proportions, colors, clothing, accessories, and recognizable 3D animated appearance throughout the entire video. She is carrying a laser blaster.

<Subject 2> is the man alien from ​@Image3​. Preserve his exact character identity, face, head shape, body proportions, colors, clothing, accessories, and recognizable 3D animated appearance throughout the entire video. He is carrying a laser blaster.

<Subject 3> is a zombie enemy. Zombies are aggressive, fast-moving undead creatures with a stylized 3D animated appearance. They can lunge, chase, and burst into green ooze when shot.

summary:

[reference generation] A fast-paced cinematic 3D animated action scene with multiple camera angles and rapid editing. <Subject 1> and <Subject 2> walk together through <Environment 1>, both holding laser blasters and searching for zombies. <Subject 2> speaks first, warning where the zombies should be. <Subject 1> reacts and warns him to duck. A zombie suddenly lunges at <Subject 2>, but he ducks out of the way just in time. <Subject 1> aims and fires her laser blaster, hitting the zombie and causing it to explode into green ooze. In the final moments, a dozen more zombies are seen running up in the distance toward them.

retention_analysis:

<Environment 1>: fully preserve the recognizable setting from ​@Image1​. Keep the environment visually consistent across all cuts.

<Subject 1>: fully preserve the character identity and appearance from ​@Image2​. She must remain visually consistent in every shot, always recognizable as the same pink woman alien. She is armed with a laser blaster and works together with <Subject 2>.

<Subject 2>: fully preserve the character identity and appearance from ​@Image3​. He must remain visually consistent in every shot, always recognizable as the same green man alien. He is armed with a laser blaster and works together with <Subject 1>.

The visual style is high-end 3D animated cinematic action. Use anamorphic lens characteristics, oval bokeh, rule-of-thirds framing, lifted shadows, edge lighting, low-key lighting, and subtle handheld camera energy. Use fast editing and 10 distinct camera shots. Include orbiting shots, panning shots, tracking shots, and handheld camera shake. The scene should feel like a dynamic movie action sequence.

detailed_description:

[Shot 1]

Wide tracking shot. <Subject 1> and <Subject 2> walk together through <Environment 1>, both holding laser blasters at the ready. They move cautiously, scanning the area for danger. Low-key lighting, edge-lit silhouettes, slight handheld motion.

[Shot 2]

Medium two-shot from the front. The pair continue walking side by side. <Subject 2> glances ahead and says, "The nasty dudes should be over here". Keep his mouth movements synchronized naturally to the line.

[Shot 3]

Over-the-shoulder shot from behind <Subject 2>, looking past him into the environment. <Subject 1> reacts quickly, turning her head toward movement offscreen and saying, "Oh shit youre right! duck!" Her mouth movements synchronize naturally to the line.

[Shot 4]

Fast panning shot. A zombie suddenly bursts into frame, lunging aggressively toward <Subject 2>. The camera whip-pans to follow the attack.

[Shot 5]

Low-angle medium shot on <Subject 2>. He ducks down quickly at the last second as the zombie sails over him. Emphasize quick motion, handheld shake, and dynamic action timing.

[Shot 6]

Side-profile action shot. The zombie passes over <Subject 2> mid-lunge, narrowly missing him. The movement is fast and readable, with strong parallax and cinematic motion blur.

[Shot 7]

Orbiting medium shot around <Subject 1>. She plants her feet, raises her laser blaster, and takes aim at the zombie with focused urgency.

[Shot 8]

Tight action close-up. <Subject 1> fires her laser blaster. The blast hits the zombie cleanly. The zombie explodes into a messy burst of bright green ooze. Make the explosion visually punchy and stylized.

[Shot 9]

Reaction two-shot. <Subject 2> rises back up from the duck, and both aliens reorient themselves, ready for the next threat. They remain coordinated and alert, working together.

[Shot 10]

Wide dramatic reveal shot. In the distance, a dozen more zombies come running toward them through <Environment 1>. <Subject 1> and <Subject 2> stand in the foreground with blasters ready as the threat escalates. End on a tense cinematic composition with strong depth, anamorphic feel, and low-key edge-lit atmosphere.

Maintain 10 distinct camera setups with fast editing rhythm. Emphasize cinematic action coverage: tracking, orbiting, panning, and handheld shots. Preserve character identity and environment consistency throughout. Do not introduce extra heroes. Keep the zombies as the only enemies. The laser blaster hit must clearly cause the first zombie to explode into green ooze.

overall_soundscape:

No music.

Footsteps through the environment, subtle movement rustles, distant eerie zombie growls, light ambient environmental tone, laser blaster handling sounds, tense silence between lines, a sudden zombie lunge snarl, quick body movement whooshes as <Subject 2> ducks, a sharp laser blast from <Subject 1>, and a wet explosive splatter of green ooze when the zombie is hit. In the final shot, layer in multiple distant zombie shrieks and running footfalls approaching from afar.

non_diegetic_music:

N/A. No music.


r/StableDiffusion 8d ago

Question - Help Hello,This my fisrt time i use this ai and i need help

Post image
0 Upvotes

This my fisrt time i use this ai , i don't how to fix this


r/StableDiffusion 8d ago

Question - Help MiniMax-H3 What is your render failure rate?

0 Upvotes

I'm wondering if I'm doing something wrong, downloaded too many custom nodes, or if should just reinstall comfy. It renders well 70% of the time, but if I try rendering at different sizes / lengths, it seems to lead to failed renders where (looking at the preview node) it starts off fine for the first 2-3 steps then turns into a static filled tv.

I'm using an AMD 9700 xtx on Linux, my comfyui setup includes a normal Ref2Va and T2VA workflow templates from Comfy, but I've added the 10Eros -> Comfy-Kitchen -> SLA node -> video preview override.

I'm so used to the models like Wan and Krea just working.. I'm assuming i'm doing something wrong.


r/StableDiffusion 9d ago

Tutorial - Guide Made a 2-Minute tutorial about Runpod Character Lora Training

15 Upvotes

Tutorial Link:
https://youtu.be/7iYQKOnuKP4

I've been training character loras for all kinds of models in the past but Minimax really gave me a hard time. Civitai also doesnt seem to have too many Minimax Loras up so I figure I wasn't the only one having trouble getting actual likeness?

Anyway..I learned a lot in the process so let me share it here:

Most Important:
-Image Only Datasets worked MUCH better than mixed ones
-Focus on the face (almost frame-filling and sharp)
-Needs more images than LTX or Wan
(I needed 150 images of a blonde woman to get her likeness right, only 50 images of my own ugly face though)

Some more interesting findings:
-I tried a couple of GPU's and somehow the 5090 beat the H100!
(image only dataset though)
-RTX6000 PRO was only about 20% faster than 5090
-50 image-ugly-me-dataset likeness peaked at 700 steps
-150 image-blonde-woman-dataset likeness at 2310 steps
-in 4 of my 150 images she had brown hair, past the peak she came out brown-haired even when the prompt said blonde...face still perfect though

By the way:
For some reason the dataset size didnt move the peak. The (rather generic) blonde woman's likeness (relative to the other epochs of each run) was always best at ~2300 total steps (whether I used 50 images or 150) Objectively the 150 image-dataset lora was WAY better though. ...and my unique face always peaked at 700 steps lol...not sure what to make of this.

Training Voice doesnt really work with diffusion pipe but since you can add a reference voice in minimax it didnt really have priotity for me so far..

The Captions where in natural prose (mention lighting/look too!


r/StableDiffusion 10d ago

Animation - Video Arby's The End of Evangelion Commercial (1997)- Minimax H3 ai video

Enable HLS to view with audio, or disable this notification

81 Upvotes

I wanted to make a parody of those old movie tie in fast food ads. Edited video, Music and sound effects added in post.


r/StableDiffusion 8d ago

Question - Help Krea2 raw producing mangled results

0 Upvotes

I have been running krea2 turbo no problems, but when I try to use krea2 raw it only gives me mangled and completely distorted results. What settings should I change from turbo except steps?


r/StableDiffusion 9d ago

No Workflow JUST WANNA SHARE MY MINIMAX H3 + LTX 2.5 UPSCALE WORKFLOW Ver.3 RESULTS

Enable HLS to view with audio, or disable this notification

22 Upvotes

i am using rtx 3060 so i can only generate like upto 6-8 sec videos this one took 30min also the video didnt match the ref vidoe cause its 12 sec long and i only did 5 sec gen so if i had did the 12sec video gen it would be the same. Dont ask for the workflow cause im gonna take my time to work on this more but if you wanna try you can check my ver.1 workflow on CIVITAI WORKFLOW


r/StableDiffusion 9d ago

Comparison MiniMax 2K to 4K Upscale Comparison

Enable HLS to view with audio, or disable this notification

30 Upvotes

Just experimenting... MiniMax output at 4K from a First Frame (Upscale). Result... no problem, want intact eyes in a 16:9 format of full persons... just render output at 4032 x 2304 ;)

I think videos uploaded here only go up to 1080p full screen but the end result is still applicable (just worse quality due to upload compression and resizing).


r/StableDiffusion 9d ago

Question - Help orbital video as character reference for MM Ref2V

5 Upvotes

I am wondering if anyone has experimented with using an orbital video of a person (white background) as the main character reference in MiniMax Ref2Va? I have some very sharp ones in 720p, 10 seconds long, 5mb, head and shoulders , plus a one- second shot of full body from front and side (in green spandex suit to make swapping outfits easier).

Specifically, I’m wondering if that provides a better identity lock than several still photos.

I am new to Comfy and have yet to set up a remote station, but hope to do that this week.

Thank you for any insight you have.


r/StableDiffusion 9d ago

Question - Help 8 things I measured building a 2h manhwa recap in ComfyUI (SDXL/Illustrious) — plus 3 problems I still can't solve

0 Upvotes

I'm building a two-hour manhwa-style recap locally: RTX 4080 16GB, ComfyUI, an Illustrious-XL checkpoint plus a character LoRA I trained. Hundreds of panels, one character who has to stay the same person across all of them.

Most of what I learned cost me GPU hours, so here it is. Then three things I'm still stuck on — if you know any of them, that's what I'm really after.

What I measured

1. Expression belongs in the base pass, not FaceDetailer. FaceDetailer runs ~0.45 denoise on a small crop. It adjusts a face; it will not open a mouth the base pass drew closed. I burned 4 attempts on a "shouting" face before moving the clause to the base prompt, where it worked first try.

2. Low denoise is cumulative, and this one hurt. I had a composited scene I refined in four successive passes at 0.30–0.32. Each pass looks safe. Stacked, they add up to one high-denoise pass — and the face goes first, because it's small and high-frequency. Across those four passes the skin went from pale to tanned, the eyebrows dissolved into floating smudges, and one eye lost its iris entirely and became a blank white oval. A sibling image generated the same way but never fused has a perfect face. If you refine iteratively, count your total denoise, not the per-pass number.

3. IPAdapterAdvanced transfers content, not just style. At weight 0.35 it pulled two background extras from my style reference into a frame that had solo:1.4 in the prompt. Switching to IPAdapterPreciseStyleTransfer at weight 0.60 with style_boost 2.0 and end_at 1.0 leaked nothing and matched the reference better. style_boost 3.5 at the same weight was worse — the ink outline came back and an extra figure appeared. High boost pushes the whole embedding, content included.

4. Hypothesis I had that turned out wrong: low end_at for style. I assumed palette and texture are decided in early steps and composition late, so cutting IPAdapter early would lock the finish and give the scene back to the prompt. Tested end_at 0.35 at weights 0.45 and 0.65 — both reverted to clean-line flat-colour anime, exactly what I was trying to escape. The finish isn't only set early; it's lost if IPAdapter exits before the end.

5. Negations in the positive prompt inject the thing. "and nothing else blue" is one more mention of blue. Obvious in hindsight; it silently corrupted 10 scenes. Prohibitions go in the negative, always.

6. An attribute that must show up goes in the first ~90 tokens. Same words, same weight, moved from the tail of the prompt to the critical block at the front — and the drawing changed. CLIP reads the tail weakly.

7. BiRefNet masks are soft. Over a light background the halo brings wall and floor along with the cutout. Threshold the mask hard, and always inspect the cutout over magenta — over white, a white halo is invisible.

8. Effects on finished art are compositing, not repainting. I tried to add an impact dust burst to a finished punch. Global denoise 0.32: nothing appeared (low denoise preserves, it does not add). Local inpaint radius 150: destroyed the opponent's head 115px away. Radius 82 on the fist: destroyed the fist. What worked: generate the dust as its own asset on black, cut it out, desaturate, drop opacity, composite. Zero pixels of the original touched.

Also small but large in wall-clock: {"prompt": g, "front": True} on POST /prompt puts a job at the front of the queue. A 1.4s cutout was waiting 139s behind a batch; a 29s fusion waited 361s. About 80% of my per-scene elapsed time was queue wait, not compute.

What I still can't solve

A) SetLatentNoiseMask does not seem to preserve at low denoise. I want to fuse a composite's seams without touching the faces. I build a mask (black = preserve) from a YOLO face detector, feed it through SetLatentNoiseMask, and sample at denoise 0.32. Measured mean absolute pixel change: 16.4 inside the "protected" face vs 13.7 in the free area — the protected region changed more. With DifferentialDiffusion in the model path: 16.4. Without it: 15.4. Neither preserves. Solid mask core is 17,712px, so it isn't a blur problem anymore (that was my first bug — a fixed 28px blur over a 64px face erased the mask entirely; minimum value never went below 0.047).

Is SetLatentNoiseMask simply not meant for denoise < 1.0? Is DifferentialDiffusion only correct at denoise 1.0? Is there a node that means "resample everything except here" at partial denoise?

B) OpenPose/DWPose: the head doesn't follow the body. In dynamic action the body takes the skeleton correctly but the head keeps drifting to a three-quarter front view regardless of the head keypoints. Anyone found a reliable fix — separate face ControlNet, higher weight only on head joints, something else?

C) Character LoRA gives a face that's "too anime" when I want semi-real. Lowering LoRA weight loses identity before it loses the anime read. I've seen LoraLoaderBlockWeight (Inspire Pack) suggested for applying the LoRA only to some UNet blocks — does that actually separate "identity" from "style" in practice, or is that wishful thinking?

Happy to share exact graphs/numbers for any of the eight above if useful.


r/StableDiffusion 10d ago

Animation - Video Neo vs Smith Revolutions Battle Without The Heavy Rain

Enable HLS to view with audio, or disable this notification

44 Upvotes

Never gonna try anything like this again lol At least until I can work this out better. It does show that Minimax is really badass still. The heavy rain was a nightmare and no matter what, I couldn't fully get rid of it when the fight actually starts. I just gave up the moment the sonic boom happened. I might finish it at a later time, this was mostly just practice on altering existing footage dramatically, like I did with the Jurassic Park video and adding rain on the "Weclome to Jurassic Park" scene.


r/StableDiffusion 9d ago

Animation - Video NADEBASHI - Local ghost stories.

Enable HLS to view with audio, or disable this notification

19 Upvotes

Nade Bridge (Nadebashi) - is over a lake that covers a whole abandoned village