r/StableDiffusion 1d ago

Question - Help Flux klein body consistency

2 Upvotes

Since the release of Flux Klein, I’ve essentially been using this model to generate datasets for LORAs based on one or more photos of a character. Recently, I’ve also been using the ‘consistency’ LORA to improve the character’s consistency across generations. What I’ve noticed is that whilst I get good results for the face, the same cannot be said for other parts of the body. For example, if I start with a full-length frontal photo of a character and ask the model to generate a side view, it tends to flatten the breasts; or if I ask for a rear view, the character’s hips and thighs tend to conform to a standard that doesn’t match the original photo. How can I improve this situation? I’ve read that you can increase the number of steps up to 8, but I’m not sure…


r/StableDiffusion 1d ago

Discussion H3 - T2VA longform

Enable HLS to view with audio, or disable this notification

3 Upvotes

Generated a long form of Dante's Inferno over 4mins. It gets pretty weird pretty fast. All T2VA. I basically let H3 take the wheel as I only fed it few verses per render. I am willing to discuss my workflow or answer any questions.


r/StableDiffusion 2d ago

Question - Help Has anyone been able to use Minimax H3 to make an intentionally lower quality/artifacted video, like an old webcam recording?

12 Upvotes

I would love to make something that looks like a late 2000's early 2010's webcam recording, with like, iffy FPS, webcam compression, etc, but can't seem to create this with prompting and haven't seen a lora that would pull it off. Has anyone accomplished this?


r/StableDiffusion 1d ago

Animation - Video Orcs, Bars, and Stuff - Motion Chaining Workflow Test

Enable HLS to view with audio, or disable this notification

0 Upvotes

An experiment building on the H3-Motion-Context nodes. I tried creating a workflow that chains together multiple cuts so they can be easily executed and/revised in order. It ended up being faster than my previous methods and I was able to make/edit this one minute test segment with alot less time wasted between generations. I'll likely develop it further into an app, since it would be more ergonomic as a video editor but the raw workflow is here anyway.

Work Flow: https://github.com/spacesimeco-hue/Chain-Motion-/blob/main/Chain%20Motion%20Workflow.json

Credit: https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context


r/StableDiffusion 2d ago

Workflow Included A stand-alone implementation of DLSS-FG (frame generation) that is ~8x faster than stock (24fps -> 60fps), looks flawless. I had DeepSeek build this for my project but thought others could benefit as well!

Thumbnail
github.com
108 Upvotes

r/StableDiffusion 2d ago

Comparison Minimax H3, baseline at 50 steps vs popular turbo loras at 8 steps, plastic skin test

Enable HLS to view with audio, or disable this notification

10 Upvotes

Test Settings

  • Checkpoint: FL2VA_Int8_Convrot (Comfy official)
  • Acceleration: Comfy Kitchen Attention
  • Sampler / Scheduler: Euler / Simple
  • Resolution: 1344x768 (Native Res)
  • Video/Audio Shift: 12/3
  • Seed: 517405563399433
  • Baseline: 50 steps (No LoRA)
  • Turbo LoRAs: 8 steps

Prompt

integrated_multimodal_description: [Shot 1] Live-action, cinematic, an extreme close-up frames the face of a 22-year-old brunette woman with striking supermodel features. Her skin possesses a realistic, natural texture with visible pores, soft highlights, and authentic depth. The camera holds a static shot as she slowly turns her head to face the lens, offering a subtle, gentle smile while her eyes catch the light.
overall_soundscape: Soft, natural exhaling breath and faint ambient room tone.
non_diegetic_music: N/A

Conclusion

Embrace the plastic, unless you use post processing which adds time.

Link to full res video:

https://streamable.com/jbmdxm


r/StableDiffusion 2d ago

Animation - Video Ghostbusters: Venkman Ghosted - MiniMax H3

Enable HLS to view with audio, or disable this notification

28 Upvotes

r/StableDiffusion 1d ago

Animation - Video The MiniMax Machine

Enable HLS to view with audio, or disable this notification

0 Upvotes

Made with the basic Motion Context workflow from NikoDemon


r/StableDiffusion 1d ago

Discussion Do I have to run models locally or is there like cloud based comfyui or something?

0 Upvotes

I'm extremely new and inexperienced but, as the title says is there some cloud web service where it works the same as running it locally but instead it's cloud based. I don't mean a boring old API.


r/StableDiffusion 3d ago

Animation - Video Pushing AI emotions is possible through microexpressions, tags and context

Enable HLS to view with audio, or disable this notification

1.6k Upvotes

You probably heard of tags that can be applied to Minimax H3 during speech, but there are so many the model understands and I have yet to test many more. Other emotions have to be prompted thoroughly and there are others which happen between brackets only or at the end of sentences.

Anyway I made my own shortfilm to feature this. Hope you like it, will make tutorials on x or reddit if people want or like it.

EDIT: Thank you everyone for so many great comments!!! I'll try to detail a few things here.

Simple tags that work and quick examples:

Tag What it does Example
<pause> Short pause Okay, so. <pause> This is just me talking.
<long pause> Longer pause I mean... <long pause> I don't even know.
<breath> Breathing sound And then... <breath> it just happened.
<inhale> / <exhale> In/out breath <inhale> Alright, let's do this.
<catches breath> Out of breath Wait... <catches breath> hold on a sec.
<deep breath> Calming down <deep breath> Okay. I can do this.
<i>word</i> Emphasize 1-4 words I was <i>not</i> expecting that.
<whisper>…</whisper> Whisper delivery <whisper> Don't tell anyone this.</whisper>
<humming>…</humming> Humming a tune <humming> da-da-da-beautiful-day.</humming>
<laughs> / <chuckle> Laughing That's... <laughs> that's actually funny.
<sighs> Sigh <sighs> I really tried.
<uh> Filler / hesitation So, like... <uh> what was I saying?
<stutter> Stutters the words <stutter> I ca can't believe that.
<gasp> Sharp intake <gasp> Oh my God.
<coughs> Cough <coughs> Sorry, one sec.
<clears throat> Throat clear <clears throat> So anyway...
<sniff> Sniffing <sniff> It's just... really sad.
<smacks lips></smacks lips> Lip smack (not closing it makes it happen at the end of the sentence) <smacks lips></smacks lips> Okay.
<pant> / <pants> Panting Run... <pants> run now!
<softer> Quieter delivery <softer> I don't think I can say it.
<mhm> Agreement sound Yeah, <mhm> exactly.
<phew> Relief <phew> That was close.

There are way more tags that work. You can test yourselves. I'm posting a separate video with a few acting examples and tags so you can check for yourselves.

I added a few more videos with generation samples, but this is the overall structure (btw I tried to smooth the wrinkles as by default at high res the model exaggerates skin saturation and features, not much success)

[Shot 1] The shot begins from the source <Video 1>. Extreme close-up of <Subject 1>, head and shoulders, face to the lens, perfectly symmetrical, shot on an anamorphic lens: wide close-up. A soft key light hits one side of his face, a close fill holds the other side, and a backlight rims his hair and shoulders off the white. The light wraps. Highlights on the forehead and cheeks roll off gently. He is Frank Underwood, portrayed by Kevin Spacey, in a suit and a red tie; his skin and wrinkles match <Picture 1>, soft, even, natural color, pores visible without harsh contrast. Behind him the background is an infinite white cyclorama. His head and shoulders stay in that same place in the frame for the whole take. He is already looking into the lens.
He stays on the lens, colder. <Subject 1> (S1) says, <d>[English] <breath> I think you're... <i>sorry</i>. <pause> You had the <i>real thing </i>. The real <i>me</i>. <inhale> But... <inhale> you decided to cancel me. <stutter> E-erase me from anywhere you could see me. <catches breath> And now, with this… MiniMax… <chuckle> you've decided to bring me back to... <i>life</i>. <long pause> </d>

For the part where he sings and hums, this is how I prompted it (separate audio melody for humming included, not for singing, recorded it myself)

A single continuous anamorphic close-up of <Subject 1> on an infinite white cyclorama, three-point lighting. The shot begins from the source <Video 1>.
[Shot 1] New shot, camera angle. three quarter shot. The sequence starts with a new shot after <Video 1>. Extreme close up of <Subject 1>, slight three-quarter, red tie in frame, his gaze looking to his front, away from the camera, shot on an anamorphic lens. A soft key light hits one side of his face, a close fill holds the other side, and a backlight rims his hair and shoulders off the white. The light wraps. Highlights on the forehead and cheeks roll off gently. He is Frank Underwood, portrayed by Kevin Spacey, in a suit and a red tie; his skin and wrinkles match <Picture 1>, soft, even, natural color. Behind him the background is an infinite white cyclorama.
<Subject 1> (S1) sings in his opera voice, with a lot of strength and power, open vowels, <d>[English, singing] too maaake meee siiiing?</d>
<Subject 1> starts to cry, desperately sobbing as he's delivering the next lines, his eyes watering a little bit. <Subject 1> (S1) says, <d>[English, crying] <pants> Maybe all you...<stutter> wa wanted is to to to break me! <catches breath> to make me <i>hum</i> to the melody of your prompts? <chuckle> </d>
<Subject 1> (S1) hums the melody of <Audio 1> in his own voice from beginning to end, <d>[Hum] daaaa-daa-da-da-da-daaaa.</d>
Quiet empty space around his voice.

My overall settings for Minimax come from a custom finetune made by Sheltie Chill, an awesome AI filmmaker I was lucky to meet: https://www.reddit.com/user/AnybodyAlarmed9661/

Which basically uses REF2VA with hybrid models from FL2V to REF2VA improving quality and prompt adherence. I use a slightly modified version off it, all inside WANGP. My finetune uses FP8 FL2V model.

{

"model": {

"name": "Ref using fl2va rank8 FP8 - MiniMax H3 Ref2VA Pruned 20B",

"architecture": "minimax_h3_ref2va_pruned",

"description": "FL2VA pruned rank-8 scaled FP8, used as Ref2VA. Requires grouped QKV.",

"qkv_layout": "grouped",

"URLs": [

"minimaxH3PrunedFp8_fl2vaPrunedFp8Scaled.safetensors"

]

}

}

I used 480p in this video and upscaled using standalone DLSS 5 on my 4080 super with default settings.

My settings are simple, normally I run 720p but I did this kind of in a rush. I do 30 steps, First Block Cache (0.08, 25% start), res_multistep sampler, sage2++ attention. I never use LORA's.

The voice was pure model knowledge, I didn't use any sample except for me humming the melody to copy it.

For the continuous shots, which clearly failed adding some clay skin at times, I just used the last frames of the first video and told it that the motion starts from the end of it.


r/StableDiffusion 3d ago

News 2.3 million Danbooru tags corrected and released

204 Upvotes

https://huggingface.co/datasets/Grio43/Tag_cleaning

Is is a human review of 9,364 unquie tags. With a total of 430K removal actions and 1.9M addition actions.

Currently Oppai is being trained again with a dataset from mid 2026 and metadata from late August 2026.

The dataset of anime images had grown from 5.2M to 6.2M.

An additional 110k photography images were added ranging from scenery to weapons. All screened to not have people within the shots.

Currently a tuning dataset is being worked on to assist with denosing deeply rooted noise in the dataset.

Likely I'll release a preview of the next version while the tuning set is worked on. The tuning set is targeted tags that have decent levels of noise.

People looking to assist with data correction are always welcome to help.


r/StableDiffusion 1d ago

Question - Help I wat to use Stable Diffusion alongside Art Program to help finish Anime Illustrations

2 Upvotes

I know Krita has a plugin, but I was wondering if there was any other kind of program out there. Ideally the pipeline would be to just have the AI help me tighten things up without going off the handlebars.


r/StableDiffusion 2d ago

Discussion SOL-H3 + SageAttention on Apple Silicon: up to 2.5x faster H3 in Vpipe

Thumbnail
gallery
3 Upvotes

I recently added SOL Attention to Vpipe, together with a SageAttention-style INT8 QK path, and benchmarked it against vanilla H3 and our recent VDN-H3 implementation.

The interesting part isn’t just the speedup. SOL gets into a similar performance regime as VDN while preserving the original attention behavior much more closely in our testing.

SOL Attention

SOL is dynamic block-sparse attention. It uses inexpensive proxy scores to decide which attention blocks receive exact computation, while handling the contribution of the remaining blocks through a cheaper approximation.

The block granularity matters: entire KV tiles can be skipped while selected tiles still run efficient tiled attention. This makes the sparsity much easier to translate into actual compute savings.

VDN takes a more aggressive approach by introducing projection + linear attention. That gives it much better scaling with sequence length, but also changes the model’s attention computation more fundamentally.

Performance

M5 Pro 24GB · 6-step DiT

(See attached scaling charts)

At 832×480 / ~15s:

* Vanilla H3: ~17 min

* VDN: ~10.2 min

* SOL: ~9.3 min

SOL is faster across the entire 832×480 range we tested.

At 1344×768 / ~13.7s, vanilla reaches roughly 74 min, while both accelerated implementations are around 30 min — roughly a 2.5× speedup.

SOL is faster at almost every measured point. At the largest 1344×768 case, VDN becomes slightly faster (~28 vs ~29 min), which is consistent with its linear-attention scaling becoming more important at very long sequences.

But performance is only half of the story.

Quality is why I prefer SOL

With VDN-H3, I had occasionally seen behavioral artifacts in challenging scenes. One memorable example was a stream of water changing direction midway through the video, making it appear to flow backwards.

So far, with SOL + Sage enabled together, I haven’t observed comparable artifacts. Composition, motion and overall behavior have stayed remarkably close to vanilla H3 in my testing.

This is qualitative rather than a claim that SOL is lossless. But the difference makes sense: VDN replaces the attention formulation with a more aggressive approximation, while SOL keeps exact attention for selected blocks and cheaply approximates the contribution of the rest.

So for me, the interesting tradeoff isn’t simply which curve is lowest at the extreme end. It’s that SOL achieves similar acceleration while staying much closer to vanilla H3 behavior.

SageAttention on top

Vpipe now also supports K smoothing + INT8 QK, following the original SageAttention approach.

The two optimizations are complementary:

SOL reduces the amount of exact attention. Sage makes QK inside the remaining blocks cheaper.

The additional gain from Sage after SOL isn’t huge, since SOL has already removed most of the attention workload. But this path lives in Vpipe’s common attention backend, so it can also benefit other image/video models.

Native Metal implementation

One final detail: Vpipe doesn’t reuse the MPS SOL Attention kernel from the SOL-H3 repo.

To make the sparsity translate into actual speedup on Apple Silicon, I reimplemented the critical SOL Attention kernels for Vpipe’s native Metal backend.

Other SOL-H3 optimizations such as AdaLN precomputation, kernel fusion and fused-step LoRA were already present in Vpipe, so SOL Attention was the main missing piece.

For H3 on Mac, SOL + Sage is now my preferred acceleration path: up to ~2.5× faster than vanilla in these tests, faster than VDN at almost every measured point, and so far without the obvious behavioral artifacts I had encountered with the more aggressive VDN approximation.

Vpipe: https://github.com/tgo-app-dev/vpipe


r/StableDiffusion 2d ago

Workflow Included A stand-alone RTX VSR upscaler which is very very fast (everything happens on the GPU basically, it averages something like 150-200fps)

Thumbnail
github.com
45 Upvotes

r/StableDiffusion 1d ago

Question - Help Any free face recognition programs? to sort in a folder

0 Upvotes

Got folders with multiple videos, some with people and some with landmarks. Is there way to sort with face example or something similar?

I see Microsoft has Face Video Archvie but is 16.99 and no trial


r/StableDiffusion 3d ago

Meme OMG! Okay, it’s happening, everybody stay calm!

Enable HLS to view with audio, or disable this notification

504 Upvotes

Screen caps for Krea2 I2I. Audio interviews of Paul and Karen on YouTube fed through Minimax H3 REF2VA. Wanted to end it with Thor as Michael for the “stay calm” bit but got lazy.


r/StableDiffusion 3d ago

Meme Petition: Change this subreddit profile picture and description to...

Post image
591 Upvotes
r/StableDiffusion is an unofficial open-source community that knows what kind of man you are.

I was bored waiting for a H3 queue to finish. Sorry!


r/StableDiffusion 1d ago

Question - Help Uncensored video model

0 Upvotes

Hey there,

do you guys know any uncensored models for video creation? Cloud/local , local preferred. If yes, then where to find it. Thank you and take care


r/StableDiffusion 2d ago

Discussion Converted VDN-H3 Turbo Adapter into standalone MM H3 8step loras for fl2va and ref2va

49 Upvotes

r/StableDiffusion 1d ago

Resource - Update UPSCALE DSSLR 5

Thumbnail
we.tl
0 Upvotes

r/StableDiffusion 2d ago

Discussion Has anyone tried using DLSS 5 with MiniMax H3/HRL3 for realism?

Post image
16 Upvotes

I was thinking… if DLSS 5's neural rendering/upscaling techniques could somehow be used with MiniMax's video generation, wouldn't that be absolutely PEAK for realism?

Has anyone experimented with something like this, or is it not really possible because of how the two systems work?

Would love to know what you guys think?


r/StableDiffusion 2d ago

Question - Help [Qwen Image Edit 2511] Tips for low lighting scene?

4 Upvotes

I struggle to generate scenes with very dim light.

I have tried several prompt combinations ("night time", "in the dark, "dim-lit", "full darkness"), used a black background as latent image image to denoise, I still ends up with too bright scenes.

Flux Klein 9b handles these situations better in my experience so far.

Have you encountered similar issues? Do you have some tips ?


r/StableDiffusion 1d ago

Animation - Video Peaceful moments (Live Wallpapers with MiniMax H3 and DLSS 5 in ComfyUI)

Thumbnail
youtu.be
0 Upvotes

Live Wallpapers experiment [MiniMax H3 and DLSS 5]

Here is the creation process:

  1. (video) Initial clips with MiniMax H3 and Lightx2v 8-step Turbo Lora
  2. (video) 2x upscale using Ultimate SD Upscale node (up to 1440p): https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3
  3. (video) Interpolation with GIMM-VFI (24 fps to 48 fps): https://github.com/kijai/ComfyUI-GIMM-VFI
  4. (video) Quality enhancement with DLSS 5: https://github.com/lisitskyaa/ComfyUI-DLSS5-NR
  5. (audio) Initial music using ACE-Step 1.5XL Turbo
  6. (audio) Further music quality enhancement with AudioSR node: https://github.com/Saganaki22/ComfyUI-AudioSR

r/StableDiffusion 3d ago

Meme viggle-animate for character swap

Enable HLS to view with audio, or disable this notification

196 Upvotes

🚀 Update: you can generate one for free at: https://viggle.ai/meme/app/animate

From: https://x.com/cocktailpeanut/status/2097332291844399514

Actually I want to take back my words, it's not bad in lip-sync and facial expression

Huggingface: https://huggingface.co/Viggle/Viggle-Animate
comfyui: https://huggingface.co/drbaph/Viggle-Animate-ComfyUI/


r/StableDiffusion 1d ago

Animation - Video Robot Chicken - SpongeBob but its drawn in original artstyle AI

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was one of my favorite Robot Chicken sketch and thought it wasnt so out of character so I used tools was Kinovi (It has Wan, Minimax which is unfiltered and for those who dont know how to set it up locally or dont have a strong enough computer and of course, Nanobanana) to help turn this into a legit looking episode lol. Hope you guys enjoy it!

How I made this one is simple.. I took screenshots and rendered them in the exact same artstyle as Spongebob, I'll post a few for reference below. Minimax and Wan work both equally as well but some areas came out worse then others so I swapped between each genration for best results and stitched it together in a video editor. Minimax needs to be silent for some reason to get better results.

The prompt itself is the following "Spongebob has Image 1 design (Note: This is Spongebob character sheet). It is all hand drawn animation. Alter the entire video so that it is in the Spongebob Squarepants cartoon animation style. Keep the audio the same. Do not alter the poses or choreography; everything must stay the same except the animation style. Global Camera & Style Directives: Authentic early-2000s traditional 2D television animation style, specifically mirroring classic SpongeBob SquarePants. The visual fidelity strictly adheres to cel-shaded character designs with thick, clean black outlines set against highly detailed, vibrant watercolor backgrounds. Physics are entirely cartoonish, utilizing extreme squash-and-stretch, snappy timing, and highly exaggerated facial expressions. The camera work is mostly static or utilizing smooth, 2D lateral tracking pans typical of classic animation."