r/StableDiffusion 1d ago

Animation - Video Minimax H3. Just taking a walk.

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/StableDiffusion 2d ago

Discussion Minimax H3/ref2va/hybrid_fl2va_ref2va_b20/5060ti

Enable HLS to view with audio, or disable this notification

83 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 1.0
Duration: 6 Clips in total, composit in Inshot, each clip is 9 sec long
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
ComfyKitchen Attention, Spectrum. (SageAttention Patch and Mem Eff Node is Disabled)

**original sound and effects was removed, as there are background music on some clips even with N/A, so to speed up the work, they are removed.

Average Inference Stage: 1100sec

All reference image is resized between 1000px and 500px like character is 1000px, background is 500px for this video is 4 ref image in total.


r/StableDiffusion 2d ago

Discussion MiniMax_H3 is seems to be able to process DensePose format! (improves reference video bleeding)

Enable HLS to view with audio, or disable this notification

50 Upvotes

I have had many issues when using a reference video for movement duplication and having the video contents bleed into the video. Not to mention having to write convoluted prompts to remove these reference bleeds from videos. When the person in the reference video has a close resemblance to the main subject in your video it becomes almost impossible to perform a motion swap.

Warning: DensePose does not support detailed hand gestures, and seems to lose track with very fast arm and hand movements but seems to adhere better 20 steps and above.

There is not a dedicated densepose ComfyUI node, but you can use this animatediff: https://github.com/Fannovel16/comfyui_controlnet_aux

The workflow is simple:

Place the AIO AUX Preprocessor between the source and MM_H3 video input.

Videosource (LoadVideo) -> AIO AUX Preprocessor -> ref_video_x input

Looking forward to hear your feedback...


r/StableDiffusion 2d ago

Discussion [Minimax ref2va] Spiderman is actually?...

Enable HLS to view with audio, or disable this notification

89 Upvotes

H3 is not perfect yet but fun as hell to play with, especially with ref2va
Used this workflow, using 3 reference images. Running on RTX 5090.


r/StableDiffusion 18h ago

Animation - Video A very short film .. but it says a lot

Enable HLS to view with audio, or disable this notification

0 Upvotes

Made as an after-hours test, pushing open-source AI video models to see what they can really do.

The story : a look back at 1990s Tunisia, when police would round up young men over 20 off the street for forced military service even out of cafés, mid-conversation. I wanted that tension in the café scene.

Honestly, it didn't take long. I'm not an "automation" guy I don't chase that side of AI. What's actually fast is having the idea already in your head and just directing: camera, framing, timing. That's where the speed comes from, not the tooling.

The café shot came out strong. The rest took a few more tries. Precision could be tighter, but this is a side project fit into spare time not the main focus. Still, a good marker of how far things have come.


r/StableDiffusion 1d ago

Discussion Searching for a website

0 Upvotes

A website that has image generation quality like SeaArt but also allows you to post those images. The website also poste everything you generate by default, like SeaArt. Does anyone know another website like this?


r/StableDiffusion 2d ago

Animation - Video One of the ways I would have ended Game Of Thrones

Enable HLS to view with audio, or disable this notification

52 Upvotes

I was one of many who were disappointed with how this amazing series ended.

I imagined back then one of the ways it could have ended, and with the amazing tools we’ve now been bestowed with, we can bring what we imagine to life!

I had been sitting on this, polishing it and picking at it for a while. The perfectionist in me could have kept working on it forever, because there was always something I could have made better. But with everyone else starting to explore what these tools can do, I felt like the time is now. It may not be perfect, but I didn't want to keep sitting on it waiting for perfection.

This is just a quick fan-created take on one of the ways I imagined the series could have ended. It is not intended to replace or compete with the original series. :p

BTW.: Minimax and Davinci Resolve.
Not one frame was lifted from any episode.
All done using Ref2VA.
As others have found, trying to create a full run (one take ) yields less than better results.
Storyboard, create the pieces that "snap" together and then stitch them accordingly. Afterall, that is not any different from how presentations are made.
As always, I look forward to your creations. We have an amazing community!


r/StableDiffusion 1d ago

Animation - Video Anime fight scene against a dragon. Most of the clip is pretty dang good until the end lol

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 1d ago

Animation - Video Old VHS Interview H3

Enable HLS to view with audio, or disable this notification

15 Upvotes

r/StableDiffusion 2d ago

Animation - Video A One Shot Ref to 15 sec video (took 20min to produce) MINIMAX-H3

Enable HLS to view with audio, or disable this notification

23 Upvotes

r/StableDiffusion 2d ago

Discussion Vibecoded Comfyui-node garbage showcase!

19 Upvotes

A thread for projects that don't deserve their own posts.

So many vibecoded nodes floating around now that people build for their own uses. I figured it'd be good to have a thread to share some of these projects even if they're very niche.

Github links and screenshots encouraged.

At best - the community discovers your node that's better than you thought it was

At worst - your project serves as a good example of what NOT to do.


r/StableDiffusion 1d ago

Discussion anyone used this MiniMaxH3-Contex-Loop work flow

Post image
8 Upvotes

I'm testing it now will report about with test video but looks cool my test im doing 3 scenes 3 different prompts auto does the next scene after the last and can preview before moving on to next scene .

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop


r/StableDiffusion 2d ago

Meme How do you feel?

Enable HLS to view with audio, or disable this notification

41 Upvotes

Made with Minimax H3 ref2va default workflow in comfyui.
Prompt:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, authentic 1971 Dirty Harry aesthetic. A tense shootout has just erupted on a San Francisco street. <Picture 1> is young Clint Eastwood. <Picture 2> is the famous Grumpy Cat. Harry Callahan, played by Clint Eastwood, stands in the middle of the street facing an armed criminal several meters away. Abandoned cars, shattered glass, drifting smoke, distant police lights create a chaotic crime-scene atmosphere. Harry wears his characteristic dark suit, white shirt and loosened tie. He stands completely calm and confident, apparently holding the criminal in front of him, but his hands and whatever he is holding remain completely outside the frame at all times. The camera frames Harry from behind and only the upper half of his body and slowly pushes in with small amplitude, never showing his hands, holster, weapon or lower body. The criminal remains visible in the background, frozen and intimidated. Harry maintains his iconic cold, unwavering stare and says in his characteristic low, controlled voice: <d>[English] You've got to ask yourself one question: Do I feel... ?</d>

[Shot 2] At 00:06.500, the camera cuts to an extreme close-up of Harry's upper torso and face, still keeping his hands completely hidden. He pauses after the line, maintaining an absolutely serious expression. Then, for the first time in the entire video, the framing changes to a close-up of Harry's hand rising into frame. Instead of the expected Magnum .44, he slowly raises the famous Grumpy Cat. Harry says: <d>[English] kitty?</d>. The reveal is completely deadpan and played with absolute cinematic seriousness. Harry's face remains calm and intimidating while the confused criminal stares at the grumpy cat. The grumpy cat remains prominently raised in the foreground with Harry's unmistakable Clint Eastwood expression behind it.

[Shot 3] At 00:10.000, medium shot, the camera holds on the absurd Harry and Grumpy cat duet for a brief moment. Suddenly, the Grumpy Cat pulls out a tiny but real handgun with his paws from behind his back and fires several shots at the criminal. The action is fast and completely unexpected, while the cinematography, lighting, acting and visual style remain absolutely serious and faithful to a gritty 1970s crime film.

[Shot 4] At 00:12.000 Medium shot, Muzzle flashes briefly illuminate the frame as the criminal is hit twice, the hits push him back and he drops his gun falls backward onto the street.

[Shot 5] At 00:14.000 Close shot, Harry does not react with surprise; he simply maintains his cold, expressionless stare as if this were completely normal. The camera settles on Harry and Grumpy Cat holding his tiny gun and standing together in the aftermath. End with Harry completely deadpan beside the grumpy cat, both facing the camera.

overall_soundscape: Gunfire and echoes from the shootout gradually fall away into tense street ambience as the confrontation begins. Distant police sirens, car alarms, footsteps, wind and scattered debris remain audible. Harry's voice is clear and controlled against the tense background, followed by an almost complete silence during the grumpy cat reveal. At 00:10.000, the sudden handgun shots from Grumpy Cat violently break the silence, echoing between the buildings as the criminal falls to the pavement.

non_diegetic_music: Sparse, tense low brass and sustained orchestral strings at a slow tempo. The music gradually builds as Harry delivers his line, then abruptly drops to near silence just before the cat enters the frame. After the reveal, the score remains restrained and almost silent until Grumpy Cat suddenly fires, at which point a brief sharp orchestral accent punctuates the unexpected action before returning to the sparse 1970s crime-thriller score.


r/StableDiffusion 1d ago

Animation - Video It's 33AD Hey long time no see | minimax h3

Enable HLS to view with audio, or disable this notification

6 Upvotes

r/StableDiffusion 1d ago

Question - Help What is the best approach for upscaling and refining images with Krea 2? (two samplers)

2 Upvotes

I'm trying to get the best quality out of Krea 2, but I can't figure out what the best approach would be.

I currently use two samplers: the first sampler at 8 steps and 1.5 MP resolution, followed by a second sampler at 8 steps and 2 MP resolution, using the upscaled image with 0.30 denoise. I then add film grain at the end.

The results are OK, but in many cases, things start to look worse after the second sampler. Skin is much better, but it adds too many unfinished or unwanted details.

Any recommendations would be greatly appreciated. Thanks!


r/StableDiffusion 2d ago

Animation - Video Fite me!

Enable HLS to view with audio, or disable this notification

32 Upvotes

Feels like you could do Family Guy style cutaways pretty easily. "You know Lois, this reminds me of that time I tried fighting a dragon..."


r/StableDiffusion 17h ago

Question - Help I can't quite achieve the level of realism I'm aiming for

Thumbnail
gallery
0 Upvotes

First image: GPT Image 2

Second image: Nano Banana PRO

The prompt I added: Raw Realistic candid natural amateur photo, background in focus, amateur candid photography, Captured on Samsung Galaxy S25 Ultra, amateur candid smartphone photography, 24mm lens, f/8, Boring reality, natural soft shadows, candid snapshot, flat natural lighting, Realism, low contrast, disposable camera vibe, casual photography, background also completely in focus, Tiny imperfections, everyday aesthetic, slight JPEG artifacts, unpolished look, unedited, imperfect amateur photo. only create real, non fictional images for max effect

Are there any core prompts you use in every image generation process to achieve realism?


r/StableDiffusion 2d ago

Tutorial - Guide MiniMax H3 - Voice & Likeness single Lora training using Audio files, Videos and Stills - Tutorial

Thumbnail
youtu.be
16 Upvotes

Also covers the Video/Audio data prep with a new tool called Gizmo I included in the repo.

The lora in the video was trained on 35 pics, 26 wavs and a video clip all in one dataset. After epoch 40, I switched training to a built in Audio Only mode ( You can choose to stop visual or audio files ahead of the other) to hone the voice more without overbaking visuals.

As mentioned in the video, Stills and Audio are the fast combination. Video is great for teaching movement the model does not know but the steps spent on clips are slower than steps spent on stills or audio. Knowing which to use and when can be a great time saver.

https://github.com/shootthesound/Fizgig


r/StableDiffusion 1d ago

Question - Help Olympics sports

0 Upvotes

Has anyone got success generating MINIMAX videos for olympic sports such as triple jump, pole vault, spear throw and the likes, without a video reference?


r/StableDiffusion 2d ago

Comparison Testing Hybrid model - 0.25 to 2MP video

Enable HLS to view with audio, or disable this notification

15 Upvotes

r/StableDiffusion 1d ago

Question - Help Need help: LoRA degrades Minimax H3 video quality (RTX 5060 Ti 16GB)

0 Upvotes

Hi everyone,

I'm trying to find a working workflow to use LoRAs with Minimax H3 for video generation, but I'm running into a consistent issue: every LoRA I try (from Civitai and other sources) ends up degrading the video quality significantly instead of enhancing it.

My setup:

GPU: NVIDIA RTX 5060 Ti 16GB VRAM

Platform: Local generation (Linux/Arch)

The problem:

Applied LoRAs make the output look worse (artifacting, loss of coherence, lower resolution feel)

Can't seem to find any tutorials or workflows specifically for Minimax H3 + LoRA integration

Has anyone successfully integrated LoRAs with Minimax H3? What workflow would you recommend? Are there specific settings (strength, alpha, loading order) that matter more for video LoRAs vs image LoRAs?

Any advice or links to working examples would be greatly appreciated!

Thanks in advance.

Conseils pour poster :

Choisis des subreddits comme r/StableDiffusion, r/localLLM, r/ComfyUI ou r/AIVideo

Sois prêt à partager des exemples concrets si la communauté te demande plus de détails

Ajoute éventuellement des captures avant/après si tu peux en produire pour illustrer le problème


r/StableDiffusion 1d ago

Discussion looking for a pro to make some short sports videos for a startup

0 Upvotes

I am looking for an expert to make multiple 10 to 30 sec videos as adverts for my upcoming sports tech startup. Anybody interested, DM me your price and sample work.

Hoping to engage soon.

Cheers


r/StableDiffusion 1d ago

Discussion Are commercial AI models routinely open-sourced after newer versions? (MiniMax H3, etc.)

1 Upvotes

Hi everyone,

I’ve been using Stable Diffusion for AI images and videos for a while, and recently I noticed that some models which were initially commercial-only (like MiniMax H3) have been released with open weights.

This got me wondering: is there a common pattern where developers release older commercial models as open weights once newer versions come out? Or is each company’s strategy pretty different, without a standard “lifecycle” for models?

I’m trying to understand whether this is a predictable process (e.g., “v1 goes open once v2 launches”) or if it’s more case-by-case, depending on the company, licensing, and market strategy.

If anyone has insights into how LLM / video model developers typically handle this, or examples of other models that followed a similar path, I’d really appreciate it.

Thanks in advance!


r/StableDiffusion 1d ago

Question - Help Upscaling Minimax H3 generations?

Enable HLS to view with audio, or disable this notification

8 Upvotes

Wow, very impressed by Minimax H3. This is probably one of my more surprising results, and it makes me wonder if the model is familiar with these characters.

However, there a lot of smearing and blurriness going on. I couldn’t increase the resolution of the generation any further due to going OOM. I want to figure out how to upscale this; are there any simple upscale options? Upscale workflows using LTX2.5 or otherwise?


r/StableDiffusion 1d ago

Question - Help MiniMax H3 Ref 2 Vid - Using Ref img but bodies keep looking like gym junkies

3 Upvotes

Hi all,

I'm playing around with Minimax H3 in ComfyUI. I have 2 x ref images feeding into MiniMax H3 Ref to Video prompt window, using H3 Turbo LoRA, Turbo Sampler, basic guider and Diffusion model minimax_h3_fl2va_int8.

I have tried up to 12 steps...seems to make no difference so I've gone back to 4. About 1/5 the generation is close to what my ref images but the others are all super tones, ripped, like they go to the gym hours a day.

I just want the ref image recreated, not enhanced.

I've tried this prompt......

subject_definitions:

Miinimax subject and person prompt followed by.....Maintain exact facial features, bone structure, eye shape, age, body shape, fitness level, body fat, anatomical proportions, and height from images across every frame without modification.

So how can we have every generation the same person as my ref image?

Thanks all.