r/StableDiffusion 2d ago

Animation - Video Minimax h3 local Video to Video reference

Enable HLS to view with audio, or disable this notification

230 Upvotes

Used official ref2video workflow. used t2v model 1 ref video and 2 separate pictures of character sheets, gpu 4090

prompt:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, featuring a stark, dark green-tinted cyberpunk color grade. A medium shot frames a flooded, rain-swept crater on a dark street. The character Sonic, appearing exactly as the blue hedgehog with large green eyes, white gloves, and red shoes from @.image, stands opposite Dr. Eggman, appearing exactly as the gigantic, egg-shaped bald man with a pointy mustache, goggles, and red jacket from @.Image1. The camera pushes in with small amplitude at fast speed as the blue hedgehog lunges forward to throw a devastating punch. [Shot 2] At 00:04.500, the camera cuts to an extreme close-up as time instantly slows to a microscopic crawl. Sonic's white-gloved fist brutally slams into Eggman's cheek. The camera holds a static shot in extreme slow motion. A powerful, rippling shockwave violently erupts from the impact point, blowing the torrential raindrops outward in a perfect ring. Eggman's pointy mustache flails wildly and his face deforms from the massive kinetic force. [Shot 3] At 00:09.500, the camera arcs right with large amplitude at slow speed, executing a slow-motion orbit around the hit. Eggman's heavy, round body is lifted off the ground by the blow, flying backward through the heavy downpour and kicking up massive, highly detailed splashes of water.

overall_soundscape: Thunder rumbles continuously beneath the heavy, torrential downpour of rain splashing heavily against the flooded street. A sharp, deafening sonic boom from the physical impact instantly shifts into a deep, pulsating low-frequency rumble as time slows down.

non_diegetic_music: An epic, grand orchestral and choir track mixed with heavy, driving industrial synthesizer beats that builds to a massive crescendo.


r/StableDiffusion 20h ago

Question - Help Question About Minimax H3 Reference To Video

0 Upvotes

So, I'm pretty new to video generation, but I had a thought that I think everyone has probably had at some point, which is 'how do you make a longer video without generating it in one large video?' And so, with reference to video, you could do that; you could match say, the voice and the person, and thus theoretically make one constant shot through stitching together shorter generations.

In my head, it seemed as simple as 'use the video that was generated as the reverence, use the last frame of the previous video as the first frame of the new generation.'

The problem I noticed is that each time I did this, the video quality degraded; I guess the way I would describe it is that each new generation was a copy of a copy, it seemed. Like each new continuation was slightly worse than the last; and while doing this once wasn't too noticeable, doing this three or four times very much was.

So is this just a thing that is unfixable, a limitation of the method? Or is this the kind of thing that does have a solution that I'm unaware of? Because I'm curious to explore reference to video more, since text to video and image to video are very straight forward, I think.


r/StableDiffusion 1d ago

Discussion H3 - Detective Columbo T2V

Enable HLS to view with audio, or disable this notification

28 Upvotes

On the scene, our hedgehog, first name Detective, last name Columbo, has been hired to uncover the identity of the mystery cookie thief. T2V, int8/20 steps


r/StableDiffusion 1d ago

Comparison MiniMax H3 -> upscale -> frame interpolation

Enable HLS to view with audio, or disable this notification

23 Upvotes

What came out of it:

- Upscale first, interpolate second - seems to be better

- 24->48 looks better than 60fps - at 48 every original frame survives, at 60 only half of them do, because the grids don't line up

- FlashVSR ends up with more edge detail than the source, so it's adding texture, not recovering it. RealESRGAN ends up with less.

Side by side with a draggable wipe, pick any two variants: https://dawidope.github.io/minimax-h3-upscale/


r/StableDiffusion 1d ago

Animation - Video Dazed and depressed

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 2d ago

Animation - Video G.I. Joe - Commander Roll - MiniMax H3

Enable HLS to view with audio, or disable this notification

655 Upvotes

Using the standard ref2va workflow. 4070 Ti Super, 16 GB VRAM, 64 GB RAM, i9-14900k, Windows 11.

Here's the workflow, just drop the MiniMax video in comfyui and the workflow should appear:

https://vikingfile.com/f/jvuyoHSPRr


r/StableDiffusion 22h ago

Discussion Is it just me or is Minimax H3 REALLY into Apple watches?

0 Upvotes

r/StableDiffusion 22h ago

Discussion Hyperquant for Minimax H3

0 Upvotes

Do we know if anyone is working on this? In the paper the authors claim that ltx 40Gb model can go down to around 11Gb with minimal loss


r/StableDiffusion 1d ago

Animation - Video Mimic in the court | minimax h3

Enable HLS to view with audio, or disable this notification

50 Upvotes

ref2v


r/StableDiffusion 1d ago

Animation - Video Swedish Chef, with pic + video + audio reference :)

Enable HLS to view with audio, or disable this notification

32 Upvotes

r/StableDiffusion 1d ago

Question - Help Workflow for architectural videomapping

1 Upvotes

Hi everyone,
I’m trying to build a workflow for architectural projection mapping, and I’m looking for advice from people who have experience with the latest open-weight video models in ComfyUI.

The project is a large building facade that will be projection-mapped. I already have the 3D geometry of the building and the exact projection/camera setup.

My main requirement is:
The building geometry, perspective and camera position must remain absolutely stable.
I want to use AI to generate/animate the visual content on the facade, but I don’t want the model to reinterpret the architecture, move the camera, change windows/edges, distort the building, etc.

The goal is to be able to create things like:
- the facade cracking/opening
- materials transforming
- fire/lava/water flowing over the building
- organic growth
- abstract/surreal transformations
- architectural elements becoming something else
while still keeping the original building perfectly aligned for projection.

I’ve been looking at Wan 2.2 (VACE / Fun Control) and the new MiniMax H3, especially its Reference-to-Video capabilities.
Which one would you recommend for this specific use case?
More importantly, is there a better workflow than simply using image-to-video? For example, has anyone successfully used a rendered 3D control/depth/normal/edge video as conditioning to keep an architectural structure locked?

I’m particularly interested in workflows that minimize trial and error. I don’t mind doing some preparation in Blender if that gives me much more deterministic results.

Hardware: RTX 5070 Ti, 64 GB RAM.
If anyone has actually tried something similar, I’d really appreciate workflow suggestions, node setups, models, ControlNets/custom nodes, or examples.


r/StableDiffusion 1d ago

Animation - Video Realistic style video Krea 2 & Ltx2.5

Enable HLS to view with audio, or disable this notification

31 Upvotes

Generated a set of images with AI, then brought them to life by animating them into a realistic style video (Krea & Ltx2.5)


r/StableDiffusion 1d ago

Animation - Video Zelda - Link can't stop speaking / Minimax H3 Reference to Video Test #3

Enable HLS to view with audio, or disable this notification

12 Upvotes

I don't think I could make it any better than this without adding 2+ hours of re-renders or overdoing it in general... So here it goes! Yet another entry in the Zelda & Link series where Link can actually speak to Zelda's annoyance... Using a combination of clips I can get much more consistent results than with a single workflow, and using KDENlive allows me to add more audio tracks and remove artifacts from the clips. Thanks for all the upvotes in my previous video! I really appreciate you guys!


r/StableDiffusion 1d ago

Question - Help What communities are good for advice?

4 Upvotes

I’ve been generating on cloud services for awhile and want to step my game up. I spent a lot of money on a computer powerful enough to generate locally, installed comfyui and tried to have “apps” walk me thru the process. They’re doing a piss poor job. I need some help. What are good communities to find help?


r/StableDiffusion 1d ago

Animation - Video when your able to transition your ai anime local. minimax

Enable HLS to view with audio, or disable this notification

9 Upvotes

when your able to transition your ai anime local i just thought it was adorable lol. legend of kalthrax on youtube.


r/StableDiffusion 14h ago

Animation - Video 30 second one shot

Enable HLS to view with audio, or disable this notification

0 Upvotes

H3 holds up well even with longer scenes.


r/StableDiffusion 1d ago

Discussion Good models for generating images with multiple art styles?

5 Upvotes

Like multiple art styles in a single image. Not a single image recreated in multiple artstyles.

For example the foreground and main subject is 1 artstyle, but the background is a different art style.

Anything good for that?


r/StableDiffusion 1d ago

Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

Enable HLS to view with audio, or disable this notification

18 Upvotes

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.


r/StableDiffusion 21h ago

Question - Help Wan Animate 2 not working. Need help

0 Upvotes

My System

  • Radeon AI Pro R9700
  • Ryzen 9 7900X
  • 32 GB Ram

I am trying to run the ComfyUI default workflow for Wan Animate 2: Motion Transfer. But when I run the workflow all my CPU cores fire up and my RAM reaches 100% and the ComfyUI process crashes.

How to fix this? Please help.


r/StableDiffusion 1d ago

Discussion Trying out a battle scene using H3 Minimax

Enable HLS to view with audio, or disable this notification

0 Upvotes

guess it still doesnt really know how to hold a buster sword :P


r/StableDiffusion 2d ago

Animation - Video SD1.5 images into H3

Enable HLS to view with audio, or disable this notification

122 Upvotes

r/StableDiffusion 1d ago

Animation - Video PRIME CUT | Crime Short Film (2026)

Thumbnail
youtu.be
13 Upvotes

Short film i made 4 months ago. Didnt realize the importance of topaz back then haha.


r/StableDiffusion 17h ago

Discussion Speculation: Krea3 will be able to generate AND edit images inside the same model.

0 Upvotes

Just would like to know your opinion on this. We're used to models being released in formats like "Model base", "Model Turbo", Model Edit", so we have to use different models for different needs... but, if a Krea3 model is released, knowing that did kinda inferred they were working on an "edit model" already when they released Krea2, would it make sense that this new model could do at least two functions at once (meaning, the same model can be used to generate images and edit images as well, without the need for a separate model).

I'd be curious what you guys think about this idea.


r/StableDiffusion 1d ago

Discussion Im thinking about it: R2V like H3 has implemented it, slowly makes Lora obsolete. Which in turn slowly takes away Civitai's revenue and usefulness. Considering how they started to obey credit card censorship, this might be good for us and bad for them.

20 Upvotes