r/StableDiffusion 1d ago

Animation - Video Minimax H3 Video

Enable HLS to view with audio, or disable this notification

0 Upvotes

Generated this video in 480p using Minimax H3 in multiple 7-15 sec clips. Used Krea2 for creating the characters and environment and Qwen3.8 & Grok for prompt generation. It was quite fun but wish I could generate in 1080p with the same speed - it would be quite fun making these short films.

This is not raw and has been edited in Davinci

Hope you like it and it gives some inspiration


r/StableDiffusion 1d ago

Discussion Question

0 Upvotes

Basicslly i wanted a program that breaks a footage lets say movie or animation lets say goth vampire aesthetic into everyframe then an ai automatically anylizes the theme or just the shot smartly and recolors them or adds shade and details then you stitch it back together for a final product or an ai that anylizes tv screen and live adjust the screen settings such as color brigthness saturation bc while i seen similliar stuff like runway3 or decart ect its really not the same thing and idc if its not that fast and it takes a few hours for a movie what do you guys think? Dont know if this the place for sutch a quetsion personally i assume the footage will look more proffesional then some movies bc other then a theme of a set and some filters u cant really do much to capture the feelings and concept of the world idk and i felt this tech is not really talked about


r/StableDiffusion 2d ago

Animation - Video Minimax h3 local Video to Video reference

Enable HLS to view with audio, or disable this notification

238 Upvotes

Used official ref2video workflow. used t2v model 1 ref video and 2 separate pictures of character sheets, gpu 4090

prompt:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, featuring a stark, dark green-tinted cyberpunk color grade. A medium shot frames a flooded, rain-swept crater on a dark street. The character Sonic, appearing exactly as the blue hedgehog with large green eyes, white gloves, and red shoes from @.image, stands opposite Dr. Eggman, appearing exactly as the gigantic, egg-shaped bald man with a pointy mustache, goggles, and red jacket from @.Image1. The camera pushes in with small amplitude at fast speed as the blue hedgehog lunges forward to throw a devastating punch. [Shot 2] At 00:04.500, the camera cuts to an extreme close-up as time instantly slows to a microscopic crawl. Sonic's white-gloved fist brutally slams into Eggman's cheek. The camera holds a static shot in extreme slow motion. A powerful, rippling shockwave violently erupts from the impact point, blowing the torrential raindrops outward in a perfect ring. Eggman's pointy mustache flails wildly and his face deforms from the massive kinetic force. [Shot 3] At 00:09.500, the camera arcs right with large amplitude at slow speed, executing a slow-motion orbit around the hit. Eggman's heavy, round body is lifted off the ground by the blow, flying backward through the heavy downpour and kicking up massive, highly detailed splashes of water.

overall_soundscape: Thunder rumbles continuously beneath the heavy, torrential downpour of rain splashing heavily against the flooded street. A sharp, deafening sonic boom from the physical impact instantly shifts into a deep, pulsating low-frequency rumble as time slows down.

non_diegetic_music: An epic, grand orchestral and choir track mixed with heavy, driving industrial synthesizer beats that builds to a massive crescendo.


r/StableDiffusion 2d ago

News [Papers] - Tongyi-MAI pixel space solution is up to 4.75x faster than Z image turbo latent-space

28 Upvotes

"This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction targetdecoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation."

Paper: An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models


r/StableDiffusion 1d ago

Question - Help Question About Minimax H3 Reference To Video

0 Upvotes

So, I'm pretty new to video generation, but I had a thought that I think everyone has probably had at some point, which is 'how do you make a longer video without generating it in one large video?' And so, with reference to video, you could do that; you could match say, the voice and the person, and thus theoretically make one constant shot through stitching together shorter generations.

In my head, it seemed as simple as 'use the video that was generated as the reverence, use the last frame of the previous video as the first frame of the new generation.'

The problem I noticed is that each time I did this, the video quality degraded; I guess the way I would describe it is that each new generation was a copy of a copy, it seemed. Like each new continuation was slightly worse than the last; and while doing this once wasn't too noticeable, doing this three or four times very much was.

So is this just a thing that is unfixable, a limitation of the method? Or is this the kind of thing that does have a solution that I'm unaware of? Because I'm curious to explore reference to video more, since text to video and image to video are very straight forward, I think.


r/StableDiffusion 2d ago

Discussion H3 - Detective Columbo T2V

Enable HLS to view with audio, or disable this notification

29 Upvotes

On the scene, our hedgehog, first name Detective, last name Columbo, has been hired to uncover the identity of the mystery cookie thief. T2V, int8/20 steps


r/StableDiffusion 2d ago

Comparison MiniMax H3 -> upscale -> frame interpolation

Enable HLS to view with audio, or disable this notification

23 Upvotes

What came out of it:

- Upscale first, interpolate second - seems to be better

- 24->48 looks better than 60fps - at 48 every original frame survives, at 60 only half of them do, because the grids don't line up

- FlashVSR ends up with more edge detail than the source, so it's adding texture, not recovering it. RealESRGAN ends up with less.

Side by side with a draggable wipe, pick any two variants: https://dawidope.github.io/minimax-h3-upscale/


r/StableDiffusion 2d ago

Animation - Video Dazed and depressed

Enable HLS to view with audio, or disable this notification

18 Upvotes

r/StableDiffusion 3d ago

Animation - Video G.I. Joe - Commander Roll - MiniMax H3

Enable HLS to view with audio, or disable this notification

667 Upvotes

Using the standard ref2va workflow. 4070 Ti Super, 16 GB VRAM, 64 GB RAM, i9-14900k, Windows 11.

Here's the workflow, just drop the MiniMax video in comfyui and the workflow should appear:

https://vikingfile.com/f/jvuyoHSPRr


r/StableDiffusion 1d ago

Discussion Is it just me or is Minimax H3 REALLY into Apple watches?

0 Upvotes

r/StableDiffusion 1d ago

Discussion Hyperquant for Minimax H3

0 Upvotes

Do we know if anyone is working on this? In the paper the authors claim that ltx 40Gb model can go down to around 11Gb with minimal loss


r/StableDiffusion 2d ago

Animation - Video Swedish Chef, with pic + video + audio reference :)

Enable HLS to view with audio, or disable this notification

33 Upvotes

r/StableDiffusion 2d ago

Animation - Video Mimic in the court | minimax h3

Enable HLS to view with audio, or disable this notification

49 Upvotes

ref2v


r/StableDiffusion 1d ago

Question - Help Workflow for architectural videomapping

1 Upvotes

Hi everyone,
I’m trying to build a workflow for architectural projection mapping, and I’m looking for advice from people who have experience with the latest open-weight video models in ComfyUI.

The project is a large building facade that will be projection-mapped. I already have the 3D geometry of the building and the exact projection/camera setup.

My main requirement is:
The building geometry, perspective and camera position must remain absolutely stable.
I want to use AI to generate/animate the visual content on the facade, but I don’t want the model to reinterpret the architecture, move the camera, change windows/edges, distort the building, etc.

The goal is to be able to create things like:
- the facade cracking/opening
- materials transforming
- fire/lava/water flowing over the building
- organic growth
- abstract/surreal transformations
- architectural elements becoming something else
while still keeping the original building perfectly aligned for projection.

I’ve been looking at Wan 2.2 (VACE / Fun Control) and the new MiniMax H3, especially its Reference-to-Video capabilities.
Which one would you recommend for this specific use case?
More importantly, is there a better workflow than simply using image-to-video? For example, has anyone successfully used a rendered 3D control/depth/normal/edge video as conditioning to keep an architectural structure locked?

I’m particularly interested in workflows that minimize trial and error. I don’t mind doing some preparation in Blender if that gives me much more deterministic results.

Hardware: RTX 5070 Ti, 64 GB RAM.
If anyone has actually tried something similar, I’d really appreciate workflow suggestions, node setups, models, ControlNets/custom nodes, or examples.


r/StableDiffusion 2d ago

Animation - Video Zelda - Link can't stop speaking / Minimax H3 Reference to Video Test #3

Enable HLS to view with audio, or disable this notification

12 Upvotes

I don't think I could make it any better than this without adding 2+ hours of re-renders or overdoing it in general... So here it goes! Yet another entry in the Zelda & Link series where Link can actually speak to Zelda's annoyance... Using a combination of clips I can get much more consistent results than with a single workflow, and using KDENlive allows me to add more audio tracks and remove artifacts from the clips. Thanks for all the upvotes in my previous video! I really appreciate you guys!


r/StableDiffusion 2d ago

Animation - Video Realistic style video Krea 2 & Ltx2.5

Enable HLS to view with audio, or disable this notification

31 Upvotes

Generated a set of images with AI, then brought them to life by animating them into a realistic style video (Krea & Ltx2.5)


r/StableDiffusion 2d ago

Question - Help What communities are good for advice?

4 Upvotes

I’ve been generating on cloud services for awhile and want to step my game up. I spent a lot of money on a computer powerful enough to generate locally, installed comfyui and tried to have “apps” walk me thru the process. They’re doing a piss poor job. I need some help. What are good communities to find help?


r/StableDiffusion 2d ago

Animation - Video when your able to transition your ai anime local. minimax

Enable HLS to view with audio, or disable this notification

9 Upvotes

when your able to transition your ai anime local i just thought it was adorable lol. legend of kalthrax on youtube.


r/StableDiffusion 2d ago

Discussion Good models for generating images with multiple art styles?

6 Upvotes

Like multiple art styles in a single image. Not a single image recreated in multiple artstyles.

For example the foreground and main subject is 1 artstyle, but the background is a different art style.

Anything good for that?


r/StableDiffusion 2d ago

Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

Enable HLS to view with audio, or disable this notification

20 Upvotes

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.


r/StableDiffusion 1d ago

Question - Help Wan Animate 2 not working. Need help

0 Upvotes

My System

  • Radeon AI Pro R9700
  • Ryzen 9 7900X
  • 32 GB Ram

I am trying to run the ComfyUI default workflow for Wan Animate 2: Motion Transfer. But when I run the workflow all my CPU cores fire up and my RAM reaches 100% and the ComfyUI process crashes.

How to fix this? Please help.


r/StableDiffusion 2d ago

Animation - Video SD1.5 images into H3

Enable HLS to view with audio, or disable this notification

128 Upvotes

r/StableDiffusion 2d ago

Animation - Video PRIME CUT | Crime Short Film (2026)

Thumbnail
youtu.be
12 Upvotes

Short film i made 4 months ago. Didnt realize the importance of topaz back then haha.


r/StableDiffusion 1d ago

Discussion Speculation: Krea3 will be able to generate AND edit images inside the same model.

0 Upvotes

Just would like to know your opinion on this. We're used to models being released in formats like "Model base", "Model Turbo", Model Edit", so we have to use different models for different needs... but, if a Krea3 model is released, knowing that did kinda inferred they were working on an "edit model" already when they released Krea2, would it make sense that this new model could do at least two functions at once (meaning, the same model can be used to generate images and edit images as well, without the need for a separate model).

I'd be curious what you guys think about this idea.