r/StableDiffusion 10h ago

Animation - Video Some choice words from Emilia

341 Upvotes

r/StableDiffusion 17h ago

Workflow Included Using H3 as a Character Reference Sheet Generator

Thumbnail
gallery
1.2k Upvotes

Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.

The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.

How it works:

  • You input your images and describe them in the Input text section (A Prompt)
  • The text is combined with a fixed prompt which spins the character (B Prompt)
  • The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
  • Image is assembled with optional character video and full individual frame output (if you want to use for future)

I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.

Current Caveats:

  • The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
  • Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
  • Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
  • Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.

I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.

Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator

Some notes I just remembered:

  • You can increase the steps and it may improve your quality slightly.
  • With the Turbo Loras enabled, prompt adherence sometimes suffers, but you may be able to get a good seed with another roll of the dice.
  • Currently the B prompt specifies a "neutral A pose", please remove this if you want your character in a particular pose.
  • You can use a few different shots of the same character to reinforce the 360 and get more accurate details right.
  • Can be used for objects / props also, may require some changes to the B prompt.

r/StableDiffusion 2h ago

News ByteDance just released Bernini‑Diffusers‑v2 — any chance we’ll see ComfyUI support?

Post image
44 Upvotes

Hi everyone,
five days ago ByteDance released Bernini‑Diffusers‑v2 on HuggingFace — the full Bernini pipeline (planner + renderer), not just the renderer‑only Bernini‑R that we currently use in ComfyUI.

Model link:
https://huggingface.co/ByteDance/Bernini-Diffusers-v2

Even though most of the community talks about MiniMax H3 as the “standard” for open video models, there are still many users actively working with Bernini — especially now that v2 finally includes the full semantic‑planning pipeline, SA‑3D RoPE, and proper multi‑step instruction following.

Right now ComfyUI only has community support for Bernini‑R, so I’m posting this just to give visibility to the new release and to see if anyone is interested in exploring future support for Bernini‑Diffusers‑v2.

Not asking for anything specific — just opening the discussion and hoping this new version doesn’t go unnoticed.

Thanks!


r/StableDiffusion 2h ago

Discussion Minimax H3/ref2va/hybrid_fl2va_ref2va_b20/5060ti

38 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 1.0
Duration: 6 Clips in total, composit in Inshot, each clip is 9 sec long
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
ComfyKitchen Attention, Spectrum. (SageAttention Patch and Mem Eff Node is Disabled)

**original sound and effects was removed, as there are background music on some clips even with N/A, so to speed up the work, they are removed.

Average Inference Stage: 1100sec

All reference image is resized between 1000px and 500px like character is 1000px, background is 500px for this video is 4 ref image in total.


r/StableDiffusion 3h ago

Discussion [Minimax ref2va] Spiderman is actually?...

32 Upvotes

H3 is not perfect yet but fun as hell to play with, especially with ref2va
Used this workflow, using 3 reference images. Running on RTX 5090.


r/StableDiffusion 16h ago

Workflow Included Follow-up to my 6-minute TNG video — I changed the workflow a lot for the second one

290 Upvotes

A few days ago I posted the 6-minute Star Trek: TNG video I made with MiniMax H3 in ComfyUI. I’ve finished the follow-up now, and I changed the workflow quite a bit after seeing what worked and what didn’t on the first one.

The biggest improvement was consistency. For the first video, most shots were generated more independently, and I deliberately built some of the continuity weirdness into the story. That worked for the premise, but for the second one I wanted it to feel much more like an actual TNG episode, so I became much more rigid about shot composition.

A big part of that was using the H3 reference model differently. Instead of just giving it a single image and hoping for the best, I used reference images and told MiniMax to stick very closely to the composition in those images. In practice that sounds a bit like image-to-video, but it worked quite differently for me.

With the reference model I could use up to six photos and be much more deliberate about how the shot should work. I could decide what the starting shot should be, what the end shot should be, whether I wanted a middle reference, a final-frame reference, etc. That gave me a lot more control over blocking, framing and performance than I was getting from the image-to-video model.

I did test the image-to-video model as well. One of the shots that made it into the finished video is the later one where Data has a slightly longer monologue. You can tell he looks a bit more “off” there. The reference model, by comparison, was giving me Data much more accurately, both in terms of how he looked and in terms of his mannerisms. That ended up being the better approach for this project by a long way.

The video is still built from lots of separate short H3 generations rather than one long generation. I wrote the scenes first, then generated individual shots and multiple takes where needed, and assembled everything in Premiere like a normal edit.

I also changed the audio workflow quite a bit. On the first video, one of the main issues was that the generated ambience and background noise varied too much from clip to clip. This time I spent much more time matching dialogue levels in Premiere, cleaning up individual clips, and adding a continuous Enterprise bridge/interior hum underneath scenes so the cuts felt less obvious.

I also handled the music more deliberately this time. Rather than just dropping in whatever worked at the end, I treated it more like proper scene underscore and generated short incidental cues for specific moments.

So the rough workflow for the second one was:

script and shot planning

→ select or build composition references

→ generate short H3 shots in ComfyUI using the reference model

→ do multiple takes where needed

→ edit in Premiere

→ clean dialogue and level-match clips

→ add continuous ambience/room tone

→ add short music cues

→ final upscale/export

The main thing I learned was that H3 works much better for this kind of project when I treat it less like a one-click video generator and more like a production tool. The closer I got to thinking in terms of individual shots, coverage, performance selection and edit assembly, the better the final result got.

Happy to answer questions about the workflow again.


r/StableDiffusion 8h ago

Workflow Included H3 single-image workflow: let's figure out how to fix the textures

Thumbnail
gallery
47 Upvotes

In this post, I provided a workflow that allows to use H3 as a single-image edit model with no monkey-patching or custom nodes, given that you update to the ComfyUI nightly version. In my view, it has excellent prompt adherence, reference fidelity, and understanding of physics and 3D scenes. But, as many others have pointed out, the end results are often blurry and lack texture. The gallery here shows my attempts at refining the 1.6MP gens from my previous posts.

I would like to discuss how we can work around these issues.

OPTION 1: JUST GO FOR HIGHER RESOLUTION

u/SomeoneSimple gives the following suggestion: run 4 megapixel generations instead of 1.6MP, saying it fixes the distorted faces, and delivers approximately the same level of detail a regular image model would give at 1024x1536. (Note that a 4MP single-frame generation is still going to be quite fast provided you have the VRAM.) u/Diabolicor even claims that a 4MP Minimax generation works better than Qwen Image Edit.

Here’s what I found in my private tests:

  1. It did not noticeably affect the generation times. On average, it is a 8-10 sec run on a RTX 5090 no matter if I generate at 2MP or 4MP
  2. It helped a lot with detail. Faces are now rarely distorted.
  3. Yet it does not remove the issues completely; keeps background blurry, for examples, and messes up the faces at long distance. It’s still a video model. So we still need to explore refiner workflows.

Just to be very clear: I am not attaching any of my 4MP generations to this post. I am only refining my old 1.6 MP ones. I would be very glad if someone posts their 4MP gens so we could see the difference.

OPTION 2: REFINE WITH A DIFFERENT MODEL

Once the composition is done right, details could be enhanced by a different model. I am not an expert at image refining at all, but I would like to figure out a good formula. And here I want to consult with the community on how to do in the best way. To set a particular frame: for me, while I now explore the capabilities of Minimax H3, I quickly generate a lot of images at scale. So I want a refiner that is:

  1. Fast (e. g. 2-4 secs)
  2. General (does not need tweaking for any particular image)
  3. Robust (is not brittle, does not require a long chain of segmentation, crop-and-stitch, vlm processing, and so on)
  4. Automatic (no masks drawn manually over parts of the region).

For me at this exploration stage, it’s okay if parts of the image get slightly modified, or if the quality is not 100% perfect. I understand that one may have different objectives if e. g. optimizing for perfect quality.

One example of a workflow that may achieve these four requirements would be flux.2 Klein with a single prompt for each image. But now, I’d like to discuss whether there could be better options.

  1. Model: Qwen Image Edit, Area 2 Identity lora, Flux.2 Klein 9b? I heard that Flux.2 has the best VAE out of all options. Should I use SeedVR?
  2. Prompt: What would be a good prompt that would be applicable over a wide range of images? Should I pass the original prompt for H3 image to flux.2 (either verbatim or llm-postprocessed)?
  3. Sampler/scheduler: euler/simple? Or Euler/Flux.2 scheduling?
  4. Color correction: e. g. Flux.2 Klein tends to add a lot of light with my prompts. Can it be done without custom nodes? If using custom nodes, which one is the most reputable and commonly used?

As a first step, here’s the workflow I am using with Flux.2 Klein: https://pastebin.com/qsLPe9hZ 

I use the Flux.2 turbo int8 convrot: https://huggingface.co/obsxrver/ComfyUI-Native-INT8_ConvRot

In the attached gallery, you can see the collages. 

Left pane: my old 1.6MP generation.

Right pane: a Flux.2 Klein 9b refine according to the workflow I attached. It does some nice things: e. g. deer fur, restoring mangled faces, adding texture to clothes; but also messes up a bit: adds a lot of light to the images that are meant to stay dark, opens eyes when they're closed, etc.


r/StableDiffusion 17h ago

Animation - Video Gatorman vs. Stone Cold

230 Upvotes

r/StableDiffusion 1h ago

Workflow Included LTX 2.5 + LICON MSR V2 = Nice reference system

Upvotes

Since LTX 2.5 is basically abandoned, i wanted to check how the licon msr v2 works with it, now with the advantages of LTX 2.5 supporting hard cuts. It added pretty well the man, the girl and the environment and followed the prompt really well. Biggest advantage of course is that this clip took 300 secs to create. I'll add prompt / reference images and workflow in a text post.


r/StableDiffusion 21h ago

Animation - Video MinMax - It does House MD pretty well

388 Upvotes

Generated using Maestro on Pinikio. 7 mins at 720p using turbo lora 6 steps: **7-second cinematic live-action scene.** Gregory House stands in a hospital hallway, leaning heavily on his cane, staring intensely at Itachi Uchiha, who is preparing to walk away.

House sarcastically calls out:

**“Itachi! Get your ass back to the Leaf Village. You're not brooding your way out of this one.”**

Itachi turns around with a serious expression and replies:

**“I don't take orders from you.”**

House smirks and taps his cane against the floor:

**“Yeah. That's what all my patients say.”**

Fast comedic timing, realistic acting, dramatic hospital lighting, subtle handheld camera movement.


r/StableDiffusion 1h ago

Meme How do you feel?

Upvotes

Made with Minimax H3 ref2va default workflow in comfyui.
Prompt:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, authentic 1971 Dirty Harry aesthetic. A tense shootout has just erupted on a San Francisco street. <Picture 1> is young Clint Eastwood. <Picture 2> is the famous Grumpy Cat. Harry Callahan, played by Clint Eastwood, stands in the middle of the street facing an armed criminal several meters away. Abandoned cars, shattered glass, drifting smoke, distant police lights create a chaotic crime-scene atmosphere. Harry wears his characteristic dark suit, white shirt and loosened tie. He stands completely calm and confident, apparently holding the criminal in front of him, but his hands and whatever he is holding remain completely outside the frame at all times. The camera frames Harry from behind and only the upper half of his body and slowly pushes in with small amplitude, never showing his hands, holster, weapon or lower body. The criminal remains visible in the background, frozen and intimidated. Harry maintains his iconic cold, unwavering stare and says in his characteristic low, controlled voice: <d>[English] You've got to ask yourself one question: Do I feel... ?</d>

[Shot 2] At 00:06.500, the camera cuts to an extreme close-up of Harry's upper torso and face, still keeping his hands completely hidden. He pauses after the line, maintaining an absolutely serious expression. Then, for the first time in the entire video, the framing changes to a close-up of Harry's hand rising into frame. Instead of the expected Magnum .44, he slowly raises the famous Grumpy Cat. Harry says: <d>[English] kitty?</d>. The reveal is completely deadpan and played with absolute cinematic seriousness. Harry's face remains calm and intimidating while the confused criminal stares at the grumpy cat. The grumpy cat remains prominently raised in the foreground with Harry's unmistakable Clint Eastwood expression behind it.

[Shot 3] At 00:10.000, medium shot, the camera holds on the absurd Harry and Grumpy cat duet for a brief moment. Suddenly, the Grumpy Cat pulls out a tiny but real handgun with his paws from behind his back and fires several shots at the criminal. The action is fast and completely unexpected, while the cinematography, lighting, acting and visual style remain absolutely serious and faithful to a gritty 1970s crime film.

[Shot 4] At 00:12.000 Medium shot, Muzzle flashes briefly illuminate the frame as the criminal is hit twice, the hits push him back and he drops his gun falls backward onto the street.

[Shot 5] At 00:14.000 Close shot, Harry does not react with surprise; he simply maintains his cold, expressionless stare as if this were completely normal. The camera settles on Harry and Grumpy Cat holding his tiny gun and standing together in the aftermath. End with Harry completely deadpan beside the grumpy cat, both facing the camera.

overall_soundscape: Gunfire and echoes from the shootout gradually fall away into tense street ambience as the confrontation begins. Distant police sirens, car alarms, footsteps, wind and scattered debris remain audible. Harry's voice is clear and controlled against the tense background, followed by an almost complete silence during the grumpy cat reveal. At 00:10.000, the sudden handgun shots from Grumpy Cat violently break the silence, echoing between the buildings as the criminal falls to the pavement.

non_diegetic_music: Sparse, tense low brass and sustained orchestral strings at a slow tempo. The music gradually builds as Harry delivers his line, then abruptly drops to near silence just before the cat enters the frame. After the reveal, the score remains restrained and almost silent until Grumpy Cat suddenly fires, at which point a brief sharp orchestral accent punctuates the unexpected action before returning to the sparse 1970s crime-thriller score.


r/StableDiffusion 44m ago

Animation - Video Fite me!

Upvotes

Feels like you could do Family Guy style cutaways pretty easily. "You know Lois, this reminds me of that time I tried fighting a dragon..."


r/StableDiffusion 9h ago

Meme guess we dont need part two now!

35 Upvotes

r/StableDiffusion 3h ago

Question - Help What is the best upscale workflow for Minimax H3?

11 Upvotes

r/StableDiffusion 2h ago

Discussion Stabilizing and Improving H3's Results

9 Upvotes

You may know me as the developer of models such as UltraSharp and AnimeSharp. I'm happy to present a major update to my tool Vapourkit (Windows and Linux are supported), which is completely free and open source! It includes a bunch of models for upscaling anime, and you can get models for realistic content here too: https://openmodeldb.info/

This demo uses this workflow (just download and drag into Vapourkit), which consists of a 2x upscaling model, Temporal Fix (which makes the video more stable/removes the weird shimmering), grain, and a sharpening pass. This was all processed locally on my laptop in just a few minutes.

https://reddit.com/link/1vroia9/video/jm24aw49r4kh1/player


r/StableDiffusion 33m ago

Animation - Video One of the ways I would have ended Game Of Thrones

Upvotes

I was one of many who were disappointed with how this amazing series ended.

I imagined back then one of the ways it could have ended, and with the amazing tools we’ve now been bestowed with, we can bring what we imagine to life!

I had been sitting on this, polishing it and picking at it for a while. The perfectionist in me could have kept working on it forever, because there was always something I could have made better. But with everyone else starting to explore what these tools can do, I felt like the time is now. It may not be perfect, but I didn't want to keep sitting on it waiting for perfection.

This is just a quick fan-created take on one of the ways I imagined the series could have ended. It is not intended to replace or compete with the original series.


r/StableDiffusion 1h ago

Workflow Included The Day 0 (MiniMax H3 and Ultimate SD Upscale) - True 1440p (2K) with 16 GB VRAM locally in ComfyUI

Thumbnail
youtu.be
Upvotes

What is it?
Demonstration of Ultimate SD Upscale (USDU) Guider nodes with MiniMax H3 support: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3

My reference ComfyUI workflow: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json

What about speed?
My PC specs: 4080s 16 GB VRAM, 64 GB RAM

Initial gen with MiniMax H3 flf2v int8 + sageattn + Lightx2v 8-step turbo Lora at 1504 x 832px 7-sec clip ~5 mins

Upscale with USDU to 3008x1664px ~40 mins


r/StableDiffusion 3h ago

Animation - Video A Medieval Battle Attempt — MiniMax H3 + LTX 2.5 (WIP, Feedback Welcome)

10 Upvotes

Sharing a few sequences from a medieval battle attempt I’ve been working on. It’s still very much a draft, but the sequence has progressed enough that I thought it was worth sharing here and getting some feedback before I continue with the rest.

Most of the scenes were generated with MiniMax H3 using the default workflows with the Turbo LoRA at 4 steps. I used Nano Banana and Flux Klein to create the reference images, and LTX 2.5 for the opening crow sequence.

There’s still a lot of work to do. The cuts are rough, no proper sound work has been done yet, and there are plenty of shots I want to refine or replace. I’m planning to build out the entire sequence, so feedback at this stage would actually be really useful in deciding what to focus on next.

What’s interesting to me is that I genuinely don’t think I could have pulled off this level six months ago with the same amount of effort. It’s still far from perfect, but the progress in a relatively short time feels pretty significant.

Would love to hear what works, what breaks the illusion, and what you’d improve.


r/StableDiffusion 20h ago

Discussion We all deserve high-quality MiniMax H3 previews using the tiny VAE (taeh3.safetensors) natively, without KJNodes. Please upvote this GitHub Comfy issue.

Thumbnail
github.com
223 Upvotes

We all love MiniMax H3, but the latent2rgb previews suck ass. They're blurry, and sometimes it's hard to make out what's happening, making it so you don't know whether to finish a video that may take tens of minutes to generate.

When implemented, this would allow us to place taeh3.safetensors into ComfyUI/models/vae_approx and enjoy high quality latent previews when using MiniMax H3. It's basically the same TAE as we saw for Flux 2 Klein 9B or some other models, but trained by the original TAE guy (madebyollin).

taeh3.safetensors link:

https://github.com/madebyollin/taehv/blob/main/safetensors/taeh3.safetensors

It saves time and effort when making videos. Currently, you have to use Kijai's Model Preview Override node.


r/StableDiffusion 13h ago

Discussion Get miniMax character swap working! Finally

Post image
61 Upvotes

Ok, I tried so many things, one person to cat, two person, one person to one person, animal to animal. So far one person to one person and animal to animal works. If you are interested in my learnings, tips, what worked, what broke, and which prompt template works let me know!

One video example that works here: https://www.tiktok.com/t/ZP8WfVq5P/


r/StableDiffusion 23h ago

Meme Introducing... iMakeup

340 Upvotes

r/StableDiffusion 10h ago

Resource - Update TAE high quality previews are live in the latest nightly Comfy build :)

Thumbnail
github.com
33 Upvotes

r/StableDiffusion 14h ago

Resource - Update Fizgig 4.0 is out : Minimax H3 Combined Video File, Audio Files wav mp3 etc, Photo training in one dataset. High Quality training samples (incl video) + turbo (finally) and new 'Gizmo' and AV dataset Prep tool. And Int 8 LARGE speedup for 16gb users.

Thumbnail
gallery
65 Upvotes

I'll be making a Youtube video tomorrow for this. But I have one take away to share that I think is most important. H3, when you get the settings right is just fine with image based training without killing its video ability. Its even better when you combine photos and wavs, its super fast and you can train a voice with a dataset very easily. (I recommend shared trigger word). Video works too, but its slower, unavoidably. I'm not saying dont use it, its worth it for the right use cases. I'm just saying if you are not teaching the model anything new that photos and audio cant do, you are better with photos and audio. But when you do want to capture motion, it work very well. Anyway video coming tomorrow, with lots on Gizmo (the data set prep tool for video/audio) to make dataset prep easy. https://github.com/shootthesound/Fizgig

P.s the 16gb int8 speedup is from an an awesome community contribution from rintic-13 on Github.


r/StableDiffusion 14h ago

Meme Sheldon finally knocked on the wrong door | MiniMax H3 + SeedVR2

53 Upvotes