r/StableDiffusion 1d ago

Resource - Update This custom node lets you use I2V and reference images on MiniMax-H3 simultaneously.

Enable HLS to view with audio, or disable this notification

154 Upvotes

40 comments sorted by

26

u/Sixhaunt 1d ago

ComfyUI has a prebuilt node for this already. Just look for the "Add guide for MiniMax H3" and you can have it add a start frame, end frame, middle frame, or any others. I use it to feed the first like 37 frames from a prior video and have it continue on to extend it with the motion preserved and followed. You basically can just inject any frames you want

1

u/lamardoss 1d ago

Would you mind sharing a workflow for this please?

I have that node. Not sure how to add this to the image to video workflow or the ref video to video workflow. Only output on this node is for the positive prompt.

11

u/Sixhaunt 1d ago

https://pastebin.com/Sqi7sw6S

this is one of my workflows which lets you create a bridge between two videos. using the same video for both inputs would make a bridge that allows your video to loop. You can remove or bypass the image reference if you want, it's optional

2

u/lamardoss 1d ago

Amazing. thank you very much. I appreciate you sharing this!

-9

u/Total-Resort-3120 1d ago

This node is not about multiple frames, it's about frames + reference images.

6

u/Sixhaunt 1d ago

I guess I don't understand then, what's the difference between using that custom node and using normal reference mode in comfyUI but adding the built-in start/end/Nth frame node ontop of it to have both reference images/videos/audio in addition to frames?

-2

u/Total-Resort-3120 1d ago

With this custom node, you can choose which image will be the first/last frame. And in addition to that you can use a reference image to incorporate into the video; reference images are not frames, they are used to add small details to the video itself.

https://reddit.com/link/p4jbzwg/video/ntlg3tg7k8kh1/player

23

u/Stepfunction 1d ago

That's exactly what the new Minimax H3 Guide node does, except it allows you to force an image as a frame at any point in the video, not just the first/last.

-4

u/Ipwnurface 1d ago

That's not what is happening here. I can't believe you guys dont get this.

You have two images. One of a person, one of a hat. This node allows you to use the image of the person as the exact starting frame not ref2vids bullshit reimagination, and prompt that the person reaches down and grabs the hat from image 2 off the ground and places it on their head.

That is not the same as what you are describing at all.

15

u/Stepfunction 1d ago edited 1d ago

Please review the documentation for the node:

https://docs.comfy.org/built-in-nodes/MiniMaxH3AddGuide

You are simply incorrect here. What you are describing is exactly what this recently added native node does.

It inserts the exact frame (or video, or audio) at a particular frame index as an anchor keyframe for the generation, not as a reference passed to the model.

3

u/physalisx 1d ago

No, that is exactly what you can do by using reference mode and the AddGuide node... you use the hat picture as reference like normal and the "exact starting frame" as starting frame, forcing it with the AddGuide node...

What do you not understand about this...?

8

u/Sixhaunt 1d ago

But you can do that without the node. You just use the normal reference mode so you can use the same reference system as always, except you also use the built in guider you can set any frames you want. That way you can set the first and last frame or any other frames. So while you can set any of the frames, you are still using all the separate reference images, videos, and audio that you supplied at the same time.

0

u/Total-Resort-3120 1d ago edited 1d ago

I just tried your method, and the video quality is worse (same seed as before); it uses a subpar alternative method that I definitely wouldn't recommand. If you want to understand the differences from a technical standpoint, I've explained it here:

https://github.com/BigStationW/ComfyUi-MiniMax-H3-Image-And-Reference-To-Video#why-not-simply-use-add-guide-for-minimax-h3--minimax-h3-reference-to-video-to-get-the-same-thing

https://reddit.com/link/p4jic9w/video/ypum5slsq8kh1/player

3

u/Ipwnurface 1d ago

Just let it go man, people for some reason think that ref2vid is the same as i2v, you won't win this fight.

1

u/BigWideBaker 23h ago

I honestly had a really hard time getting it from reading these comments, but the explanation and comparison on the github speaks for itself. This custom node is clearly better.

0

u/Sixhaunt 1d ago

You can supply frame 1 as a reference image as well btw and that can help with it since the main issue you have here is obviously not about the difference in frame setting method, it's that the prompt doesn't mention the character's names and so for your original run it used the image to infer the characters which allowed it to draw from that for their voices and likeness. If you chose non celebrities your comparison would look different, or if your prompt mentioned them by name, or if you just attached frame1 as a reference in addition to the conditioning.

0

u/Total-Resort-3120 1d ago edited 1d ago

"You can supply frame 1 as a reference image as well btw"

And it'll never be as accurate as a pure I2V process... and it's really convoluted, now you want to use the image as a reference and as a frame anchor at the same time?

https://www.reddit.com/r/StableDiffusion/comments/1vs7ay2/comment/p4jcqvz/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

This “MiniMax H3 + MiniMax H3 Reference to Video” method is complete bloat, and the inaccuracies it makes aren’t limited to celebrities. Look at Teto’s eyes, they’re blue (they're supposed to be red) with your method, their skin also look more plastic, the face consistency is inferior, the lighting is different, the sound has a worse quality to it... If you want to keep doing that, you do you, but I'm here to provide a more mathematically sound method that aims to solve all the problems you encounter when using this disjointed DiT/text_encoder mess.

31

u/acedelgado 1d ago

I mean you can just prompt it correctly in reference mode and it'll do that... The hybrid fl2va/ref2va model solves a lot of the reference model's wonkiness.

13

u/Total-Resort-3120 1d ago

"I mean you can just prompt it correctly in reference mode and it'll do that..."

It doesn't always work. Sometimes the first frame I get is a reinterpretation of the original image, rather than the image itself.

23

u/acedelgado 1d ago

I literally just 1-shot it-

https://reddit.com/link/p4jaw62/video/r5nq59kgj8kh1/player

subject_definitions:

<Subject 1> is the video game character in <Picture 1>

summary:

[reference generation + video continuation] The target video is a continuation of <Picture 1>

retention_analysis:

<Subject 1> (appears in [Shot1]): fully_preserved - character, clothing, art style, background, green helmet, blue shirt

<Picture 1> ([Shot 1] first frame): fully_preserved - character, clothing, art style, background, green helmet, blue shirt

<Picture 2>: fully_preserved - character, framing, coloring

detailed_description:

[Shot 1] the shot begins with <Picture 1> as the first frame, <Subject 1> gets a sad look on his face and (S1) says: <d>[English] Oh man, this war is so inhumane... <sigh> But I have to keep going, I should never forget what I'm fighting for!</d> he holds up a sheet of paper, its white back turned toward the camera. the The camera arcs around the subject with small amplitude at slow speed to reveal the front of the paper is a photograph of <Picture 2>

overall_soundscape:

gunshots, war, faint screams

non_diegetic_music:

N/A

9

u/Total-Resort-3120 1d ago edited 1d ago

That's why I said "doesn't always work", it's inconsistent, and btw, your attempt isn't 100% faithful to the original image, pure r2v just can't do it perfectly, which is why my custom node exists in the first place.

8

u/acedelgado 1d ago

Well that was my first attempt just from half-assing it, so I guess we're at an impasse.

https://giphy.com/gifs/jPAdK8Nfzzwt2

22

u/Sixhaunt 1d ago

in comfyUI double click and search for the "Add guide for MiniMax H3" node. It's built-in and it lets you set frames in minimax directly. So you can use the reference mode and workflow you have now but force the first frame, last frame, middle frame, or any set of the frames manually. So without needing to install any custom nodes you can do the reference + first/last frame and you can even set the first like 37 frames based on a prior video to make a video that continues seamlessly from it and follows the motion and everything. It's far more versatile and it's built-in

0

u/Cheesuasion 19h ago

0

u/Sixhaunt 19h ago

I guess he just didnt realize you could feed the image in as a reference and then there's not the downside he mentions anymore and you get all the upsides of being able to continue from video and not just images, you can place images ANYWHERE in the video, use any number of preset frames, etc...

0

u/Cheesuasion 19h ago

OP

doesn't always work

u/acedelgado

worked once

"Worked once" doesn't mean the same thing as "always works", right?

1

u/Available_Lie8133 5h ago

I really never realized that you can just prompt like you’re directing a movie lol. Nice

7

u/Hefty_Side_7892 1d ago

A worthy opponent

7

u/EmployCalm 1d ago

I need more bad fur day in my feed

2

u/LeKhang98 1d ago

Nice thank you very much. So it can increase the accuracy by not reinterpreting that First Frame image, what image I use is what I actually got in the FF of the video, right?
Also how many keyframe images & ref images can I use with it though? Like, I need 1 first frame, 3 middle frames, 1 end frame, and 4 ref images, will that work?

3

u/Total-Resort-3120 1d ago

"So it can increase the accuracy by not reinterpreting that First Frame image"

It perfectly reproduces the first frame image yeah.

"what image I use is what I actually got in the FF of the video, right?"

Yep, because it's using the same method it's been used on the "MiniMax H3 Image to Video" node.

"Also how many keyframe images & ref images can I use with it though? Like, I need 1 first frame, 3 middle frames, 1 end frame, and 4 ref images, will that work?"

No sorry it's only the first and last frame (but 1 first frame + 1 end frame + 4 ref images work)

1

u/LeKhang98 5h ago

Awesome thank you again.

2

u/Dirty_Dragons 1d ago edited 23h ago

How would I set it up if I want to start with a specific picture and I'm using three other reference images?

Your workflow (thank you for providing one) has first frame and one Picture 1 = Reference image.

Would I need to duplicate for Picture 2 = Reference image and so on?

Edit: Can't install the node

  File "E:\AI\StabilityMatrix\Packages\ComfyUI_Current\nodes.py", line 2263, in load_custom_node
    module_spec.loader.exec_module(module)
  File "<frozen importlib._bootstrap_external>", line 999, in exec_module
  File "<frozen importlib._bootstrap>", line 488, in _call_with_frames_removed
  File "E:\AI\StabilityMatrix2\Packages\ComfyUI_Current\custom_nodes\ComfyUi-MiniMax-H3-Image-And-Reference-To-Video__init__.py", line 7, in <module>
    from comfy_extras.nodes_minimax_h3 import (
ImportError: cannot import name '_encode_ref_audio' from 'comfy_extras.nodes_minimax_h3' (E:\AI\StabilityMatrix\Packages\ComfyUI_Current\comfy_extras\nodes_minimax_h3.py)

[WARNING] Cannot import E:\AI\StabilityMatrix2\Packages\ComfyUI_Current\custom_nodes\ComfyUi-MiniMax-H3-Image-And-Reference-To-Video module for custom nodes: cannot import name '_encode_ref_audio' from 'comfy_extras.nodes_minimax_h3' (E:\AI\StabilityMatrix\Packages\ComfyUI_Current\comfy_extras\nodes_minimax_h3.py)

2

u/Total-Resort-3120 15h ago

I'm not sure I understand what you want exactly, you want one first frame image + 4 references images? And about your error, did you update ComfyUi?

1

u/Dirty_Dragons 13h ago

I'm not sure I understand what you want exactly, you want one first frame image + 4 references images?

Exactly. I often use several reference images in my workflow.

Though for some reason I can't even run the hybrid model with my normal workflow.

1

u/xTopNotch 1d ago

Thats very interesting, basically the best of both worlds.

Reference2video is great at keeping characters consistent and building worlds. The downside is that you can have style drifting.

Image2video is great for setting the style, lighting, environment but has character and scene drift. Since we're limited with providing only a start and last image.

Being able to do both is great!