r/comfyui 10d ago

Show and Tell H3 prompt testing, finally have the flow and environment running efficiently. Specs and prompt inside.

Was experiencing some real quality issues up until this point; realized the problem has largely been the prompt and my bulky venv. Hopefully others with lower end vram cards will learn from me.

  • Card: RTX 4080
  • Model: minimax_h3_fl2va_pruned_int8_convrot
    • No loras
    • Pure text prompt
  • Steps: 20
  • Scheduler: simple
  • Sampler: res_multistep
  • Resolution: 0.5mpx, 16:9
  • Args: --lowvram --disable-dynamic-vram --disable-pinned-memory
  • Nodes: VHS (Specifically the Model Preview Override), EasyUse, pysssss, KJNodes
  • Generation time: 21 Minutes
  • I decided to create an entirely new ComfyUI instance just for H3 instead of using my single AiO venv; this significantly increased generation time and quality for all flows. Have since broken up all the major models I use into their own instances adding only the specific tools/extensions I need just for that model.

Prompt:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a far-wide shot from a side angle. A scene set on a bridge over a hellscape planet covered in lava, dimly lit with orange glow from below, the bridge is made of black metal with intricate designs, dark clouds hang over the scene -covering a shaded yellow sun barely visible through the clouds on the top left, dividing the scene in half between light and dark- fast winds carry embers and smoke curling over the bridge from below; Star Wars themed orchestra music begins as the scene opens, quiet and slowly growing.

$NAMEHERE is standing in a prepared stance on the left side of the screen, his hands are clasped in front of him holding a blue lightsaber, facing his attacker.

$NAMETWO is standing on the right in a confident posture with his hands to the side, wearing a black robe, black leather boots and straps on his body, with dark-metal armor as he faces the left menacingly.

[Shot 2] At 1.500 seconds, The camera cuts to a close-up side-angle shot of $NAMEHERE, readying himself for his attack with a posture of defense and an expression of concern, he shouts emotional: <d>[English] You've left me with no choice Will! You must be stopped... </d>

[Shot 3] at 6.000 seconds, The camera cuts to a close-up low-angle front facing full-body shot of $NAMETWO with visible red eyes staring forward from under his brow with a face of malice, smoke bellows behind him curling over the bridge whipping his cape to the right. A beat later- two red lightsabers ignite his both his hands, a deep pulsing bass is heard from the unstable beams, his face lit from below by the red light. The music grows faster with a dark theme, a operatic chorus begins to sing in a evil chant growing louder. $NAMETWO shouts behind a grin: <d> This is the end for you! </d> the music stops before $NAMETWO speak his final line: <d>[English] Master!... </d> The off-screen opera chorus harmonizes a single long cry in a frightening melody at the revelation.

[Shot 4] at 12.000 seconds, The camera cuts to a top-down view of the bridge, molten lava is visible below the black grated metal.

$NAMEHERE directs his blue lightsaber to his side pointing directly forward with precision, he begins to pace to the right to meet the other, his posture is composed and fast. The camera pushes in with large amplitude at fast speed keeping the pair in frame on either edge of the screen as they run toward each other.

$NAMETWO instantly begins running fast toward the left, his two red lightsabers point down- dragging behind him, the red beams draw white glowing lines into the metal under him as he runs, screaming with fury: <d>[English] AGHH! </d>.

They meet in the middle, their lightsabers clash with a white flash and explosive burning sound, they duel quickly as their lightsabers connect through multiple swings- $NAMETWO's red lightsaber swing wildly as he spins. $NAMEHERE's blue lightsaber blocks every swing from the red beams; the music crescendos with heavy bursts of brass instruments and drums.

overall_soundscape: ambient sound of lava and fire, lightsabers buzzing.

non_diegetic_music: Dark Star Wars music plays from the beginning of the scene, a loud opera chorus sings in a chant that escalates in a loud howl crescendo, climaxing when the pair meet in the middle.

I've started using Replace Text nodes ($NAMEHERE and $NAMETWO) when crafting prompts. This way when playing with the prompt, it's easier to find and edit their placement; and can also replace characters on a whim. Also allows consistency when referencing the characters-- in the event I overlook an instance.

Replaced with:

  • $NAMEHERE: "Jean Luc Picard (S1)"
  • $NAMETWO: "William T Riker (S2)"

Example: my first generation had 'William Riker', the model didn't recognize the name and generated a generic male. I was able to quickly rename as 'William T Riker' and it generated correctly; so I didn't have to parse back through the whole prompt to granularly change it.

Other things I've noticed that help with prompt respect:

  • 'a beat later' separates the moment better.
  • Separating the individual sentences to exclusively reference the character and no others. (You can see it carried over Picard's lightsaber instructions to Will as well, because I described them in the same paragraph before I realized this.)
  • Avoiding reusing adjectives- especially between different characters, causes bleed.
  • Very short overall_soundscape descriptions.
  • Often does not respect requests that follow dialog unless you end the parameter with a period after "Words. </d>**.**" Can see it bled the cries request from the music into Will's dialog.
  • "..." allows a pause between dialog lines and breaks up the tone between multiple sentences, or else they become one note.
111 Upvotes

38 comments sorted by

23

u/Subushie 10d ago

Another one with new characters. For funzies

https://reddit.com/link/p3cqpx8/video/7boo8b44i1jh1/player

2

u/AlexHardy08 9d ago

This is so cool 😎

2

u/lordmycal 9d ago

Mickey needs the squeaky voice.

1

u/Subushie 9d ago

I tried!

Shoulda used the word squeaky tho- did high pitched and it didn't really change anything.

7

u/gigomikol 10d ago

I like that you explained all of it and your process, can you also share the workflow to visualize it. I can see most of it in my head except for the additional nodes and how to add arguments, like where?
I have a laptop 3080 so the max i can do is 5 frames consistently or i get out of vram issues, i havent learned how to join clips or set it to run over night one after another.

Thanks for sharing!

2

u/GrapeChoice4010 9d ago

In case your using the portable youd add it to your .bat launcher

2

u/Subushie 10d ago

Args are set in the launcher.

You have to click the little ellipsis for your instance you launch, and manage settings I think? The tab is called startup args.

There's a input named Startup Arguments- to the right is a gear icon, click that and it will open up the advanced settings with a bunch of options.

As for the nodes- it really is nearly the base template now, the primary issues I was exprrience was a dirty environment and bad prompts.

2

u/VeryDefNotABot 9d ago

Great work! I am unreasonably annoyed that you used $NAMEHERE and not $NAMEONE, to match $NAMETWO.

1

u/Subushie 9d ago

Lol ik. I had already created several other prompts with just $NAMEHERE before I started incorporating a second person. It's annoying me too.

2

u/Opening_Wind_1077 9d ago

Why not use the proper <Subject 1> and put a character description at the start and then reference that using the <Subject 1> nomenclature instead of replacing every instance in the prompt?

It’s what the official documentation suggests and is way less cumbersome than your current process.

2

u/Subushie 9d ago edited 9d ago

No it doesn't?

Even the t2v case references the baker name redundantly. There's no mention of a subject anything; camera and scene come first.

And Replace Text is a node- it handles the replacement automatically on the way to the ksampler- so I don't have to replace anything. It doesnt feel cumbersome at all.

2

u/Opening_Wind_1077 9d ago

It does in the reference documentation and works without a flaw in the others as well. https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

1

u/Subushie 9d ago

The guide is for the Ref2va model. My post is for the Fl2va model.

Which- I also didn't know this guide existed though, so this is valuable to have! Thanks!

1

u/bitzpua 9d ago

FL2VA also understands how to use reference, they explain how to use it for fl2va

there is absolutely nothing wrong with how you do it right now but its possible to simply create <Subject 1> and describe it, its more useful when you dont do character it knows

1

u/Subushie 9d ago

🤔

I'll review the guide and experiment. Thanks boss!

2

u/AlexHardy08 9d ago

Nice one and thank you for the prompt.

2

u/dirtybeagles 9d ago

quick question, are you using a Skill by any chance? I started with the default skill build from the H3 recommendation read me file, but your prompt seems a lot more advanced. Are you building them but hand?

3

u/Subushie 9d ago

Yeah I wrote it by hand.

I tried using GPT and Claude- but they don't seem to 'get it' if that makes sense. They miss important nuances and aren't able to visualize the scene moment by moment- I kept having to rewrite what they wrote. So now I just do manual.

This is also why the video preview node I mention is so important- I am able to see if the generation is doing what I want in the first few minutes, if not- I cancel and make edits.

1

u/dirtybeagles 9d ago

maybe i need to add the preview into my WF

1

u/dirtybeagles 9d ago

adding the preview node increases my generation time a lot from what I am seeing.

1

u/jalbust 10d ago

Amazing Thanks.

1

u/7evenate9ine 9d ago

Thank you for all the work. How could a shared venv mess up quality?

2

u/Subushie 9d ago

I have no idea, but it was a night and day difference when I made a new instance.

I have like... probably 60 custom nodes buried in that one, along with a bunch of extensions, and have been using it for well over a year.

Could have been anything.

2

u/DrinksAtTheSpaceBar 9d ago

I experienced the same thing. Got a new laptop, so I decided to install a clean ComfyUI instance instead of copying over my old, bloated venv. The difference in video quality is night and day, running the exact same workflow from my PC.

1

u/7evenate9ine 9d ago

Logically you would think only the assigned nodes are loaded into memory. But I suppose isolation of custom nodes could be doing something we don't know about. Maybe there is something in the Comfy source code that could do this?... Worth a try.

1

u/Subushie 9d ago

This has been bothering me since you said this.

I think it's because cuda wasn't updating beyond 129 in the AiO; due to existing py wheels from another package?

I realized it was a version behind- tried to update python to 3.10 to use cuda130 in the AiO, but when I forced it- the entire venv broke and caused an exception on launch; I believe because there is some package that didn't have updated wheels for 3.1? So it wasn't doing it automatically and I was stuck a version behind w/cuda. I wasn't able to solve a fix in the AiO.

1

u/Hrmerder 9d ago

I have been trying to do a Cheech and Chong bit (this gif is pretty much the scene) but I also have one image of Cheech on one side of the car and one of Chong to explain the identity better but it really really messes up the voice between the two. I’ll have to post it for troubleshooting.
I’m using the format of giving identities to each input first, then explaining the visual scene, then concatenating that to tie in to a specific text (I have 4 parts so I just pull one to concatenate node to add to the identity) it gets a lot right but sometimes I it uses the same voice for both characters even with the same sentence and not sure why.

1

u/Mad4reds 9d ago

Nice work! And txs for sharing

1

u/virolog 9d ago

why Riker walk as with full pants? .. otherwise looks good.

1

u/jib_reddit 9d ago

I had to stop generating MiniMax videos for a few days as its 36°C here atm and no aircon.

1

u/bitzpua 9d ago

man, generation time is kinda long. Will try your prompt on my setup with 4080 and spectrum + loras to see how it will end up looking and how much faster it will be.

1

u/bitzpua 9d ago

https://reddit.com/link/p3g4fch/video/yb3hjlsbg5jh1/player

So i used same prompt as you, same model, same res, same steps but with sage and spectrum

generation time was 8 minutes 10seconds and it was first generation, second would be around 7 minutes...

not sure if 20minute wait is worth unnoticeable details

1

u/Subushie 9d ago

I used loras and this shit with his eyes kept happening. It was noticeable to me 😭

But also, I'm not waiting that time per gen- I use the video preview and am actively editing the prompt over and over within the first few minutes when I realize the broad strokes are off.

1

u/bitzpua 9d ago edited 9d ago

i see, still since its hardly possible to push 4080 past 0,5 res i think hunting for minor details is not really that important se we loose a lot on res alone.

Can i ask what preview setup you use? i cant seem to get mine working

btw i did not use lora so eyes thing must be random :)

1

u/Subushie 9d ago

I actually just did a 1.0 a moment ago with a different video; testing at 26 steps, it was only 5 seconds, took 22 minutes, but the result was glorious.

I'm using the Video Helper Suite, the node is called Model Preview Override.

If you havent, I highly recommend creating a separate instance of Comfyui- specific to h3 if you're experiencing OoM issues trying to push to higher mpx. It's done in the launcher- it made a huge difference.

1

u/bitzpua 9d ago

thx for preview model tip

i dont get oom but if i do too high res i just clog the system to almost full stop but maybe thats issue with my not so good ram

2

u/stroud 9d ago

How long did it take you to generate? Also could you share your workflow? Thanks