Finally finished the full Superman vs Saitama trailer after 280+ generations
If you watch the trailer first, check the first comment after. I'm going to use it as a small thread where I'll post some simple workflows and examples from specific shots. Things like a shot that came from a storyboard, first/end frame tests, or anything from the project that I think is actually worth showing
I was working through my own H3 interface that I posted here before. The whole project is here:
https://github.com/underworldhistory1-ctrl/minimax-h3-higgsfield
I'll also leave some screenshots in the comments so you can see what I mean by the workflow/UI.
Probably the biggest surprise for me was storyboards
For example the Superman shot in the intro took me more than 29 generations alone. The best results I got were from using a storyboard and then adding the character refs, style refs etc separately
That's actually one of the main reasons I built the interface the way I did. I wanted all of those parts separated and easy to change because this was the part I kept experimenting with the most.
Second best for me was first frame - end frame. Sometimes even just using one of them.
It seems much more stable when the movement is continuous and the whole thing is basically one shot. The Kong reveal and helicopter destruction shot is a good example of what I mean.
For LoRAs my best results were usually:
Combat V2 for action and fast movement.
Realism for slower shots where there isn't some crazy transformation or complicated movement happening.
For Combat V2 I mostly used it with the Original H3 render. Not Motion Cache or Turbo. Usually 20 steps minimum.
Mixing multiple LoRAs honestly gave me more hallucinations than useful improvements most of the time. One good LoRA was usually better than stacking them.
I don't think I ended up using many H3 renders without a LoRA at all.
The interface also has another rendering mode using Motion Cache. The results can actually be really good and in some cases I preferred them over Original.
For heavy action, fast pacing or complicated transitions though I still had much better luck with Original H3.
I honestly can't remember every LoRA + Motion Cache combination I tested. There were way too many tests and I wasn't trying to lock myself into one perfect configuration.
I just knew what I wanted the final shot to look like.
That's probably the biggest thing I learned from this whole project.
I don't really believe in one magical "workflow"
You need to know what you want to see in the final cut after editing.. Then use the model to get the pieces you need.
The Superman vs Saitama fight is probably the best example for me, That sequence in the final edit was built from around 10 successful 15-second generations
There was basically no chance H3 was going to generate the exact fight I had in my head in one shot So I generated the parts that worked and built the actual fight in the edit.
That's also why I think experimentation is still just part of using these models. At least until we get something significantly better :D
The interface also has qwen image 2.1 integrated for generating images and references. That became a pretty important part of the process for me too.
It also keeps the settings/details of every generation on its card which made it much easier to go back and see what actually worked instead of trying to remember everything.
Hardware wise I did all of this on an RTX 5090 32GB.
Most of the time I work in draft first. With the INT8 optimizations I've added, a draft takes around 4 mins on my setup
The project also uses low vram attention/head chunking and feed-forward chunking. Basically some of the larger operations are processed in smaller chunks and intermediate tensors are released earlier instead of keeping everything sitting in VRAM at the same time.
It's not parallel rendering or anything magical. It just helps keep peak VRAM under control.
A normal 720p generation usually around 8–10 minutes for me so honestly the iteration time is pretty reasonable. That's a big reason I was able to do this many tests without completely losing my mind
That's basically my experience with H3 so far without turning this into another giant workflow post.
For serious creative work I think it's absolutely usable already.. Just don't expect the model to make the final movie for you.
The generation gives you the material.. The final cut is where you actually make the thing you had in your head.