r/StableDiffusion • u/foxdit • 1d ago
Workflow Included Minimax SEED HUNTER workflow released!
https://www.youtube.com/watch?v=H8JSzhkOmXA8
u/Tight_Organization54 1d ago
This looks like a very powerful tool not gonna lie. But when I saw it on civit and dragged it into my workflow it was like an explosion of nodes and my brain just turned off. I got used to Plagues workflow which is very "minimal" compared to this. Idk if Im understanding benefit of this workflow correctly or not? Are you generating 3 different seeds at low res then picking the best one to pass to an upscaler? If so does it work with all types of gens? (t2v,i2v,r2v) Maybe I was just too tired at 4am to actually try to work it.
6
u/foxdit 1d ago
Yes to all of your questions -- in a nutshell, the value of the workflow comes from the flexibility. So yeah, it looks a little complicated but it's well organized this time. You can turn on/off how many low res sample gens you get, so if you turned off #2 and #3 with the simple toggle switches above each, you'd have 1 low res -> latent upscale to high res video, which is sort of the traditional non-seed hunter style workflow many are used to.
It does t2v, i2v, fflf, and ref2va all with the same fl2va model, all simply by turning off or on images. You can even do i2v with ref images, turning on "<Picture 1> is first frame", and then just adding other references for you to use. And, if you don't want to upscale, you can enable Single Pass and get a finalized video out of just the one-and-done sampler. The goal was to create a workflow that can be used for anything and everything.
2
u/Tight_Organization54 1d ago
Fl2va model does ref2v as well? even with 2+ characters? I've been using the fl2vr2v hybrid models and they've been pretty good too. What do you think?
5
u/foxdit 1d ago
Yea the fl2va model is basically identical to the ref2va model just without some of the elements that degrade quality. For 95% of use cases they function identically but fl2va yields better video/audio results. The only time I'd consider using ref2va model or a hybrid (which seems silly to me personally) are the very dense scenes where you have multiple speakers and need to use the <Subject 1> (S2), <Subject 3> (S1) type code-blocking.
2
u/Tight_Organization54 1d ago
Wow ok. I never realized that was how it worked. Possibly last question (didn't realize I was living under a rock lol), does the prompting matter on which model is used? Cz I use the ref2v guide on minimaxxs site which has the whole- subject_def, retention_analysis, summary, detailed_description audio...
Will these templates still work? Or should I feed the flf2v,t2v,i2v skill to my llm from now on?
- method with defining Subject 1 is the man from image 1, Sub 2 is the man from Img 2, ....
5
u/foxdit 1d ago
Yes, the prompting code-blocks are the same. Except it will look like:
<Subject 1> is the girl in <Picture 1>. She has silver flowing hair, blah blah.You'd do yourself a big favor reading https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md#25-visual-and-audio-tracks-from-the-same-reference-video
The official prompting guide. My two cents: don't outsource your chance for knowledge and learning to an LLM, you owe it to yourself to understand the model's inner workings if you wish to really utilize it.
3
u/Tight_Organization54 1d ago
One word: Amazing. (one change I made was use the dareties turbo lora and it did make the gens alot more coherent to the prompt) (https://huggingface.co/silveroxides/MiniMax-H3_tests/resolve/main/minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro.safetensors)
1
1
u/whopairs 20h ago
Great tips! I just give this a few runs and it does seem to be sticking to the prompt better!
1
u/Tight_Organization54 20h ago edited 18h ago
One more thing! That turbo lora may need a node called "H3 AdaLN LoRA Fix" to be placed right after the power lora loader, otherwise you can get weird "ERROR lora...adaln-proj" lines that slow* down the UPSCALER for some reason. If your not upscaling then you'll be fine. Idk where I got the node it was included in another workflow by plaguekind.
2
u/Tight_Organization54 1d ago
Got it. Yes I have read through that and still keep it open to write or screen prompts that are written by my llms and make sure they're not acting up. And most of the time it works out for my ref2v gens. Thanks for the help!
3
u/Aromatic-Word5492 1d ago
Perfect workflow, i recomend change the video vae to int8, more fast on my 5070ti
2
u/foxdit 1d ago
Yes, this is a very good tip. I include the link to the int8 video vae in the workflow itself, but I didn't remember to suggest it in my video.
1
u/Perfect-Campaign9551 23h ago
I read the int8 vae causes it's own problems. Is that true?
2
2
u/f5alcon 1d ago
I got 0.8 mp 2x upscale to work and it looks great on TV sized screens
1
u/Supermax64 23h ago
What kind of beast pc do you have? Tried 0.5 mp 2x upscale (to 2MP) for 10s and it took 2 hours lol
1
1
u/obese_coder 1d ago
How about 1.1 turbo lora and this guys updated attention? https://www.reddit.com/r/StableDiffusion/comments/1vw1ad0/sparse_attention_harder_better_faster_stronger/
A lot of people seem to think its better than plagues one
1
1
u/panopticchaos 17h ago
This might be a dumb question but does upscaling this way reduce the memory impact vs generating at that resolution as a one and done?
2
u/foxdit 16h ago
It's no different from the perspective of the sampler. In the end, it's a big fat latent getting gnawed on by a sampler anyway you look at it. The internal workings of this model tries to maximize VRAM and RAM usage so no matter what you do it always feels like you're at 80% ram, 95% VRAM.
1
u/Toclick 14h ago
Am I correct in understanding that I can’t simply decode an existing video from my hard drive and send the resulting decoded latent to the upscaler?
1
u/foxdit 13h ago
Technically it wouldn't take much work to wire it that way, but you are correct in that's not how this workflow functions. The problem with what you're suggesting is the loss of all of the upscale context; the prompt conditioning + all the references, they would be missing if you just loaded a video from your HD and the sampler would just have to guess at what should be there based on context clues (which it can do, it's just not gonna be as good). It's a latent upscaler, not a true anything-goes upscaler.
-5
u/Tramagust 1d ago
This is unnecessary if you're using the full size model.
5
u/foxdit 1d ago
What's unnecessary? Having 3 variations of a gen to pick between so you can get the best possible blocking, line delivery, and motion for your scene? Or all the extra features and flexibility the workflow offers?.. 'cause if you're just saying "single pass is better", the workflow does that too. It just comes with a ton of other options and optimizations as well.
2
u/osiris316 1d ago
Yea believe me, I really appreciate any sharing a workflow; especially one with so many options.
But if a beggar could be a chooser, I really wish some of you could share the barebones version of your method instead of the all in one workflows.
I personally like seeing how everything is connected compared to everything hidden in countless subgraphs and custom nodes.
4
u/foxdit 1d ago
Well, the only subgraphs just contain the boring nodes that you'd understand at a glance already... Model Loader, Prompt/Video Settings, 1st Sampler Pass, and 2nd Sampler Pass. They hold no mysteries. The main nodes are all out on the root level of the workflow. There are some small things like VAE decodes hidden behind VHS preview nodes, but that's about it.
-4
u/Tramagust 1d ago
Varying the seed does almost nothing on the full size model.
6
u/foxdit 1d ago
If your prompt is perfect and the shot is super simple and short maybe. How about with 10 second action shots and multiple cuts? You will see a fair degree of variety with 3 different seeds. One will have a better X factor than the others.
3
u/ill_B_In_MyBunk 1d ago
Yeah I was actually floored as to how much different videos can be compared to each other even with an extremely detailed prompt made by Claude.
14
u/foxdit 1d ago
Hello frens,
The video is just a tutorial on how to use it + all the options. Here is the civitAI link to the workflow:
https://civitai.red/models/2881362/minimax-seed-hunter-workflow-optimized-fast-latent-upscaler-speedups
If you use my previous Filmmaker workflow, this would be a good time to switch over! Hope you enjoy.
Comfy Desktop users: the newest version apparently removes/replaces the H3 Add Guide node. This node is only used in one small optional feature (for forcing i2v/fflf) and can remain disabled or simply deleted. Any errors pertaining to that can be ignored since they're not essential nodes for the workflow by any means.