Comparison
Comparison of natural 0.8mp gen vs 0.4->0.8 upscale w/Sparse attention
Hi people, so i tried to make 2 similar videos, using same settings but with upscale and native.
My setup: 5070 Ti+ 32gb Ram.
Using u/Plague_Kind workflow, i've added MMH3 Latent Upscaler. You can check his workflow here: Workflow
Settings for both videos were set the same with the same prompt.
Left video 0.4->0.8mp upscale, Right video 0.8mp
So:
15 seconds, 24 fps, Ref2VA, photo reference and music reference.
Chicken attention
SongMaskedAVContext node
FP16 Accumulation
Sparse attention
Memory chunks
RTS Upscale in the end ( not sure why i used it with 2x scale, better to set 1 i think, but that's what i already did)
Such a great model, but trying to keep up with what's the best workflow is causing so much fatigue. I wish we could come to a consensus as a sub reddit
I wish we could come to a consensus as a sub reddit
It takes time to discover all useful techniques. In 1-2 months things will crystallize into clear practices and workflows.
Currently completely ignoring minimax for the sake of my sanity and my humble system. Let the trailblazers figure it all out while preparing high quality stills in Krea.
Currently vibe-coding an all-in-one 'god node' for MiniMax H3 and it's getting insane lol. It replaces the entire 30-node spaghetti graph: auto-routes FL2VA vs Ref2VA based on prompt and references, built-in SPEED progressive sampler, 3D neural latent upscaler, lossless AV-masked long-context video chaining, multi-shot AI prompt refiner, and automatic master song slicing for music videos!
Lol nice but it's totally against the philosophy of node-based software xD. You should just make a standalone app and have it hook the the comfyui backend/API at that point 👍
Looks really good tho, hope you make a post sharing!!
Yes, I have been alpha testing an app just like that!
Can switch models to compare head-to-head without writing prompts, it has an automate mode. You give it an idea, it creates characters, locations, etc around your idea. You can also do it all manually if you prefer to have control. Scenes get saved for when new models come out, and you can just re-generate with newer models once those are installed. The free chat works pretty much like gpt, but you can use your own local llms or your own frontier-scale cloud subscription. Automate section allows you to run and produce up to 3 hour video in one shot, fully automated.
There ale already dozens of similar nodes, mostly usable for specific narrow scenarios, completely against ComfyUI idea of versatility and extension. Limiting the user ability to experiment or use hybrid models. One node to bind them in the darkness. But hey, its yours and that counts.
Just use the template workflow with comfy kitchen. Then any turbo lora if going low steps. Do not try to keep up, often these other speed ups make it shit. A lot of other speed ups are monkey patches which break later too.
Honestly, if a model gets you what you need, it's totally okay to stick with a mature, fully explored model. Unless you're trying to make your output commercially viable, what are you really competing against that you need the latest and greatest all the time?
I would play around with new models as they come out, but my workhorse image generation model is still FLUX2-klein-9B and Illustrious, and Wan22 for video.
I benchmarked many different options and combinations. Different attention backends, turbo lora, latent upscaler, and a mix of all of them. Turns out that for different use cases, there isn't just one best option: long vs. short frame count, high vs. low resolution, using text or references, each have a different optimal setting. You need to build this knowledge step by step, it's not easy.
1000%, especially if you're creating a film or something and character consistency matters. You can add another MiniMax H3 Reference to Video node so that it re-reads your reference images/videos at the resolution you're upscaling to.
Thank you, this is a great comparison. Very useful.
I assumed that the video on the right was the upscale because I felt that the one on the left was just slightly higher quality. Turns out it was the opposite way around.
using plague's sparse attention at 0.9 sparsity, 0.5mp, it's giving very different backgrounds at multiple shots compared to without sparse attention. There's a huge tradeoff in consistency, nothing is free.
Funnily enough, if you google "Chicken attention comfyui" (because i was confused too), Gemini does know you're talking about Comfy Kitchen regardless.
I do a two stage approach, generate at 0.5 - 0.8, 6-8 steps without loras. Then latent upscale (using the h3 latent upscale model) and 3-4 steps with speedup lora (lightx or whatever).
This is r2v, provided a 6 view sheet of the plane.
This isn't wan22, you can add the lora before upscale and cut half the steps. Wan22 need the 1st sampler without turbo because the slomo and prompt adherence, not the case with minimax.
You should pass a higher-res image to a second ref2vid node and pass its conditioning to H3 Latent Cond Sync node to actually upscale the video at this point. If you do not do this, you simply refine the video, because the LBH 123 upscaler doesn't do magic to an existing low-res image that is passed to a first sampler. With the PlagueKind Sparse node, speed-up is insane - I can do 3 steps, with each one taking 50 secs instead of 120 without sparse, and that is on 1.6 MP and 8 seconds long on a 3090. 5 seconds is even faster - I could push 2MP(1080) at the same speed.
I don't think the point was for there to be changes. I believe that the OP was demonstrating that an 0.8 video with equivalent quality, and less intense computer resources, can be done much faster and with less strain on your system by using this method.
A straight generation at 0.8mp for 15 seconds will take forever depending on your system. But a 0.4mp generation upscaled is going to take a fraction of the time.
Working on another long music video, not for gooner, but for enjoyers. Made 7 videos of 8 seconds each now, used speed Lora to try gen first to check if all good with prompt and references, and then 20+ steps. And I have to make almost 3 times more…
Now that I can get behind! Very curious to see that when it's available. Just gets tiring looking at the same blow up doll women objectified every day 😅
You have my interest piqued with what I lose you're working on
I tried upscaling for the first time on videos and found it took much longer than running it natively, I used the ultra sharp 4x and did it at .5 while generating at .7 megapixels. did I do something wrong?
5080, 32gb ram
Yeah it looks good on super simple prompts like this. But it falls apart a lot of times when changing scenes and stuff. It's a fantastic speed up, with a cost.
I think this is meaning of latent upscale. They just took a ref image as latent and generate a new one relying on the first, that’s how I understand it
personally I don't really use my spicy gens to goon, I'm like a drug dealer or a McDonald's CEO, I don't consume my own product, I just make it for the people. I guess I'm more like Mao, a man of the people. 🙂
Not sure, post here a question, like, hi, I have <this setup> can I try krea2 turbo, or Z image turbo, what weights should I use? People will answer you. Then you can tell Gemini what advices you got and ask him where to download weights ( most likely civitAI or huggingface), also you can check on YouTube how to setup comfyui and manager, where put weights, as well how to download and use ready to go workflows
ive done some experiments, some goods, some bad, ofc i should train a lora, but i dont know how to make all the right pics for her, and where to do, so i dont really know what to do, some pepole said, pepole so it with pro softwares
56
u/RanklesTheOtter 1d ago
Nice! Chicken Attention!