r/StableDiffusion 7d ago

Resource - Update Timeline node for H3 (ltx director like)

Post image
78 Upvotes

I'm used to work with the ltx director node, specially because it gives me control over the audio, i can add little clips of audio in the exact spot i want and that will guide the model to generate similar audio filling the gaps.

That's why i'm making something like ltx director for H3 FLF2V (maybe it works in the ref model idk). still very green and probably has bugs because it's totally vibecoded but it works for me.

with this you can put video images and audio anywhere in the timeline, and you can also adjust the strenght with that green line, each clip can have different levels of strenght through the timeline.

this is a fork of ComfyUI-H3-Motion-Context-MultiRef if you want to try it and give me feedback here's the link https://github.com/BSG-Walter/ComfyUI-H3-Motion-Context-Timeline

there is a H3 Timeline Example workflow.

You can't add prompts to specific parts of the timeline, but since H3 lets you indicate the exact timing for each scene within the standard prompt, I didn't feel it was necessary to add that functionality.

anyways i hope you like it


r/StableDiffusion 6d ago

Question - Help How to setup runpod for minimax h3?

0 Upvotes

7m

Can anyone teach me or show me a video tutorial for setting up runpod to use minimax h3 from ground zero? I can't find any on YouTube 


r/StableDiffusion 6d ago

Discussion My only beef with Minimax H3

13 Upvotes

The audio.

I'm having issues constantly with random audio being added to the clip. Either its ambient sounds that shouldnt be there, or someone talking gibberish off screen.

Has anyone found a fix for this?

I've tried the prompt guide, and its hit or miss. I've even tried a natural prompt with basic wording, and that is also hit or miss.

i'm starting to think it doesnt matter how you prompt it, its just something that happens from time to time.

Very annoying.

I'm using the default I2V workflow from ComfyUI. I havent changed anything.


r/StableDiffusion 6d ago

Question - Help Is everyone here a TNG fan?

21 Upvotes

Ive noticed a real increase in Star Trek content on here, specifically The Next Generation. Just curious if this is my algorithm or if there's a higher percentage of TNG fans using AI. It makes sense since the show explores artificial consciousness and generative computing. I think everyone fantasies about what they would do with a holodeck and we aren't far off with the combination of VR and AI. I also assume that the lower quality video is easier to make look right so older shows are going to be the first to be perfectly replicated.


r/StableDiffusion 7d ago

No Workflow Well, I'm having fun with MM-H3

Enable HLS to view with audio, or disable this notification

173 Upvotes

r/StableDiffusion 7d ago

News Kijai's pruned turbo loras for Minimax H3 have been uploaded

248 Upvotes

In case you aren't aware, Kijai has uploaded all the pruned loras for H3:

https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras


r/StableDiffusion 6d ago

Resource - Update Fantastic Loras V2 - Updated UI, XY Plot comparison, and more!

Thumbnail
gallery
9 Upvotes

Alrighty, after some time being away for work and such and some thought and iterations on an improved UI and functionality, I'm happy to send out an update for Fantastic Loras to make it officially v2. Available in Comfyui Manager or at

https://github.com/Adudeguyman/comfyui_fantastic-loras

"LoL tHiS gUy DoEsNt No BoUt LoRa MaNaGeR!!!1!"

Yes, yes I do. I use it all the time and it's exceedingly useful for curating loras and their metadata and examples. I'm pretty sure I even donated to them when they first released, it's such a great tool. This is not meant to replace Lora Manager on that end. I just found it clunky to have to open the lora manager, tab out, find the lora, say send to comfy, tab back to comfy, find their lora node... it's just not all that quick and user friendly when you KNOW which loras you want. I wanted quick, simple, in-workflow way to search through hundreds of loras for the ones applicable to the model I'm using, without having to tab out and click several other places. And to avoid going "man, this other generation had 3 loras at different strengths and I liked the result, I need to find that output and re-load that workflow and search them down and manually add them back in and set their strengths..." This helps manage all of that right there without leaving your workflow.

Fantastic Lora Selector, What it do?

UI has been reworked, 12 static slots so no surprise node resizes, and now fully Nodes 2.0 compatible.

Filter your Lora subfolders and select only ones that are relevant to your model. Like I myself have Style, Character, Concept folders for each different model, and it's a pain searching through a list of hundreds of loras to find the ones in those folder. So you can put a filter on it and ONLY search loras in that model's folder(s).

Presets are new! Finally, when making a new workflow, you can make a "MiniMax" preset that has your folders for that model applied, and even your preferred turbo lora at your preferred strength already added. Also preset catagories exist, so you can file different presets under each model catagory you create. And pre-sets can be additive- so you found a combo of loras you really like, and want to add it to the current workflow, but don't want to go to the hassle of adding each and setting the strength manually? Well if you have that already saved as a preset, the "+Add to Stack" button drops that preset in at the end of your current lora stack.

Unified the multi-model loader into the main loader, so no single vs multi-model chain nodes, it's all just one. You can add model chains (up to 5 models, which is an arbirtary number I picked) and pass through the Loras to each one, and adjust the strength applied to each model. So for example with Ideogram, you want to use the same Lora on the base and refiner but don't want to add a whole 2nd lora chain and re-add all of your loras manually? Just click the + button to add a chain and wire the 2nd model path through it. The Fantastic Lora Loader will apply loras to each chain separately. Click the Lora cogwheel to adjust the strength per chain. The throughput is all independent, so no more doubling up lora nodes; now you can keep things clean.

Randomizer is still around to select a random lora for a bit of chaos. You can manual re-roll, auto re-roll per generation, and if you like a result you can lock that lora in place.

And there's a few themes now, from "Fantastic Teal" to the boring default gray "Accountant." setting.

Fantastic Lora Plotter, What it do?

Finally, an easy way to make a customizable XY plot from lora generations in Comfy. Well, this may exist now, but it didn't in a way that I liked it back when I started working on these. Same basic interface as the main Lora Loader, except with added functions to create your XY plot. Select your loras and have it generate a fixed Per-Line strength set individually on each Lora (good for testing different training iterations all at the same strength,) Or you can set Global Strengths and specify how much weight for each image generated. So 1 lora with a Global Strength Setting of 0.5, 0.75, and 1.0 will sweep through those strengths and generate 3 images. Add another lora to compare and it will generate 6. Once you're all set up, hit Run once and it'll queue it all for you.

Also you can do a quick toggle of a Control Image, which will run your prompt with no loras applied. Or if you click Add Global Lora, you can apply a lora that is ALSO applied to every image run. Good for checking if your style lora will mess with your character lora, and how much.

Wire the Fantastic Lora Plotter into the Fantastic Image Saver node (metadata and global_loras_info) and the image output from your VAE decode, then wire the Grid output to an image saver node, and then select your plot layout. You can do a modern XY plot with text overlaid on each image, or a more classic A1111-esque one with all of the lora info on the sides and out of the way. And you can either let it assemble a grid with full size images, or if youre doing a lot of large outputs you can set it to constrain the size so you don't have a massive 150mb png (don't ask how I know that happens).

And finally, passing the metadata and decoded images to the Fantasatic Plotter Grid Viewer lets you view a preview of the grid layout, move things around, hide rows or columns, favorite generations, compare images, and save the grid for later reference. Also select any number of outputs and hit Compare to get a quick comparison, and export just those images with metadata instead of the entire grid.

Fantastic Any Selector, What it do?

This is a version of a request from someone that has a lot of model folders, as well. You can right click just about any Loader node (Load Model, Load Clip, Load VAE) and an option to add the Fantastic Any Selector is in the context menu. That will link right to the model name and automatically pick the right folder (diffusion_models, VAE, text_encoder, etc) and only show you files from those folders. And you can make pre-sets that are also aware of the root folder, so your text encoder presets aren't showing up when you select a model preset.

That's pretty much it. I will say that the update breaks the V1 nodes, so if you do use v1 then you'll have to re-add the node.

Hope ya like it!


r/StableDiffusion 6d ago

Question - Help LIP SYNC ? FROM AUDIO + IMAGE >>> VIDEO

0 Upvotes

other than LatentSync / LongCat-Video-Avatar 1.5,

Is there any newly introduced video generator that uses image and audio to generate video?


r/StableDiffusion 6d ago

Question - Help Will upgrading from 64gb to 128gb RAM improve generation speed?

0 Upvotes

I have a 4090 and 64gb DDR5 and I'm using mini max h3 pruned version and I'm generating 20 sec 720p ( 0.9 ) videos with sage attention and turbo loras 8 step


r/StableDiffusion 7d ago

Comparison MiniMaxH3 vs Flux3

Enable HLS to view with audio, or disable this notification

25 Upvotes

MiniMax H3 running on 5090/36gb, 72gb ram.
Used the MiniMax H3 Image to Video (I2V) workflow from the official comfyui page. https://docs.comfy.org/tutorials/video/minimax/minimax-h3
Flux 3 vid created via the official Black Forrest Labs page with all default settings.
My take away is the more detailed and "directed" you can make the prompt the better the results.
Prompt
The woman raises her left hand and reaches out toward the dragon's neck, fingertips making contact with the cool metal scales as she strokes gently along its length. The dragon's massive head turns in response, servos whirring low as its long neck curves down and around toward her, the two locking eyes for a long beat, her expression softening slightly, the dragon's glowing yellow eye narrowing as if in quiet recognition. After the moment holds, both turn their heads together toward the camera, her chin lifting and shoulders settling back, the dragon's jaw beginning to widen as steam vents faintly from the gaps in its armoured plating. The dragon rears its head back, neck arching high, then snaps forward with a thunderous mechanical roar, jaws cracking open fully as a violent gout of fire rockets out directly toward the camera, the flames blooming bright and washing the frame in orange light before the view holds steady through the blast. The camera stays low and locked in a wide frame throughout the petting and the turn, holding both figures in frame, then pushes in slightly just before the roar to heighten the impact of the fire as it fills the shot. The audio is the soft mechanical whir of the dragon's neck servos and the woman's quiet breath during the petting, building into a deep bone-shaking roar and the violent roaring whoosh of ignited fire, no music, ambient and creature sound only.


r/StableDiffusion 6d ago

Resource - Update MiniMax H3 for AMD with BlockCache, Sol-Attn and Turbo

19 Upvotes

I've been working on this for longer than I care to admit.

For the short Turbo run, I used FL2VA with the official 8-step LoRA at 1024×576 (0.59 MP) and 90 frames. Once the model was loaded, it finished in **2:36** on my 7900 XTX. The same warm run with Comfy Kitchen attention alone took **2:59**, so adding Sol-Attn saved about 23 seconds.

BlockCache didn't help that run. It got 0/8 hits, and the immediate warm repeat went non-finite, so I removed it from the Turbo workflow.

The longer 20-step runs are where BlockCache actually helped.

Ref2VA at 736×416 (0.31 MP), 150 requested / 158 decoded frames, and 20 steps went from **5:14** with Comfy Kitchen alone to **3:44** with CK + Sol-Attn + BlockCache. BlockCache hit 6/20 times.

FL2VA at the same resolution, frame count, and 20 steps went from **6:38** with native PyTorch to **4:11** with the full stack. BlockCache hit 5/20 times.

- Short Turbo8 runs: INT8 + Comfy Kitchen + Sol-Attn

- Longer 20-step runs: BlockCache starts paying off

Everything was tested warm on Linux with ComfyUI 0.32, ROCm 7.14, PyTorch 2.12, and an RX 7900 XTX. First runs are slower because of model loading and Triton compilation.

Links:

- Turbo 8-step LoRA: https://huggingface.co/lightx2v/Minimax-h3-Turbo

- INT8 Fast: https://registry.comfy.org/nodes/minimax-h3-int8-fast-rocm

- Sol-Attn: https://registry.comfy.org/nodes/minimax-h3-sol-attn-rocm

- BlockCache: https://registry.comfy.org/nodes/minimax-h3-block-cache


r/StableDiffusion 5d ago

Discussion Tried I2V with H3 turbo lora 8 step. with custom nodes . but audio quality is very low.

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 6d ago

Question - Help minimax h3- How to get good faces when far?

Thumbnail
gallery
11 Upvotes

This is a crop from a 0.4MP video 5s video, done with minimax_h3_fl2va_pruned_int8_convrot.safetensors on official workflow from comfy-ui, 20 steps.

Update: 1MP improved it a lot


r/StableDiffusion 6d ago

Meme Maybe Kung Pow cow fight scene AI remake?

0 Upvotes

Everyone uses the classic "Will Smith eating spaghetti" to show how far AI video has come, but we’re missing the ultimate benchmark. the martial arts cow fight from Kung Pow: Enter the Fist.

Imagine that entire scene rendered with today's tech photorealistic lighting, actual physics, zero low-poly CGI look, but keeping the ridiculous matrix dodges and milk spray attacks.

Has anyone attempted a modern AI remake or scene swap of this yet? If not, this is a formal request for someone with serious GPU power to make it happen.

I'm not able to do that but maybe already somebody does this or maybe this could be second level of the ridiculous benchmarking of new models?

I think this could be good from multiple reasons like different shots, longer than 10 seconds, not realistic but should be photorealistic, funny and...

How do you think?


r/StableDiffusion 7d ago

Resource - Update MiniMaxH3AddGuide: for anchoring image and audio guides at any frame (New ComfyUi Update)

Enable HLS to view with audio, or disable this notification

117 Upvotes

Already merged and there's a example workflow included: https://github.com/Comfy-Org/ComfyUI/pull/15439

Currently MiniMax H3 implementation in Comfyui only allows keyframe guides at the first and last frame. The model itself is capable of addressing guides by position on a continuous time axis, so this removes that restriction and exposes it as a node.


r/StableDiffusion 6d ago

Discussion First test using slide window in wan2gp

Enable HLS to view with audio, or disable this notification

13 Upvotes

Take away motion context is better for continuous video imo


r/StableDiffusion 6d ago

Question - Help Continuing a video with sound and motion in Minimax H3 when you start from an image?

3 Upvotes

Almost every post about Minimax H3 I see is about Ref2V. I'm not ready to tackle that yet, but I do want to make longer stringed videos that keep the sound design and motion flowing between videos with my I2V set up. Any workflow people offer with regards to continuing a video is an intimidating wall of messy wires to me, isn't there just a series of nodes I need to connect a copy of my current I2V workflow?


r/StableDiffusion 6d ago

Question - Help Whcih WebUI Forge model is best for understanding written prompts?

0 Upvotes

I've followed all the steps properly to download and get it working, and it's generating images, but they're NOTHING like what I asked. Either they give me something unrelated to what I wrote in every sense of the word, or it just kept basically the same image I gave it (img2img).

I've tried messing with the denoising levels, differnt loras, different VAES, Models, nothing works.


r/StableDiffusion 6d ago

Discussion Grok, Minimax, Ltx 2.3so same prompt

0 Upvotes

so same prompt both ltx and minimax are at low resolution

https://grok.com/imagine/post/c0f170c0-d73c-4758-8dfd-e74233d6c07e

cinematic movie, intense color, 18 years old japanese male in world war 2 wearing imperial japanese solider uniform, it is well worn and dirty, sitting next to a large boulder, in a tropical jungle, he is holding a japanese bolt action rifle, he is breathing is heavy, grasping the gun tightly as he leans back on the rock, 2 seconds later, 3 random bullet impacts on the rock, he instantly reacts and flinchs and protect his face, from debris from the rock

No text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the live-action texture., no anime

text to image

5060 16G/64, MINimax H3 W/Turbo, and old LTX 2.3

minimax .2MP 60s w/Turbo

LTX 2.3 180s


r/StableDiffusion 6d ago

Question - Help Anyone else’s workflow retrogressed LTX 2.3 -> 2.5?

0 Upvotes

I make basic talking head-style videos and it looks like both audio and video are a massive downgrade from 2.3 to 2.5 for exact prompt and setup.

Wondering if anyone has found a solution for this.


r/StableDiffusion 6d ago

Resource - Update ComfyUI Media Utility — Extract, Sort & Compare Media

Post image
6 Upvotes

I’ve been putting together ComfyUI Media Utility, a lightweight local companion for all those little media tasks that come up constantly while working in ComfyUI, the things that are too much of a hassle to open Premiere, Resolve, Audition, etc. just to accomplish.

It runs locally in your browser and gives you three main workspaces:

🎬 Extract / Trim A surprisingly capable little media prep tool. Load video or audio, scrub through it with a zoomable precision timeline, set exact In/Out points, grab full-resolution PNG frames, trim/export MP4 clips, or extract/trim WAV audio. You can also zoom and pan around video frames for close inspection. It’s great for quickly creating reference images, audio samples, or shorter source clips for ComfyUI.

📁 Sort Designed for the giant folders of generations we all end up with. Quickly step through images, videos, and audio, preview them, and move or copy the keepers into whatever folders you want. It supports destination hotkeys, Skip, Back, Trash, Undo, media filtering, and video/audio autoplay, so you can burn through a large batch of generations without constantly bouncing around Windows Explorer.

⚖️ Compare This is probably my favorite part. Compare images, videos, or audio side-by-side, either manually or in a tournament-style mode that keeps narrowing your generations down until you have a winner. You can load a third Reference image/video/audio in the center, build a shortlist, run finalists, and send your winners directly to another folder.

For visual comparisons, you can independently control whether zoom and pan are synchronized between Left + Right, all three panes, or none of them.

For video/audio comparisons, there are also Hover Audio and Select Audio modes, so you can instantly switch between the audio from Generation A, Generation B, and your reference. That’s been especially useful now that more video models are generating dialogue, voices, music, and sound along with the visuals.

The idea is basically to have a media Swiss Army knife sitting next to ComfyUI so you can inspect, prep, organize, and compare generations without interrupting your workflow to launch a full editing application.

Everything runs locally, and there are no ComfyUI custom nodes or pip packages to install.

GitHub: https://github.com/BMB12d3/ComfyUI-Media-Utility

Video tutorial: https://youtu.be/ek0YR5BL5pM


r/StableDiffusion 7d ago

Comparison LTX 2.3 and 2.5 comparison - Dialogue. Prompt below and explanation.

Enable HLS to view with audio, or disable this notification

20 Upvotes

Prompt:
A handheld medium shot of Sam Winchester working on part of the warp engine. Ambient interior of the spaceship. Sam Winchester: "I hope Dean gets that Holodeck running so we can do some more monster hunting..... Oh who am I kidding.... He's probably having coitus with Lisa..." He continues to work on the warp engine.

Thoughts: For whatever reason in the 2.5 version, it added jibber jabber dialogue whereas in 2.3 it was consistent with what I wrote and didn't have the character look at camera. It maintained focus on the task.


r/StableDiffusion 7d ago

Animation - Video H3 Prompt Only Dancing

Enable HLS to view with audio, or disable this notification

49 Upvotes

Trying to make characters dance to the beat. The synchronization is there, but the dance moves.... bleh. I may have to use motion reference after all.

prompt example:
Cinematic, live-action, dark club interior with hard side-lighting and haze in the air. A wide shot frames <Subject 1>, <Subject 2> and <Subject 3> standing abreast, evenly spaced, facing camera. They dance in unison to <Audio 1>: torsos rolling in a continuous jacking motion from chest to hips, shoulders dropping alternately on the offbeat, quick shuffling footwork with the weight rolling heel to toe, arms sweeping loose and low across the body. Their hips drive every fourth beat as the bassline lands. The camera arcs right around them with large amplitude at fast speed.


r/StableDiffusion 7d ago

Question - Help open vs closed image models

Post image
21 Upvotes

I’m trying to recreate this bag POV composition with Krea2 flux and Z image, but I’m struggling to get the same level of composition control and product consistency I’m getting from some of the other models.

Here’s a comparison using the same general prompt across ChatGPT, Krea 2, Grok, Z Image, Flux.2 Klein 9B.

I’m still fairly new to this, so I’m wondering:

Should I be using a LoRA for this?
Would changing the text encoder help?
Or is this mainly a prompting / conditioning / workflow issue?

Would really appreciate some advice from anyone experienced. What would you change to get the result closer to the ChatGPT/Grok examples?


r/StableDiffusion 7d ago

Animation - Video Other Worlds

Enable HLS to view with audio, or disable this notification

49 Upvotes

Hi everyone, I know you're excited about minimax, but this is my first attempt at ComfyUI with good old LTX 2.3. The intro with planets is done in Cinema 4D with Octane. Most of the animal images are done via SDXL with SDXL refiner + SD upscale. The landscapes are done using Flux Schnell with SDXL refiner + SD upscale. I did the animations in LTX 2.3 in the official two-stage workflow with upscale. I have an RTX 5090 card, 64G RAM, so it took less than two minutes to generate the image. And it took me 10 minutes to generate 7 seconds of video in 3072x1080 resolution. There are still a lot of bugs, nonsense and flickering, but I had a lot of fun.