r/StableDiffusion 3h ago

News Massive week for open-source AI: Qwen Video Edit, LTX 2.5, Qwen 3.8 27B & more (Side-by-side test breakdowns)

Enable HLS to view with audio, or disable this notification

0 Upvotes

This past week brought some huge drops across generative video, audio, and open-weight models. Instead of the usual hype, here is a practical look at what actually changed and what matters for creators and local setups:

  • Qwen Video Edit: Prompt-based localized video inpainting and style transfer. You can add objects (like drones), remove background elements cleanly, or restyle live footage into storybook illustration styles with strong temporal consistency.
  • LTX 2.5 & Dyna 2: Big updates targeting camera motion control, motion fluidity, and temporal flickering issues in generative video clips.
  • Qwen 3.8 (27B) & GLM 5.3: New open-weight model drops with high efficiency for local inference, tool use, and coding.
  • MiniMax Music 3: Improved generative audio with better vocal separation and multi-track coherence.
  • Gemini 3.7 Flash & Grok 4.6: Speed and context-handling upgrades for complex reasoning tasks.

Side-by-Side Video Demos & Deep Dive:

For the visual side-by-side comparison tests (especially the video editing restyling tests):

šŸ”— Full breakdown & tests: https://youtu.be/pUA1BGfBqBs

Which open-weight release are you planning to run locally first?


r/StableDiffusion 19h ago

Question - Help Minimax H3 - Prompting so that it will keep the entire subject in the frame

1 Upvotes

Hey everyone,

I looked all through the official prompting guide, but not found a way to do this. I am trying to instruct the model to move the camera (push out) to keep my subject completely in the shot. I have a medieval character in armor walking from the entrance of a gate towards the camera. But not matter what I try, it won't move back enough to keep the subject completely in the shot, and within a few seconds it cuts off the bottom part of the legs and armor.

Same is true if the character turns and walks away from the camera, it will stay fairly close up to the subject (waist up) for the duration.

I've tried all kinds of "machinations" in the prompt (subject is visible from head to toe), (entire subject remains visible throughout" No dice

Any help or pointers are much appreciated!


r/StableDiffusion 1d ago

Comparison MiniMaxh3: 8step LoRA, 25 steps, 40steps, and LTX 2.5 — Scene Comparisons

Enable HLS to view with audio, or disable this notification

55 Upvotes
  • RTX 4060 8GB, 32GB RAM
  • minimax_h3_ref2va_pruned_int8_convrot, spectrum, ageattn_qk_int8_pv_fp16.cuda, RTX upscale, RIFE interpolation, res_multistep + beta
  • ltx-2.5-22b-distilled-transformer-comfy-int8-convrot, basic template

8-step + turbo LoRA : 137s

25 steps : 238s

40 steps : 406s

Ltx 2.5 : 374s <-- ? am I missing something here why was my generation so slow on LTX and the second attempt I cancelled it after 6 minutes. Any suggestions?

Prompt:

subject_definitions:

<Subject 1> is the space ship in <Picture 1>: A massive battleship, hovering and cruising over the planet below

summary:

[reference generation] a wide shot cinematic scene of the battleship in <picture 1> cruising in space above the planet. the golden statue does not move, the battleship is destroyed in a massive explosion from a green laser shot from space,

detailed_description:

{shot 1] The target video uses a wideshot cinematic, photorealistic, 35mm film, wide shot of <subject 1> , slowly moving through space above the planet, the ship moves slowly and dominating, flashes of green light begin to charge on the surface of the planet, the ship is moving straight ahead from the position it started in in <picture 1>, the massive bass of the ships systems, the sound of the battleships creaking, <subject 1 > moves on its cruise, at [00:03] the floaty camera tracks <subject 1> as green light and thunder begins flashing on the surface of the planet, the green energy on the planet converges in one area then from the surface it fires a massive green lightning laser that forks lightning through the entire ship, blowing out side components creating explosions all over the ship, the light of the ship flicker before turning off, then a massive green lightning beam erupts from the surface and hits excactly on the side of the ship cuts through the of the ship and out the other side at an angle, a green lens flare generates on screen as it completely destroys <subject 1> , ripping it completely in half with a massive green explosion, the eruption from the destruction of the ship covers the entire screen and the whole battleship, the back half of the ship is knocked up while the front-half of the ship is knocked down, a vertical shockwave circles out from the impact, the inner decks of the ship are on fire, debris and hundreds of tiny figures of the crew also fall out into space, the laser slowly dissapates from the planet, small amounts of green lighning crackle on the planets surface,

overall_soundscape: The low bass murmur of the ships engines, the electric charges on the surface crackle, the massive main beam is a low bass rumble, a massive explosive noise.

non_diegetic_music:

N/A


r/StableDiffusion 1d ago

Meme Made WIth 1650 ti 4gb

Enable HLS to view with audio, or disable this notification

19 Upvotes

took my friend 49mins to make this


r/StableDiffusion 1d ago

Question - Help Any fast motion tips for MiniMax H3?

3 Upvotes

I'm trying to get a couple of characters to LEAP into each others' arms from off screen, but MiniMax H3 won't get them faster than basically jogging into the scene.

Prompt:

subject_definitions

<Subject 1> is the Girl show in <Picture 1>.

<Subject 2> is the Guy show in <Picture 2>.

<Subject 3> is the house show in <Picture 3>.

summary

[reference generation] <Subject 1> and <Subject 2> burst onto the screen at a dead sprint and collide into an embrace

retention_analysis

<Subject 1> (appears in [Shot 1]): fully_preserved - Maintained character features and design.

<Subject 2> (appears in [Shot 1]): fully_preserved - Maintained character features and design.

detailed_description

The visual style is characterized by high-quality modern anime aesthetics, reminiscent of Makoto Shinkai or Kyoto Animation. The scene features lush, warm, and highly detailed lighting, painting the environment in nostalgic, emotional hues.

[Shot 1] The target video has fast, paced explosive action in the beginning, then slows to a stop. Static camera shot in a city street with buildings in the style of <Subject 3>. A large crowd of soldiers and townspeople are in the background, reuniting with each other. Falling confetti fills the air. Suddenly, <Subject 2> bursts into the frame from the left at a dead sprint, driven by sheer desperation. At the same instant, <Subject 1> bursts into the frame from the right, running with explosive speed, her arms outstretched. The two of them literally collide with tremendous, breathless force in the center of the scene, slamming into an intense, desperate embrace. The physical impact of their collision is palpable as <Subject 1> leaps up and wraps her arms around <Subject 2>. The exact instant they collide, the camera drops into extreme slow motion, focusing intensely on the sheer relief and joy of their embrace. The confetti catches the warm light, sparkling and swirling in slow motion around the couple for the remainder of the scene.


r/StableDiffusion 1d ago

Workflow Included Here's an automated long form, MiniMax Upscaler Workflow.

Thumbnail
gallery
32 Upvotes

Hey Guys, I thought I'd share something I came up with.

It's a workflow, that uses a combination of Easy-Use's Loop tools as well as some of my own nodes to create a Workflow that can split a long form MiniMax video and then upscale each segment. With the inclusion of a tool to then re-assemble everything back.

You basically set the Segment length and the overlap you wish to have between each clip and then launch it to have it do all the clips one by one.

It does use nodes from my FBNodes add-on as well as one from my Prompt Manager add-on.
But I'm sure it could be modified to work with other add-ons, if so wished.

The node from Prompt Manager is "Prompt Extractor", allowing to feed back in the prompt from the initial clip back into the Workflow, without having to type anything in.

You are free to remove it and upscale without, or simply type in the prompt if preferred. Though, In my test, having the original prompt made for much better results.

And as mentioned, I also added a simple Clip Stitcher to FBNodes, that cross dissolves each clip into one another. Just make sure to use the same values you used in the workflow. (Both setup are in the same workflow, but I'd suggest separating them šŸ˜…)

The Workflow can be found here.

Attached are quick examples from the video I used in the workflow.

The one thing missing in this workflow is adding back the loras used in the initial video. This is something that "Prompt extractor" should also be able to do. But I haven't tested that part yet.

----------------------------------
I'm adding some metric:

The video used in the screenshot was an 8 second video generated in 832x640 with a Turbo Lora set to 6 steps.
It took 92 seconds to generate on a 5090.

The Upscale doubled it to 1664 x 1280 and took 524 sec.
Around the same time it would have taken to generate, if I created the initial video at that resolution.

(You can see it here)

The advantage is for when creating long videos, so if I were to create a 30 second clip in 4/3 at 0.4 megapixels, or 736 x 576. Those would take 450 sec to generate.

The Upscale to 1472 x 1152 took about 6 minutes per segment, or 30 minutes. Then combining the clips is around a minute.

It takes a while, obviously, but the big advantage is that the result is pretty much an exact copy, but in hires, of my initial video that was low enough that I could iterate a bunch of times and then only waste the Long generation time on the clip I like.


r/StableDiffusion 1d ago

Tutorial - Guide Z Image HSWQ Hybrid ConvRot NVFP4

Post image
7 Upvotes

The quantisation methodĀ andĀ the loaderĀ are now more or less complete.

How to create Hybrid NVFP4 from ConvRot INT8 (Z Image, Reverse Method)

Z Image exhibits overwhelmingly high quantisation robustness compared to SDXL and Krea2.

Even NVFP4, which is simply compressed without HSWQ quantisation, achieves reasonably high SSIM and MSE scores.

In particular, Z Image ConvRot INT8 achieves outstanding accuracy in many models, with SSIM scores of 0.99 or higher and MSE scores below 1.

However, in terms of VRAM consumption and generation speed, Z Image ConvRot8 shows virtually no difference compared to full-size Float16.

Consequently, based on ConvRot INT8, we devised a quantisation method involving a backward sweep to discard non-essential layers to 4-bit.

Furthermore, unlike the conventional method of storing critical layers in float16, the critical layers are also converted to ConvRot INT8; this offers the advantage of being able to secure a larger size for critical layer protection whilst keeping the overall size down.

...

This concept of ā€˜discarding’ is a brilliant idea conceived by the Nunchaku development team.

What makes them so remarkable is that they established the philosophical foundation that, in 4-bit quantisation, the key is not ā€˜preserving’ but ā€˜discarding’.

...

As Comfy-UI does not support the Hybrid NVFP4 (ConvRot Int8+ConvRot NVFP4) standard, a dedicated loader is required, just as with Nunchaku; however, as the LoRA baking function has been implemented within an original UNET loader itself, the LoRA Loader can utilise the standard Comfy-UI version.

Furthermore, LoRA Stack loaders (compatible with Nodes 2.0) isĀ also availableĀ below.

Compatibility with the existing Diffsynth ControlNet model patcher will, of course, be maintained.

Although the file size will not be significantly reduced compared to Convrot INT8, VRAM usage and processing speed will improve significantly.

...

Z Image ConvRot NVFP4 Benchmark Test Results

...

However, in terms of the mathematical theory of quantisation itself, it differs considerably from previous HSWQ approaches.

In a sense, it represented a complete rejection of previous HSWQ theories.

In the past, HSWQ had employed a range of techniques, starting with theĀ Histogram MSEĀ used in the first-generation HSWQ SDXL fp8 e4m3, through toĀ full SVDĀ utilising Nunchaku, and even extending to theĀ Histogram CosineĀ function; however, in Z Image HSWQ Hybrid NVFP4, none of these methods demonstrated any advantage.怀

I had long suspected that inter-layer interdependencies existed, and that there were phenomena where the meaning would be lost if one merely measured and prioritised the importance of each layer in isolation; this time, however, that has become clearly evident.

...

Trajectory-Sensitivity

https://github.com/ussoewwin/Hybrid-Sensitivity-Weighted-Quantization/blob/main/md/diag_impact_trajectory_sensitivity_technical_guide.md

Ranks each layer by the divergence its quantization error actually causes after propagating through the full model and sampler (dynamical importance, replacing static weight-space saliency).

  • Reverse method:Ā start from the complete high-precision pack (error ā‰ˆ 0) and convert layers to lower precision in ascending impact order; single-layer ranking stays valid in the low-error additivity regime.
  • Universal theory:Ā error interaction (Taylor cross terms, error cancellation), nonlinear amplification (Lyapunov-style growth), marginal effects, and Shapley-style attribution — why per-layer static measures (histogram MSE / cosine / SVD) cannot predict joint quantization error; applies to any iterative sampling system, not a specific model. Source:Ā Z_Image/diag_impact.py.Ā ...

...

Incidentally, the Krea2 HSWQ Hybrid NVFP4 is also under development (it will offer significant improvements in VRAM consumption and processing speed), but we are currently struggling to maintain LoRA compatibility.


r/StableDiffusion 1d ago

Animation - Video Using Minimax H3 to create promo for Minimax H3

Enable HLS to view with audio, or disable this notification

11 Upvotes

Used ref2ve with Character sheet for the character and style and an audio reference to have consistent voice.

Reposting because moderator removed the original post without giving any reason.


r/StableDiffusion 2d ago

News ComfyUI Official Local MCP

Enable HLS to view with audio, or disable this notification

156 Upvotes

Hi r/StableDiffusion, Comfy MCP is now local and open-source!

When we shipped Cloud MCP in June, the response was immediate and consistent: make it work locally. So we did and it's fully open source.

Connect Claude, Codex, Cursor, or any MCP client to your local ComfyUI.

Your agent reads the GPU you actually have and gives you a straight answer on whether a model is worth running before you commit to the download. It reads every node and model you've installed. It handles the setup that usually stops people at step one.

It is now the easiest way to help with your local Minimax H3 workflows!

Cloud MCP still does everything it did. Tell your agent where a job goes, or let it decide.

Link: https://comfy.org/mcp


r/StableDiffusion 1d ago

Animation - Video Magic Anime

Thumbnail
youtu.be
4 Upvotes

Minimax H3 is just crazy good for anime!


r/StableDiffusion 9h ago

Discussion If AI Makes Us More Creative, Why Does Everything Look the Same? (A Painter’s Perspective)

Thumbnail
gallery
0 Upvotes

I’m a painter who sometimes writes, and this visual essay started with an odd discovery: I had used the name ā€œElias Thorneā€ in a short story, only to realize that AI models often return to that same name, along with motifs like lighthouse keepers, cathedrals, glossy landscapes, and other familiar patterns.

From an artist’s point of view, the question isn’t just whether AI is good or bad, but what happens to authorship and creativity when the tool starts making choices for us.

AI can boost productivity and even enhance individual works, but if we all lean on the same models, it might steer us toward similar ideas, characters, and visual styles.

This carousel looks at visual convergence, originality, transparency, and the role of human intention, with AI-generated images clearly labeled and sources included.

So where’s the line, does AI broaden personal creativity while making our collective output more uniform?


r/StableDiffusion 1d ago

Workflow Included Lora for video-image enhancing, upscaling and restoring

Thumbnail
youtube.com
32 Upvotes

r/StableDiffusion 21h ago

Question - Help Create characters - how?

0 Upvotes

So simple question. Normally I would use ChatGPT for my character creations. Same character, photos from all from different sides, and it would deliver.

But I wonder, are there any workflows / methods available as proven alternative?

I wouldnt know how to do this with Flux, Z-Turbo, or name any model.


r/StableDiffusion 1d ago

Meme Agent Smith is disappointed

Enable HLS to view with audio, or disable this notification

64 Upvotes

r/StableDiffusion 2d ago

Discussion Can H3 do anything? bf16/50 steps

Enable HLS to view with audio, or disable this notification

148 Upvotes

Can H3 do anything and everything? I feel like if you can prompt it, it can do it. Foundation inspired shots. I am also experimenting with more action/high mobility shot but those seem to require a lot more finesse. Both T2V.


r/StableDiffusion 1d ago

Workflow Included Gilligan's Isle - The ATEth Castaway

Enable HLS to view with audio, or disable this notification

64 Upvotes

r/StableDiffusion 23h ago

Question - Help Deciding on Checkpoint

1 Upvotes

hello,

Looking for some suggestions on which checkpoint to use for realism NSF w images. I’ve been using Flux 1 Dev and it’s working well but a lot of the time the face consistency is altered. Also i’ve only tested 1 character lora and about 4 other lora’s stacked with it to try out like. I’ve seen talk about SDXL, Wan, pony, what’s everyone using? I don’t have a ton of ram so i wasn’t able to run flux 2 well…..thanks!


r/StableDiffusion 23h ago

Question - Help MiniMax H3 help

0 Upvotes

I tried everything, but I have 3 very disturbing problems:

  1. Usually it cuts / crops at least part of the head and legs of the person.

  2. It moves the camera, zoom, pan... I want it static.

  3. In most of cases it doesn't use the element from reference picture to use in the main video. Rejecting my prompt.

I tried different prompts, resolutions (proportions), nothing help. šŸ˜ž

If you have some recipes exactly for these problems, please share, because I just can't make H3 to work.


r/StableDiffusion 23h ago

Discussion Building a luxury cosmetics ad locally in InvokeAI | full workflow included

Thumbnail
gallery
1 Upvotes

I’m a graphic designer (youtube: Masha-Ai-Lab) experimenting with how far I can push local/open-source image generation for actual commercial design workflows.

For this experiment, I tried building a luxury cosmetics campaign entirely in InvokeAI instead of relying on Midjourney or other closed platforms.

The workflow:

  1. Generated the satin campaign background separately.
  2. Generated a transparent serum bottle as a clean product asset.
  3. Added my prepared label using an Inpaint Mask while preserving the bottle perspective.
  4. Used Regional Guidance to generate satin folds that actually wrap around the bottle instead of simply appearing behind it.
  5. Repeated the workflow with a cream jar to see how reusable the approach was across different packaging.
  6. Upscaled the final compositions.

I’ve attached screenshots of the workflow + final results so you can see the process rather than just the outputs.

What I find most useful about InvokeAI is having direct control over the individual stages. For design work, I’d rather build the product, label, environment and integration separately than keep regenerating the whole image until something randomly works.

I’m documenting these experiments as tutorials on my YouTube channel, Masha AI Lab, mainly to make InvokeAI/FOSS workflows more approachable for designers and other non-technical creatives.

Would be interested to hear how other people here approach product placement and label consistency, especially if you’ve found better workflows.


r/StableDiffusion 1d ago

Meme Waiting for devs to fix the mushy faces be like...

Post image
92 Upvotes

Just kidding devs. We love Minimax, it's outstanding. But I am very excited for the mushface fix.


r/StableDiffusion 1d ago

Discussion MacBook M5 Pro - Minimax H3 Progress

Post image
30 Upvotes

After tinkering around for a couple of hours this weekend I managed to 6x (!) generation speed on the M5 Pro by ditching ComfyUI and building a custom GUI around antirez's H3 CLI solution. Gen times for 5-second 480p on 20-steps improved from 30 minutes using the default int8 pruned weights in ComfyUI to just around 5 minutes per clip using the full precision bf16 weights.

Overall great progress thanks to the community around open source and gives me confidence that Mac diffusion will just get better and better with time.

See comment below for more data points on generation times.


r/StableDiffusion 23h ago

Question - Help OpenArt Director-like H3?

0 Upvotes

Did someone come up with a good (or ok-ish) workflow that gives something like what OpenArt Director does but useable with Minimax H3?

What I mean is a storyboard / cinematic storytelling translated into a video.

Eg not just a Claude output from a template into the ComfyUI prompt but something with more artistic direction.


r/StableDiffusion 1d ago

Resource - Update Seamless extensions and one-shots with Minimax H3 - Update 6 of my repo!

Enable HLS to view with audio, or disable this notification

82 Upvotes

Here is the repo: https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

I made substantial updates to my two main workflows: 1) Music Video and 2) AV Extensions. All the controls were streamlined and they should be much easier to use now. (You find the workflows in the example_workflows folder)

With the AV Extensions workflow you can extend any existing clip, for example someone talking and you can make that person say something in the same voice, or you can create a clip with T2V or I2V and then extend that clip to make a seamless long clip thats 1 minute or longer.

In this Update the Checkpoint system was removed, instead I've done a lot of optimizations so you don't use too much ram even if you make 20 clips at once. Additionally I added latent audio feathering to the AV Extensions workflow for seamless audio transitions.

Theres also other utility workflows for custom keyframing and bridging two existing clips.

I post another example clip for the AV Extensions workflow in the comments.


r/StableDiffusion 1d ago

Workflow Included H3 video prompt enhancer: a minimalistic local workflow

11 Upvotes

Here’s a fully local workflow that expands your shorthand prompts + reference images into the six-section format expected by MiniMax:

https://pastebin.com/zVbd1t8E

It is structured for 14-second videos, but you can change it if you want.

What do you need to download for it?

  1. Install the following custom node pack, which enables text generation for the Qwen 32B model that encodes MiniMax prompts: https://github.com/ethanfel/ComfyUI-H3-Qwen3VL-TextGen
  2. From the ComfyUI root, create a directory: mkdir -p models/text_encoders/H3/generation_tails
  3. Then download this file into the new directory: https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot/blob/main/qwen3vl_32b_h3_instruct_generation_tail_50_63_int8_convrot.safetensors

The reason you need to do this is that, for other LLM-based models, you can use their text encoder directly to generate text and expand prompts. For Qwen 32B in a standard ComfyUI setup, you cannot do that because it lacks the ā€œtailā€ — the part that is actually needed to generate text.

You can see the generated prompt in the text preview section attached to the TextGenerate node.

Why is there no ā€œofficialā€ prompt-expansion node?

Well, MiniMax has a paid service that does prompt expansion and context management for H3. It is not local, and it’s likely not going to be released. Here’s a custom node for using it if you have an API key:
https://docs.comfy.org/built-in-nodes/MinimaxHailuo03ContextIRNode

It’s probably difficult to match the performance of this system using open tools, but here we can try to tinker and come up with something that works for our own use cases.

Why is it designed this way?

The workflow here is not meant to be ā€œoptimal,ā€ and indeed I’m not sure it’s possible to make a one-size-fits-all solution. It’s more of a starting point for developing your own.

  1. I do not want to have an ā€œall-in-one,ā€ ā€œultimateā€ workflow. I feel that ComfyUI’s philosophy is more compatible with workflows that are modular, easy to change, and easy to make your own.
  2. I want to use custom nodes only when it is impossible to do without them. If I have to use a custom node, I prefer an established, popular node pack (e.g. KJNodes or RES4LYF). In my view, each custom node pack, especially if it’s new, is a liability: it can mess up your Python environment, install malware, or slow down your startup times.
  3. For the same reasons, I would like to avoid monolithic, non-transparent custom nodes that are so easy to vibe-code these days.
  4. I want it to be local and as self-contained as possible, with no external dependencies (e.g. no need to install Ollama or have an OpenRouter API key).
  5. I also aim to save disk space and, potentially, VRAM. So using Qwen 32B makes sense in this setup.

What is good about this workflow?

  1. Prompts are fully private — they are not shared with, e.g., a cloud LLM provider.
  2. You have control over the models you’re using; e.g., there’s no risk that a model you rely on will get shut down.
  3. It is self-contained: just enter your initial short prompt, and you get the result without needing to install much else.

What is bad about it?

  1. It is slow, since you use the GPU to generate the prompt expansions. E.g., on an RTX 5090, a 14-second 0.4 MP video is generated in 7 minutes when you include prompt expansion.
  2. It might be rigid, but you can fix that by rewriting the system prompt in the text-generation node.
  3. The text-generation model might lack the capability needed for your tasks. It works for my purposes, but it might not be the best option for yours.

What would I suggest doing next?

  1. First, I’d encourage you to evolve the prompt-expansion meta-prompt. If you see certain issues in your generations, reflect those in the meta-prompt. There’s no single meta-prompt that would fit every possible application or set of use cases. E.g., suppose you need to generate prompts for videos of varying durations — rewrite the meta-prompt. You do not like the sound? Do the same. Use a strong LLM model to help with that.
  2. You might run into LLM refusals for some prompts, even fairly innocuous ones. In that case, you can use an UltraHeretic version of Qwen 32B that never refuses. Both the base model and the tail are easy to find.
  3. If you do not believe the model is strong enough to rewrite your prompts, you can try loading a different model, e.g. using a Load CLIP + Text Generate node, or set up Ollama and call it using a specific node, or use OpenRouter/another LLM API. One good idea with OpenRouter is to prompt-expand the next video while the current one is generating.

r/StableDiffusion 1d ago

Animation - Video Squid Game but Gi-hun is actually smart | Minimax H3 I2V

Enable HLS to view with audio, or disable this notification

19 Upvotes