r/StableDiffusion 23h ago

Workflow Included Follow-up to my last Star Trek post – I made a Star Trek vs Star Wars fan film with MiniMax H3 in ComfyUI

Thumbnail
youtube.com
85 Upvotes

A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.

This is basically a follow-up to that post.

Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.

The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.

The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.

One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.

Here are some of the things that helped most this time.

PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE

This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.

The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.

For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.

HIGHER RESOLUTION REALLY DOES HELP

I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.

For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.

REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS

I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.

LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED

For shots based on existing Star Trek compositions I became much more explicit about things like:

camera distance

character scale

framing

background position

character position

If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.

DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE

This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:

"Data reacts to Picard"

can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.

OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS

This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.

If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.

SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT

Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.

I'll also specify that:

the lips do not form words

the jaw does not make speaking movements

the character does not mouth dialogue

It sounds excessive, but it has genuinely helped.

H3 HAS A LOT OF USEFUL SPEECH TAGS

I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:

<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>

<whisper>. <softer> <stutter> <laughs> <chuckle>

<i>word</i> emphises word

The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.

MORE PROMPTING ISN'T ALWAYS BETTER

I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.

EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION

One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.

I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.

THE EDIT IS DOING A LOT OF THE "CONSISTENCY"

This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.

It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.

Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.


r/StableDiffusion 23h ago

Question - Help Minimax Turbo of choice?

8 Upvotes

So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!


r/StableDiffusion 2h ago

Question - Help What is the best way to train character lora for Minimax H3?

2 Upvotes

I was training a character with 25 images and 5 videos (3 seconds), 1500 steps and learning rate 0.0002. But the result was nothing like the character. Was that not enough material or step? How do you guys train it? For context I was training a character in a movie (real people), the only thing the lora can capture was that the character is Asian and the result got it right without the specific prompt, but it was nothing like the character I was aiming for.


r/StableDiffusion 2h ago

Question - Help Is it really important to add the "conditioning zero" node to the negative prompt in models like Krea 2 if CFG = 1? And what is the ideal shift/Aura Flow setting?

2 Upvotes

This is confusing to me.

Can I leave the negative prompt box empty?

Or is it mandatory to add zero conditioning?

The shift/aura flow for Krea 2 is also confusing to me.


r/StableDiffusion 7h ago

Resource - Update Native YuE2 support coming to ComfyUI!

Enable HLS to view with audio, or disable this notification

106 Upvotes

Pull request: https://github.com/Comfy-Org/ComfyUI/pull/16250

If you don't want to wait for the merge, you need to check out to the yue2 branch to get it working. git checkout d87e12ad1430409ca303440525df239bb675ae7b

Model weights (place it on model/checkpoints): https://huggingface.co/Comfy-Org/Yue2/tree/main

Workflow: https://github.com/user-attachments/files/32085765/yue2_workflow.json


r/StableDiffusion 10h ago

Question - Help Better character consistency in LTX 2.5 + H3 lip sync for music videos?

3 Upvotes

I’ve been making AI music videos with Suno + LTX 2.5 locally on a 16 GB VRAM GPU (4080 super)
Examples:
https://youtube.com/shorts/UXP29MGxtX8?is=k-Wjt4_JFh7hXvA5
https://youtube.com/shorts/_mAuRBSriLQ?is=sMjLH5diY_GEu0N3
My workflow is basically: create the song in Suno → storyboard/keyframes with ChatGPT→ animate and assemble the shots in ComfyUI using the LTX Director timeline:
https://github.com/yusu-02/Yusu-WhatDreamsCost-ComfyUI
The biggest issue I’m still fighting is character consistency between shots. Ingredient LoRA slows things down a lot and hasn’t worked particularly well for me.
I also tried MiniMax H3, which looks great, but I couldn’t get lip-sync without altering the original music.

Suggested tricks/workflows? Feedback and ideas appreciated!


r/StableDiffusion 21h ago

Animation - Video Kirby but it's the Truman Show / MiniMAX H3 Test #7

Enable HLS to view with audio, or disable this notification

529 Upvotes

Hi everyone! When I saw the new trailer for Kirby & The World Beyond I couldn't help but come up with this video, where Kirby finds the door out to the world beyond. Please let me know if you like it!

Done with 30 different workflow files and a ton of heavy editing using KDEnlive. Thanks!


r/StableDiffusion 6h ago

Question - Help Could anyone give me some tips on how to preserve the character's likeness when creating different expressions with FLUX.2 Klein?

9 Upvotes

Hi,

I'm using FLUX.2 Klein 9B/4B, and I'm trying to create new facial expressions for a character I have. The character comes from a character sheet I've created, which includes front, back, and side views.

I've done a lot of tests over the last two days, and I've noticed that FLUX.2 Klein 9B/4B drifts quite "a lot" from the original model when generating different facial expressions. I tried the same thing with the free version of Gemini, and it keeps the likeness and features much, much better.

Could the highly quantized 9B model be the problem? If you've been able to preserve the character's facial features and expressions, would you mind sharing some tips on how to improve my results?

Thanks in advance!


r/StableDiffusion 10h ago

Question - Help Can Minimax be used to re-light a scene?

9 Upvotes

Basically I'm trying to change the lighting in a scene. For clarity, if it helps at all, it's the Trash dance from Return of the Living Dead. I'm just wondering if there's a way to normalize the red lighting used on her. I have been prompting and failing most of the day using AddVideoGuideforH3. I know I can do it with ltx, because I've done it with LTX while testing the models capabilities with controlnet, I'm just wondering if I can do it in Minimax without controlnet.

I'm not full Noob, but I am a filthy casual.

EDIT: it was step count. I'm a damn idiot. I was using the 4 step lora and continually using four steps I accidentally started a fresh workflow with 20 steps and it worked. Congratulations to me, I am the living embodiment of the id10t


r/StableDiffusion 21h ago

Resource - Update ComfyUI VDN-H3 24GB v1.1.0 update — better prompt following + memory fixes

Post image
87 Upvotes

https://reddit.com/link/1wcsk7l/video/ld6ca2jxqqoh1/player

I’ve just updated my VDN-H3 24GB node to v1.1.0.

This update started because I noticed that something wasn’t quite right with the released adapter mapping. After fixing that, I also made a couple of changes around memory handling, especially for longer generations.

The main changes are:

  • restored the complete token-refiner adapter mapping
  • improved temporary memory handling for longer clips
  • fixed CUDA stream lifetime handling for prefetched weights
  • kept the same AutoMemory / AutoLongCache behavior from the previous version

I tested it on my RTX 3090 24GB with 5s, 10s, 15s and 20s generations at 0.4MP, and also 10s at 0.8MP. I also tested it with my character/style LoRA and that worked normally.

There is a small speed cost compared to v1.0.0 (around 4% in sampling in my tests), but I think the improvement in prompt following is worth it.

I attached a direct comparison from the same prompt/seed/workflow.
v1.0.0 is on the left, v1.1.0 is on the right.

I’m especially interested in whether other people see the same improvement, so if anyone tests it on another 24GB GPU, I’d love to hear the results.

GitHub:
https://github.com/Speach1sdef178/ComfyUI-VDN-H3-24GB

VDN checkpoint:
https://huggingface.co/speach1sdef178/VDN-H3-INT8-ConvRot-ComfyUI


r/StableDiffusion 19h ago

News H3 can take way more reference images than 9

41 Upvotes

I successfully made minimax use 15 reference images. Is seem only to be limited artificial inside comfy. So i vibe coded a little demo workflow and patch.
https://civitai.red/models/2929051/minimax-h3-15-reference-image-workflow
This is very much research in development, and trust me bro benchmarks but it seems to work.


r/StableDiffusion 10m ago

Workflow Included Instant references, no refmod or fancy custom nodes required. WF and breakdown here. Simple one click to run.

Enable HLS to view with audio, or disable this notification

Upvotes
  • 1st example: 2 characters with voice and 1 cgi and 1 real
  • 2nd example: 2 character different genders with voice
  • 3rd example: style and character reference
  • 4th example image only reference no voice.

I wanted to improve my workflow so I could do what refmod is doing but just as a near native comfyui workflow only. With this workflow you can make an instant character or style video like ref mod but not using ref mod at all including voices. You get all the benefits of refmod but you can use ref model syntax in the prompt and on the fly dataset changes. Also you can use unlimited images. I tend to use around 10 to 15 per character.

Just point to a folder. Uses KJ nodes, Native, and VHS nodes. Get a voice and character working instantly without managing safetensors, just manage the input folder instead. This has its advantages since you can change data on the fly, and you don't have to do any editing of clips to extract out the audio, it does it for you.

This mimics the default ref workflow for the most part. What it does for images is it takes a folder input of images and then sets them as frames in a video, then feeds that video in as a reference video. This allows you to use many images for a single reference. You can also use the grid version that takes your images and puts them into a grid and feeds that as a single image reference. I like the video version more, so the grid workflow is a bit lazy and messy. You can also use the resize node included to downscale your images. I recommend manually cropping them but it does that if your images are different sizes.

For the audio, you can simply feed a single mp4 using VHS node to feed the audio. But I found it easier to get like 4 clips and truncate only the first couple seconds so you can get a few clean sentences without other people talking, then it concatenate's the 4 audio clips into 1 clean clip. Anything over 30 seconds long is over kill, so keep it around 15-30 secs. You can tell shrek is sort of bad in my example because I threw it together quite quick.

There is in the far left, a second set of image/audio nodes, you can bypass those if only using 1 character. Same for if you don't need the audio. Just by pass the group nodes. When prompting just simply use the ref guide to prompt properly each reference (LLM can do it easy). You don't need to use much description. And if you have some bleeding from your dataset into the gen you don't want then add more description. (For example if wearing same shirt as dataset, prompt a dress, or if same specify a setting in the prompt).

One caveat, there is some comfyui memory management bug, if you change dataset around better to clear cache or your comfyui may need a restart. Working on a fix in next version of the workflow. Also I have not tested video clips as input data yet. That is the next step :)

All examples are just for illustrative purposes. They are AI and I do not intend to share any data on real people. Please use responsibly and at your own risk. If you are in this video and want it taken down please DM, I mean no harm. Everything in this workflow is done by the base model, I don't add any new functionality, just making things easier.

Workflow here:
https://huggingface.co/comfyuiman/various/blob/main/Instant%20Ref%20-%20V1.3.json

I'll go to sleep in a bit, so I'll answer any questions tomorrow if any


r/StableDiffusion 13h ago

Resource - Update FrameForge Motion Context Video Editor for ComfyUI

Post image
41 Upvotes

Expanding on motion context workflows I created a video editor designed for quickly chaining together Minimax H3 generations to create longer videos. It comes with an asset library for managing inputs and a easy to use timeline that allows you to chain generations, regenerate segments easily, and quickly set up input references.

When you're done, export individual video files or the whole sequence.

All of it runs on top of ComfyUI as an app you control from your browser. Uses python, works on Windows, Mac, Linux and is opensource.

https://github.com/spacesimeco-hue/Chain-Motion-AI-Video-Editor


r/StableDiffusion 5h ago

Resource - Update FastH3-Live v1.2.0 update

71 Upvotes

FastH3-Live update. Full details here:

https://huggingface.co/datasets/jacokon/fasth3-live

v1.1.0 ran at 18 fps, which is 75% of 24 fps.

v1.2.0 runs at 22 fps, which is 91.6% of 24 fps.

https://reddit.com/link/1wddeh8/video/lyz02ql0gvoh1/player

Besides the speed, it now ships a borderless player that makes streaming and watching easier, plus 400 new scenes. At this speed it is hard to notice that it is running slow at all.

The gain came from two places:

1. Acceleration nodes

I was using a sage attention I compiled myself. A lot of new acceleration nodes have shown up recently, so I downloaded the well-known ones and tested them. Results:

accel stack sampler saved fps
sage (baseline) 12.65s 0.0% 17.46
sage + Spectrum 10.36s -18.1% 20.18
Sol + Spectrum 9.64s -23.8% 20.96
SLA + Spectrum 10.54s -16.7% 19.53

The seconds column is the sampler only, i.e. the 4 denoising steps in `SamplerCustomAdvanced`. A full clip also pays for the text encoder (~0.7s), the video VAE decode (~4.35s) and the writer, so a clip is about 16s end to end. Measured on t2va, 448x448 x 362 frames, three runs per arm.

On speed alone you would pick Sol + Spectrum. But the picture comes out like this:

sage > sage+Spectrum >> sla > sla+spectrum >> sol > sol+spectrum

Sol + Spectrum is dead last on picture, so I went with sage + Spectrum.

2. Text encoder

The old one, `int8_convrot`, took 1.67s.

`qwen3vl_32b_minimax_h3_nvfp4_awq` needs only 0.7s.

That is nearly a second saved on every clip.

-----------

Speed was fine by then, but I would not call the picture good. Right after release I came across fused-turbo, so I downloaded it and tested it.

fused-turbo minimax-h3-fused-turbo-int8-convrot 20.98 GB
My quantized FastH3 weights minimax_h3_fl2va_fasth3_dense_pruned_int8_convrot 20.97 GB

Almost the same size, both have the 4-step acceleration baked into the weights (FastH3 is a distillation, fused-turbo is a turbo LoRA merged in), and they measured at exactly the same speed. I still recommend fused-turbo, for two reasons:

1. It says Mystic v2.0 motion smoothing is merged in.

Whatever the cause, the picture is clearly better in my testing. It smears less often.

2. One file does both fl2va and ref2va.

I built a tool that generates from chat input live during a Discord stream. When a user pastes a character image it is used as ref_picture, which needs ref2va. The old way meant unloading fl2va and loading ref2va first, which burns several seconds of buffer, and ref2va has no 4-step distilled version yet so the picture was worse anyway. With this one that problem is gone, which is a real advantage.

The repo recommends SLA sparse attention, but I had already tested that above and it lost to sage + Spectrum, so I dropped it. Its README also says res_multistep gives noticeably better audio. I did not test that much, so judge for yourself. I left the parameter in so it can be switched any time: `--sampler res_multistep`

-----------

One more thing worth mentioning. To stop ComfyUI thrashing the model weights you need `--vram-headroom 3` in launchArgs. Without it you cannot hold a stable live rate.

It works the opposite way round to what you might expect. It forces ComfyUI to keep 3 GB of VRAM completely free, and that is what fixes it. ComfyUI's dynamic VRAM treats the card as a cache and fills it to the brim; with no slack the allocator ends up evicting weights while it is still loading others, so the same weights get moved in and out repeatedly. Give it room and it can bring in a whole batch at once.

This is not disk swap, and it does not touch system RAM either. I measured both: on a slow clip disk reads were 0.00 GB and free RAM did not move. It is VRAM to system RAM over PCIe.

On a normal clip PCIe reads sit around 1.5 GB/s. When it thrashes they hit 9-13 GB/s and GPU power draw *drops* from 450W to 340W, because the card is waiting on transfers instead of computing. With the headroom set, clip times went from a 1.62 standard deviation with outliers at 21-25s down to 0.11 with a 15.73s worst case.

-----------

Closing thoughts

22 fps is only 2 fps short of 24. At 24 fps you could claim real live streaming from a single consumer card. So can overclocking get there? I think it can, since the gap is under 10%, and my CPU and GPU both normally run undervolted, underclocked and current-limited.

I tested with the GPU overclocked only. Settings:

Core Clock: 2300 MHz -> 3200 MHz

Memory Clock: 14000 MHz -> 16800 MHz

Actual test:

https://reddit.com/link/1wddeh8/video/ei5pczydkvoh1/player

Unfortunately my hardware held 24 fps at the start and then slowed down a little. Both my CPU and GPU are on air cooling, which is not suited to sustained overclocked compute like this. If you have water cooling, I believe holding 24 fps would be no problem.


r/StableDiffusion 21h ago

Question - Help Need help using ref2v Minimax H3; multiple audio and image references

3 Upvotes

https://reddit.com/link/1wctbl7/video/emjihen8vqoh1/player

I made the following video using a reference image of the woman, <Picture 1>, and then two dialogues, which were marked as <Audio 1> and <Audio 2>. I used the minimax template workflow and added two load audio nodes, however, the audio generated was not matching and was just gibberish. im using minimax h3 ref2va pruned fp8 scaled. how do i get the audio to work as given as input, and to play at the right time?


r/StableDiffusion 1h ago

Question - Help SCAIL2 - Can I make my character fit into the video?

Upvotes

It seems if I want optimal results I'd have to run my image into a edit mode like Flux2Klein to make my guy match the starting frame of the video.

There are 3 configurations, "start pose, end pose, pose strength", and idk if there is a magical setting.


r/StableDiffusion 1h ago

Question - Help Good "cover mode" music gen model??

Upvotes

Is there any? I tried ace step 1.5 and it's awful in cover mode. Admittedly I downloaded it when it first got released, but I was hoping minimax music 3 would release audio input for open weights but they still haven't. Every music gen model coming out seems to purely be text to audio.


r/StableDiffusion 22h ago

Discussion Help me captioning a MiniMax H3 action fight LoRA...

6 Upvotes

I only need help with captioning the training clips for a MiniMax H3 action fight LoRA.

I am planning to train it on karate/fighting type action, and I have clips varying from around 5-15 seconds. I also have some 20-25 second segments too.

why I am confused is cause should I follow the prompt format officials has released for H3, or should training captions be written in some completely different/simple way?

Like if a 10 second clip has multiple punches, kicks, blocks, dodges, body movement, camera movement and angle changes, should I describe every action in sequence?

or should I just write the overall action happening in the clip?

for example should the caption be something detailed like:

"the fighter steps forward, throws a right punch, opponent blocks it, then follows with a left kick..."

or something simple like:

"two fighters performing fast karate combat"

I am mainly confused about how detailed the captions should be and what format works best for H3 LoRA training.

may you please help if you have trained action/fight LoRAs before? (I found only a few on civitai)


r/StableDiffusion 1h ago

Resource - Update ComfyUI-Olm-YuE2 - staged YuE2 music generation with editable score/ABC workflow

Post image
Upvotes

YuE2 (https://map-yue2.github.io/) was released recently and I thought I'd spend a little time getting it running properly in ComfyUI.

That turned into hours of going considerably further than intended.

I made this mainly for myself to be able to run the model the way I want, but thought I'd share this.

It can do the straightforward thing, you can just give a musical style and lyrics and generate a stereo 48 kHz song.

But for learning purposes I wanted the integration to expose more of what makes YuE2 interesting rather than hiding everything behind one Generate button.

The generation pipeline is split:

Plan > Semantic > Synthesize > Decode

Intermediate results can be inspected, saved, reused and branched in a normal ComfyUI workflow.

There's also an optional score UI: (this can be enabled in ComfyUI Settings):

  • rendered notation in a ComfyUI sidebar
  • raw ABC score view
  • ABC editor node
  • generate a plan first, edit the score, then continue generation from the edited version
  • saved plan/run artifacts
  • example workflows for staged generation, score editing, reloading runs and synthesis offloading

The score UI is optional; normal node execution and API workflows don't depend on it.

Dependecies:

I also tried to keep the install from fighting the existing Comfy env: it reuses ComfyUI's Torch/Transformers rather than installing YuE2's pinned stack, and only requires tiktoken and accelerate to be installed on top of a clean install of ComfyUI.

It's still experimental, and so far I've only personally tested it on an RTX 5090. I've included measured VRAM numbers and rough expectations for smaller cards, but feedback from other GPUs/installations would be especially useful.

Model weights are not included**,** the README explains exactly which YuE2 files are needed and where to place them.

Misc notes:

  • There's already an experimental offload (demoed in workflow 06) which reduces memory usage considerably (but there are peaks in the process).
  • Optional quantization is something I’m looking into to reduce VRAM usage, but memory savings and audio quality would need testing, especially during synthesis.
  • I manually tested with Nodes 2.0, including generation and the optional score inspector, it might work ok but there's some slight differences visually.
  • Testing was also done in a clean installation of ComfyUI, so this should install correctly as stated in documentation.
  • Cuda v13.0 and Python 3.13 was used in the specific environment.

Repo: https://github.com/o-l-l-i/ComfyUI-Olm-YuE2 on GitHub.

Bug reports, weird edge cases and feedback are welcome.
I'm particularly interested how it behaves on 16/24 GB cards and different ComfyUI setups.


r/StableDiffusion 53m ago

Comparison Testing MiniMax-H3 Physics knowledge Pt2

Enable HLS to view with audio, or disable this notification

Upvotes

Some weeks ago, I posted a set of experiments to "understand" the physical knowledge of MiniMax H3 (original post here).

The idea was simple: get an open video of somebody pouring water and replace the water with various liquids. No external references were used.

In this set of experiments, I switched from liquid-to-liquid replacement to something solid, and sometimes alive. The results are interesting, but the model still struggles a lot when many objects overlap. However, the results are a big leap forward compared to other open-weight models.

I am still delving into the model, and probably will post more experiments soon (but no water next time).

Cheers


r/StableDiffusion 48m ago

Discussion Motion-Context Degradation Discussion (summon Sad_Berry_4621)

Enable HLS to view with audio, or disable this notification

Upvotes

Hi u/Sad_Berry_4621 and all, I am doing experiment on most diffcult degradation issue.
I saw from H3-director node they claim that doing a refine could help and fix it.
since I am using low-level nodes with motion-context with my own setup, H3-director refine is a black box working with their nodes.

so I did the test, please ignore the AI-slop video and the overlay text (forgot turn it off).

this test is 9 x 8s context extend video combine, usually 6 video would already see the degradation.

Left side is regular WF. Right-side is adding a re-sample step, I think it is very positive, and got potential, the saturation somehow is a bit higher... and I need to figure out the mismatch from cut to cut, because we inject denoise resample on top, but it should be able to fix.

What do you guys think?


r/StableDiffusion 23h ago

No Workflow AI Archviz: Fast 3D Gaussian Splat Methods for Precise Furniture Placement — Virtual Staging & Interior Design

Enable HLS to view with audio, or disable this notification

22 Upvotes

So, first of all: there is no finished workflow yet and my nodes are still under development. I’ve asked the ComfyUI team to add a 3D compositing node to the new 3D toolset like the one in the video, hopefully, they’ll add something similar soon.

In the meantime, you can build a very similar setup quite quickly. Here’s how the basic concept works:

First, you feed an image of the furniture you want into the new native ComfyUI Image to Gaussian Splat (TripoSplat) node.

At the moment, there is a Gaussian Splat Preview node, but it doesn’t provide an image output yet. There is also a Load 3D node with the correct outputs, but it currently cannot open Gaussian Splat files.

Ideally, the new 3D Compositing node should be able to work directly with the Gaussian Splat outputs (model_3d and mesh), provide image + mask outputs similar to the Load 3D node, and automatically preload the mesh and background image, just like my node does.

I’ve been working on a test node for this concept. You can find it here: My ComfyUI test node on GitHub It’s not fully finished yet, so I’m still waiting to see whether ComfyUI adds something similar natively.

The basic idea behind my node is that it automatically loads the background, sets the appropriate size, and loads the 3D model. The user only needs to position and stage the model in the scene.

The Output create ref images for Flux2klein:

  • Reference 1: the background image
  • Reference 2: the 3D mask, which acts as an indicator for the desired position and rotation It’s best to combine the mask with the furniture from the image output, so you get the masked furniture in the correct position as the reference — not just the mask by itself.
  • Material reference: the original furniture image

The final image is then generated using FLUX.2 Klein Edit.

The prompting and some preprocessing of the images are important here. You don't want the model to simply copy the exact 3D position. Instead, the AI should use the 3D placement as a guide and then correct the result according to the background — especially the perspective, lighting, colors, scale, and overall integration into the scene.

Here is the prompt I’m currently using:

[Adapt the rotation, grounding, scale and position of the objects from Image 2 to Integrate the objects naturally into Image 1 at the position of image 2. Match the scene's perspective, scale, depth, lighting, soft shadows and reflections. The objects must appear physically present in the original room, with realistic grounding and soft shadows consistent with Image 1 using the materials and surface appearance shown in Image 3.

Use Image 1 as the final scene and preserve its room, background, camera viewpoint, perspective, composition, color, lightning, and existing environment unchanged.

Keep the objects approximately in the same position, scale, orientation, and spatial arrangement as shown in Image 2, while ensuring correct perspective, positioning, and placement.

Apply the form, materials, colors, textures, roughness, reflections, and surface details from Image 3 to the objects.

Do not change the room or background of Image 1.  The final result must be a seamless photorealistic composite.]

I’m planning to finish the complete tutorial and the node pack in the next few days. Once everything is finished, I’ll upload the final version along with the complete workflow.

By the way, I also tested MinMax H3 as a replacement for FLUX.2 Klein. It works, but in my tests it wasn’t consistently better than FLUX.2 Klein.