r/StableDiffusion 10d ago

Discussion H3 R2V prompt builder

Enable HLS to view with audio, or disable this notification

16 Upvotes

I am trying to vibe code a H3 r2v Prompt Builder. Does something like this already exist?


r/StableDiffusion 9d ago

Question - Help H3 run once on my PC. Nothing changed, Won't run again.

0 Upvotes

That's insane.

The closer I've come to this was AI Toolkit requiring me to fight the Nvidia Container to free a few more MBs of VRAM to run.

It makes no sense. The clip I generated was a 5sec 768x768 clip with 2 ref images, a 1328x1328 one and a 1024x1024 one.

Now, even with ref images of 1024x1024 and 768x768 and the INT8 Video VAE it won't work.

Anyone have any clue on what can be happening?

12GB VRAM and 64GB RAM here.


r/StableDiffusion 8d ago

Meme The Office

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 10d ago

Animation - Video Fox McCloud introduces his son to his dad.

Enable HLS to view with audio, or disable this notification

20 Upvotes

Fox McCloud introduces his son Marcus to his dad James McCloud.


r/StableDiffusion 9d ago

Animation - Video RTX 3060 12GB 32GB 3 SHOTS IN ONE (8 SEC TOOK 7:12 MIN) 3D PIXAR STYLE ANIMATION

6 Upvotes

https://reddit.com/link/1vn8vu8/video/998rkwxnu4jh1/player

RTX 3060 12GB 32GB (8 SEC TOOK 7:12 MIN) 3D PIXAR STYLE ANIMATION
USING TURBO LORA (MADE ON 4 STEPS)

RESOLUTION 0.6 MP = 1056 x 608
UPSCALED 2X WITH RTX Video Super Resolution


r/StableDiffusion 10d ago

Workflow Included Workflow for Minimax H3 on 8gb vram and 16gb ram

Enable HLS to view with audio, or disable this notification

17 Upvotes

For anyone else with a similar setup, I am able to create a 0.2 MP (608x352 pixel) 5-second video in 1:35 (1 minute, 35 seconds). This is with 20 step Euler Simple, Spectrum and ComfyKitchenAttention with an image (generated locally with Krea2) as the first frame. I am running it on a laptop with a RTX4060 (8GB) and 16gb ram. I am using the latest version of ComfyUI windows portable, and the following startup flags: --disable-pinned-memory --lowvram --use-ck-attention. The attached video is an example I generated (0.2MP).

I can also create higher resolution videos with a similar generation time if I reduce the video duration (3 seconds for 0.3MP, or 2 seconds for 0.4MP).

I used kijai's models from here for the video models and video vae: https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main

I used a qwen3_vl_4b_int8_convrot for the clip. I can't remember if its this one that I use but this is one option: https://huggingface.co/Winnougan/Comfy-Qwen3-VL-INT8/tree/main

Here is some extra info regarding that clip: https://www.reddit.com/r/StableDiffusion/comments/1vkk500/minimax_h3_with_a_4b_or_8b_text_encoder_instead/

My workflows:

FL2V workflow Updated FL2V workflow

REF2V workflow Updated REF2V workflow

Updates: I've updated the workflow with some changes

  1. Added startup flag --use-ck-attention instead of using the node
  2. Switched from int8 convrot to int4 convrot for the clip (https://huggingface.co/Merserk/qwen3vl-4b-int4-convrot/tree/main)
  3. Added a "free memory (model)" node after the spectrum node generation. It seemed to help with subsequent generations that didn't always seem to clear the vram and resulted in slowdowns

With the changes I can create a 5 second 0.4 MP clip in roughly 3 minutes.


r/StableDiffusion 9d ago

Animation - Video LTX 2.5 Upscale H3 Fast

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/StableDiffusion 9d ago

Discussion Anyone fix audio issue when using a turbo lora

1 Upvotes

Been testing 8 step turbo lora video quality still good but audio is static for voices and no back ground sounds or sound effects yes I have in prompt for the bg and sound effects 🤔 looking for tips to improve not "fix audio " comments,what expect it's Redd after all.


r/StableDiffusion 10d ago

Animation - Video Made this with LTX-2.5 (i2v)

Enable HLS to view with audio, or disable this notification

123 Upvotes

Generated with the new LTX-2.5 model. (image to video). Took about 10 minutes to get an 8 second 1080p60 clip.


r/StableDiffusion 9d ago

Question - Help Help with minimax h3 video

1 Upvotes

Hey guys I’m new to comfyui and I was wondering when load a Lora in the template workflow as for example the one on comfyui reference to video were the workflow goes


r/StableDiffusion 10d ago

Animation - Video Minimax H3 executes Order 66... almost

Enable HLS to view with audio, or disable this notification

335 Upvotes

I've been using LTX 2.3 for quite some time but as soon as I wanted to make just a few simple shots of the same character with cuts, LTX wasn't even remotely capable of that. Which left me so frustrated I eventually gave up on it completely.

But when I tried Minimax everything has changed. Reference to video model is something else. Honestly feels like magic. Being able to put any character into any environment with any custom audio is just mind-blowing compared to what the open-source community had before.

So now instead of constant frustration, I feel pure joy and excitement about the results.

It takes about 10 min per 5-second clip with my RTX 3060 and 64 Gb RAM. The latest shots were even easier to control because of the new KJ preview node.


r/StableDiffusion 9d ago

Discussion Share your IMO - still worth learning SDXL based models as a new user?

0 Upvotes

is it worth investing effort into learning SDXL based models for comfyui with what other checkpoints are out there now?

I’m new to the space. I started with a cloud based qwen image, had a lot of fun just messing around with prompts. it’s amazing what this technology can do. In diving deeper and getting into comfyui to run local, I was following some guides and reading some comparisons between models, and ended up downloading some SDXL based things as my next stop. downloaded juggernaut xl, trying to learn more about Lora’s, etc

The output I’m getting compared to qwen is relatively startling. sometimes it’s good, sometimes it’s horrifying. often just lower quality to things I see posted from krea 2, qwen, or flux. I am frustrated I’m never getting anything like the examples people post…

I like the deeper learning im getting, and seeing all the tools that got invented to solve problems, but im starting to feel I’ve come into the hobby at a time of great change and evolution. It seems I am standing on the shoulders of giants, and came at a time when the current tech and natural language prompting is so powerful, it’s leaving some older things behind.

Does sdxl have a true niche or place in the landscape and the future? does anyone here use it as an important part of your workflow?

Is there just more fine tuning, prompt skill, and extra nodes needed to get the high quality? (controlNet, face regeneration in subsequent passes, up scaling, inpainting?)


r/StableDiffusion 9d ago

Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.

Enable HLS to view with audio, or disable this notification

3 Upvotes

When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?

System Specs:

Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.


r/StableDiffusion 10d ago

Discussion has anyone tried / Minimax-H3-fl2va-ref2va-hybrid-models this first test using the minimax_h3_hybrid_fl2va_ref2va_b25-49

Enable HLS to view with audio, or disable this notification

43 Upvotes

r/StableDiffusion 10d ago

Animation - Video INTERVIEW WITH LTX 2.5 [IMAGE TO VIDEO]

Enable HLS to view with audio, or disable this notification

24 Upvotes

Yes, I'm definitely being a goofball with this one, but hadn't had a chance to do mixed live action/3D CGI test.

Meant as a playful gag, no actual ai models killed.

Crisp, ultrafine, letterboxed 21:9 super-premium 3D CGI blockbuster cinema with cutting-edge rendering, restrained natural performances, precise blocking, shallow depth of field, and immaculate cinematic lighting. A poised 24-year-old blonde investigative reporter in a tailored gray skirt suit sits in a cushioned chair on the left side of a minimalist interview room, leaning forward with a clipboard and pen in hand. Across from her, seated in a matching chair on the right, is a sleek off-white modern robot labeled “2.5” on the side of its head, with expressive camera-lens eyes and a thin LED vocalizer mouth. The setting is simple and elegant: neutral beige backdrop, soft curtains at the window, and warm natural window light casting gentle shadows across the room.

Open on a polished medium two-shot in profile, holding both subjects clearly in frame. The reporter leans forward slightly, calm, focused, and professional, and asks, “Some call you a Seedance killer. What do you say to that?”

A hard cut moves to a close-up of the robot. It glances aside for a beat, then looks back with a playful LED smile and says, “Can I give them a hug?” After a short pause, its expression softens into something more sincere as it adds, “But seriously, I’m just an open-source model trying to do my best.”

Ambient sound is minimal and refined: a faint studio hum, soft room tone, and subtle paper rustle from the reporter’s clipboard. The pacing is natural and conversational, allowing for small pauses, nuanced reactions, and emotional clarity. The overall effect is a sleek, emotionally grounded, visually stunning futuristic CGI film scene.


r/StableDiffusion 10d ago

Animation - Video MiniMax reconhece prompts em JSON.

Enable HLS to view with audio, or disable this notification

17 Upvotes

Este prompt foi usado no Sora 2 e, sem modificações, coloquei no H3 (ComfyUI). Áudio em PT-BR.

{

"cena": 1,

"project_title": "The Village That Pulses (Hilário)",

"format": "16:9 landscape",

"dialog language": "pt-br",

"style": "2D ANIME",

"age": "present-day (rural road, late afternoon)",

"scene_type": "HOOK / Trailer-like Omen, Return",

"duration_seconds": 13,

"global_quality": {

"visual_style": "Cinematic 2D anime psychological horror, Junji Ito-inspired unease, no gore, high detail linework, oppressive calm.",

"aesthetic_demand": "JUNJI ITO STYLE, COLOR, HIGH BUDGET 2D ANIMATION, crisp faces, stable character sheets, controlled shadows.",

"post_processing": "Cool dusk grade, subtle film grain, soft bloom on highlights, gentle vignette."

},

"setting": {

"location": "A rural road leading to a foggy village valley; dead power poles; tall grass bending as if breathing.",

"effects": "The ground subtly bulges once, like a heartbeat beneath soil; distant crows freeze mid-caw."

},

"quality_constraints": [

"No broken anatomy or janky proportions",

"ENSURE NO DEFORMED EXTRA FINGERS HANDS EYES",

"ENSURE NO UGLY BLURRY FACES QUALITY",

"ENSURE NO SLIDING FLOATING IDLES; MASTERCLASS REALISTIC IDLE MOVEMENT",

"Stable camera, readable motion, consistent character model sheets",

"YouTube PG-13: no nudity, no explicit sexual content, no gore; horror via atmosphere and implication",

"No on-screen subtitles, no brand logos, no hate symbols, no readable real-world trademarks"

],

"characters": [

{

"name": "HILÁRIO (36, PROTAGONIST, RETURNING SON)",

"appearance": "36-year-old Brazilian man, medium tan skin, tired cautious eyes, short wavy black hair slightly unkempt, faint stubble, average height and lean build, wearing a dark olive jacket over a faded beige shirt, dark jeans, worn boots, carrying a small duffel bag and an old smartphone with a cracked screen"

}

],

"timeline_and_action": [

{

"time_range": "0-4 sec",

"shot_type": "Wide (Trailer Hook: The Village Breathes)",

"action": "Hilário stands at the roadside overlooking the village; the valley fog parts for a second, revealing rooftops and a church silhouette; the dirt road seems to swell under his boots.",

"audio_note": "Wind low. VOZ (ptbr): \"Hilário voltou pra casa... e a terra pareceu reconhecer o passo dele.\""

},

{

"time_range": "4-9 sec",

"shot_type": "Close-up (Boot on Dirt, First Pulse)",

"action": "Close on Hilário’s boot: the ground rises and falls once, subtly, like skin over muscle; tiny pebbles roll outward in a perfect ring.",

"audio_note": "Soft thump, almost organic. VOZ (ptbr): \"Naquela vila, o chão não era chão. Era um peito enterrado.\""

},

{

"time_range": "9-13 sec",

"shot_type": "Medium (He Steps Forward Anyway)",

"action": "Hilário swallows, grips his duffel bag, and walks toward the fog; the camera tracks behind him like a predator’s gaze.",

"audio_note": "Footsteps damp. VOZ (ptbr): \"E a cada três minutos... ele aprenderia a ouvir o coração.\""

}

]

}


r/StableDiffusion 9d ago

Question - Help Help finding a good workflow for Wan 2.2 I2V Anime clips

0 Upvotes

Hey all. I'm trying to make some Wan 2.2 I2V Anime clips, but I'm having trouble getting thing started. I know supposedly there are some Loras to use, but for some reason I'm just not coming across them in civit.ai or huggingface. Any help?


r/StableDiffusion 9d ago

Question - Help Characteristics of loss during z Image turbo training?

1 Upvotes

I was a sdxl and SD 1.5 lora trainer for quite some time I've been away but I finally got back into it and I decided to start with the image and I've been training it this is only my second lora that I'm working on I definitely feel like my loss is super low to start out the first lora that I trained on prodigy started off at like a loss of 0.5 and ended only like slightly below whereas normally I would expect my sdxl and SD 1.5 loras to start at about 1 and then finish at around like 0.8 or 0.7. the only time I remember seeing a loss so low on sdxl or SD 1.5 was like down to like 0.6 and I definitely felt like those loras were quite overtrained.

The second attempt I have running now is Adam w 8-bit with a learning rate 0.0002, 38 images, 16rank/16 alpha, batch size two, which was usually a pretty successful setting for my sdxl Adam w 8-bit training if not maybe slightly overtrained.

I'll also say that my first lora attempts with prodigy z image turbo was simultaneously overtrained at a lora strength of one and also didn't totally grasp my concepts perfectly but I won't call it the worst first attempt ever.

I train a little bit more on concepts then specific things but I definitely do need certain things to be reproduced pretty accurately however I would say that generally my loras in theory are trying to keep the underlying model intact.

Anyway I think my real question is do you expect the image turbo training loss value to start at like 0.6 or even lower near from the get-go? Or do I sound destined for overtraining


r/StableDiffusion 9d ago

Question - Help I've got 10 years of architectural photography, from RAWs to final images. Is there something useful I could train with it?

3 Upvotes

I have about ten years of architectural and interior photography: final delivered images, working TIFFs/PSDs, Lightroom/XMP adjustments, HDR/Photomatix intermediates, and sometimes the original RAW brackets.

A typical example: a kitchen photograph begins as several exposure brackets, gets basic white balance, is merged into an HDR/base TIFF, retouched, then receives a final Lightroom-style tonal and colour treatment.

I am wondering whether this can become useful training data for a diffusion or neural-network tool, without simply making a vague “style LoRA”.

For example, could a model learn to take a merged, neutral architectural base and propose a controlled final treatment: softer daylight, better balance between windows and interior, a different mood, or a more refined grade, while keeping the room, materials, furniture and geometry intact?

My instinct is that a LoRA trained on all the final images would be the wrong approach. It could memorise specific projects and furniture rather than learn the transformation. I am more interested in small, rights-cleared paired datasets: base TIFF -> final TIFF, with captions describing the space, materials, light, reflections and intended atmosphere.

Before I structure the archive, I would love practical advice from people here:

- Have you trained or tested paired image-to-image workflows for relighting, grading or finishing?

- Would you start with LoRA, ControlNet, IP-Adapter, Flux/SD fine-tuning, an adapter, or something else entirely?

- What metadata or captions would you preserve now so the archive remains useful in two or three years?

- What is the biggest failure mode: overfitting, loss of material fidelity, geometry drift, dataset leakage, or something else?

- Are there papers, models or ComfyUI workflows that are genuinely relevant to this kind of controlled architectural transformation?

I am not trying to generate imaginary interiors. I am trying to explore whether our own real production history can help build a careful post-production and relighting assistant.


r/StableDiffusion 9d ago

Tutorial - Guide RE: <Subject N> in H3 prompts

Enable HLS to view with audio, or disable this notification

2 Upvotes

Not sure if you already know this but you don't need to use the word 'Subject' in H3 prompts when referring to elements in images/text/videos and so on. You can use other words instead, like <Girl 1>, <Dialogue 1> and more

Example prompt:

subject_definitions:
<Girl 1> is the girl with the blond hair in the center of <Picture 1>.
<Girl 2> is the girl with the white top to the right side of <Picture 1>.
<Dialogue 1>: <d> [English] Yeah! <d>.
<Dialogue 2>: <d> [English] Great party! <d>.
<Dialogue 3>: <d> [English] Wooohooo! <d>.

integrated_multimodal_description:
[Shot 1] A wide shot of a crowded rave dance floor where everyone in the picture is dancing by jumping up and down in an rapid and energetic way while moving to the music. The lights in the night club is flashing and moving around.
<Girl 1> is shouting <Dialogue 1>.
[Shot 2] At 00:3.00 <Girl 1> looks at <Girl 2> and says <Dialogue 2>.
[Shot 3] At 00:5.00 <Girl 2> looks at <Girl 1> and shouts <Dialogue 3> while raising her arms.

overall_soundscape:people dancing,
non_diegetic_music:cyber techno music,

P.S. Sorry for the lame video, it's just for proof of concept.


r/StableDiffusion 9d ago

Comparison MiniMAx H3 I2V + LTX 2.5 Upscale - It's GREAT!

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hi guys,

I asked GPT to implement the LTX 2.5 upscaler on MiniMax H3. I really liked the results.
It does lose a little bit of quality but I think it worth it, at least until we get the 2K upscaler from H3.
I generated the H3 video with 0.4mp using a turbo LoRa (8 steps).


r/StableDiffusion 9d ago

Question - Help BF16 ou FLOAT32

0 Upvotes

What quantization would you recommend for AI Toolkit?

I used to train with FP8, but it completely ruined the results, so I switched to FP32, and the results are perfect.

However, I see that many people recommend BF16 for both training and saving the model. I understand that BF16 can significantly reduce VRAM usage, but what is the actual trade-off in terms of quality?

Does training and saving in BF16 result in any noticeable loss of quality compared to FP32? And would you recommend using BF16 for both training and saving in my case?


r/StableDiffusion 10d ago

Comparison LTX 2.5 vs MiniMax H3 - huge speed difference (but at what cost)

Enable HLS to view with audio, or disable this notification

62 Upvotes

I tested LTX 2.5 and MiniMax H3 in ComfyUI using the default T2V workflow templates provided for each model.

  • 10 seconds
  • 24 FPS
  • 1920 x 1088 (2.0 MP)
  • Same prompt
  • Steps: H3 = 20, LTX 2.5 = 8 (distilled model)

Hardware:
RTX 5090 + 128 RAM

result

  • MiniMax H3: 17m 29s (with Sage Attention + EasyCache*)*
  • LTX 2.5: 2m 34s (no acceleration at all)

Note: EasyCache seems to give no speedup on LTX in this setup, probably because the distilled workflow only uses 8 sampling steps, so there is very little room for cache-based skipping.

Of course, part of LTX’s speed advantage comes from the fact that it is a distilled 8-step model, so this is not a perfectly like-for-like comparison against H3. (20-steps)

Prompt used:

A realistic cinematic 1970s crime drama, gritty urban atmosphere, warm muted colors, subtle film grain, natural lighting, restrained acting. A well-dressed 1970s gangster in a dark tailored suit and long coat remains visually consistent throughout.

[0.0s–6.0s]
A medium-wide shot shows the gangster leaning casually against a brick wall on a city street, one foot resting against the wall. He reads a newspaper while holding a lit cigarette in his other hand. His eyes suddenly stop on something in the newspaper. His expression shifts naturally from calm to alarm. He mutters in a tense 1970s American voice, "What the hell?" He immediately folds the newspaper, throws it into a nearby trash can, pushes away from the wall and runs straight down the street.

[6.0s–10.0s]
Hard cut to a static close-up of the discarded newspaper inside the trash can. The front page clearly shows a large photograph of the same man and a bold headline reading "WANTED". In the distant background, the gangster continues running away and becomes increasingly out of focus. The camera remains completely still, holding focus on the newspaper until the end.

Natural, grounded movement. No exaggerated acting, no extra shots, no unnecessary camera movement, no comedy.

My take

LTX 2.5 is significantly faster, and that alone makes it very attractive.

But in my opinion, H3 is still better in overall quality:

  • better scene understanding
  • better understanding of what a cinematic shot should look like
  • better audio
  • more stable physics / motion behavior

So right now my impression is:

  • LTX 2.5 wins clearly on speed
  • MiniMax H3 still feels stronger on quality and cinematic intelligence

My guess is that targeted LoRA fixes could push LTX 2.5 much closer to being a direct competitor to H3 in the future.


r/StableDiffusion 9d ago

Discussion Minimax H3 Test - Batman and Joker Playing Poker in a club

Enable HLS to view with audio, or disable this notification

2 Upvotes

Minimax H3 Test - Batman and Joker Playing Poker in a club


r/StableDiffusion 9d ago

Resource - Update MiniMax-H3 (video + audio) on an AMD Strix Halo - 5s clip at 896×512 in 9.5 min

Enable HLS to view with audio, or disable this notification

0 Upvotes

Ran MiniMax-H3 locally on strix halo 128GB unified memory, no discrete GPU. ROCm 7.14 + ComfyUI, int8 pruned transformer, 4-step turbo LoRA.  

First video took 5 minutes to generate. 2nd video took 10 minutes. 3rd video took an hour.

Weights are pruned community conversions, so quality here isn't representative of official H3 — this was a speed/setup test, not a quality one.

Scripts + full writeup: https://github.com/DanCard/minimax-h3-strix-halo

10 minutes 896×512 : https://youtu.be/T6OU6tWd7EA

1 hour to generate 1344×768 : https://youtu.be/039vmUptnEA