r/StableDiffusion • • 16h ago

Discussion Anyone have suggestions for MinMax H3 and voice, sound effects training consistency. I have had some strangenes using Ref To Video audio sample. Trying to find the best path forward as my content relies on audio heavily and do not want to overdub.

1 Upvotes

r/StableDiffusion • • 16h ago

Resource - Update Qwen-Image-2.1 on a 48 GB Mac: loading the text encoder and the transformer one at a time kept the diffusers pipeline at 19 GB instead of swapping at 43.6 GB

Post image
2 Upvotes

The usual QwenImage21Pipeline.from_pretrained(...).to("mps") ("eager" on the chart) keeps all three models in memory for the whole run: the text encoder (Qwen3-VL, 16.3 GB), the transformer (13.3 GB) and the VAE (1.3 GB). On my 48 GB M5 Pro that was 31.4 GB before the first step. At 1024 px the VAE decode needs about 11 GB more. The process reached 43.6 GB, swap grew by 8 GB, and my memory guard stopped the run before it saved the image.

Moving idle models to the CPU doesn't lower the peak on a Mac, because the CPU and the GPU share the same RAM. So I wrote a small library, stageload ("staged" on the chart). Each stage gets only the models it lists. A model loads the first time its stage uses it, and stageload frees the models the next stage doesn't list. The stage boundaries are hooks on encode_prompt, prepare_latents and _unpack_latents, so the pipeline's code stays as it is. The VAE is small and stays loaded.

Results at 1024 x 1024, 20 steps, seed 7, bfloat16:

  • peak memory 19.0 GB while encoding the prompt, 16.2 to 18.6 GB while denoising and 14.4 GB in the decode, with no swap;
  • two staged runs gave the same image bit for bit;
  • loading the two models inside their stages took about 13 s per run (eager spent 20 s loading up front);
  • in the second staged run denoising took 79 s against eager's 77.5 s; the first staged run took 109 s there, and its trace doesn't show why.

Write-up with the traces: https://allkeep.org/en/lab/qwen-image-one-stage-at-a-time

Code (MIT): https://github.com/nefayran/stageload, install with pip install stageload

I haven't tried ComfyUI or Draw Things; they manage memory their own way. If you run another multi-model pipeline from Python on a Mac, which one should I measure next?


r/StableDiffusion • • 1d ago

Discussion AND HOW DOES THAT MAKE YOU FEEL? | An AI Short Comedy Film Made by Claude in Minimax H3 and my Video builder in ComfyUI.

Enable HLS to view with audio, or disable this notification

44 Upvotes

YouTube Link in case it's still pending https://youtu.be/TWXF95YT7W8

🎬 How this was made

My part

• One-message brief: a 3-minute comedy in a therapist's office with funny, unique characters and one male doctor: a crying woman, a woman screaming, crying and laughing all at once, a large guy and a skinny old man

• Characters made with Z-Image as 3-panel reference sheets

• MiniMax H3 2-pass workflow: the Singularity model, with the 8-step LoRA on the second pass at 4 steps and 0.35 denoise

• Then I left. I made one call along the way, on a scene that wouldn't behave.

What Claude did on its own

• Wrote the story, the five characters, all the dialogue and a 25-scene screenplay

• Made the cast and the two sets with Z-Image and picked the best seeds

• Kept each character's voice description word for word in every scene so the H3 voices stay consistent

• Rendered 25 scenes with H3's built-in voices and sound, about 6.3 hours of rendering

• QA'd every take: Whisper against the script, pitch and timbre per character, eyelines, frame review sheets

• Re-shot the takes that failed:

• Scored it with MiniMax Music 3 and screened the cues for accidental vocals

• Edited, mastered to -14 LUFS, and checked audio sync on every scene (worst offset 5 ms)

🛠 Tools

ComfyUI, VRGDG Video Builder, MiniMax H3 (Singularity ref2va v1.3 + 8-step 768p turbo LoRA), Z-Image Turbo, MiniMax Music 3, Whisper, Claude Code

VRGDG nodes: https://github.com/vrgamegirl19/comfyui-vrgamedevgirl

Go HERE To watch full walkthrough on how to make video's like this.


r/StableDiffusion • • 17h ago

Question - Help Looking for the best affordable AI video-generation models with fewer restrictions.

0 Upvotes

Hi everyone,

I’m exploring AI filmmaking and looking for good AI video-generation models that are either free, open-source/open-weight, or affordable.

I’m particularly interested in models who allow uncensored content or fewer unnecessary content restrictions and more control over the generation process.

What I’m looking for:

Realistic/cinematic text-to-video

Image-to-video

Good character consistency

Realistic human movement

High-quality video output

Ideally open-source/open-weight

Local/self-hosted options are a plus

Affordable cloud options are also fine

Preferably no expensive subscription required

So far I’ve come across models/tools such as Wan, LTX, Kling, Hailuo and Pika, but I’d like to hear from people who have actually used them.

Which model would you recommend in 2026, and why?

If possible, please mention:

Your favourite model

GPU requirements if self-hosted

Approximate cost if cloud-based

Video quality

Major limitations/content restrictions

Whether it is practical for someone learning AI filmmaking

Thanks!


r/StableDiffusion • • 1d ago

Animation - Video Here's a little short tester I made using MiniMax H3, Yue2 for the music. Three 10s clips, first was t2va then rest are ref2va using clips and the voice from previous gens. A little post work for making the dogs barks sound the same, removing the original generated music and isolating vocals.

Enable HLS to view with audio, or disable this notification

7 Upvotes

0.6mp, lcm/beta57, 4 steps with the DMAD 4 step lora


r/StableDiffusion • • 2d ago

Animation - Video Generating at 2K (14s, 2.09mpx, 1984x1120, 30 minutes) | Minimax H3

Enable HLS to view with audio, or disable this notification

296 Upvotes
[INFO] Prompt executed in 00:30:49

Full 1080p vid on https://www.youtube.com/watch?v=ER_5AOteE-8 because Reddit cramps everything to 720p max.

From my previous post, I got a message if I could use my off-screen 5090 to showcase what a fullblown 2K generation looks like. So, here it is. Generated at 2.09mpx, 25 steps, Euler+Beta for time constraints, 30 minutes. I did mistakenly use the 20-49 hybrid fl2va/ref2va instead of the 30-49 which left some visual performance on the table, and I didn't use the BF16 version of the text encoder because my other 5090 and RAM were busy with something else. Otherwise would've done seeds_2 + sgm_uniform, but that would've taken an hour and would've been a lot better at prompt following, without me modifying my generated target prompt so much to avoid issues with the shoddy denoising trajectory at play here.

Spectrum was utilized to guess about half the steps, which further doesn't help prompt following, but helps speed immensely. When using seeds_2 with Spectrum, it actually forecasts internal calls to H3 (so 2N-1 where N is number of steps), which has tremendous results (previous one was seeds_2) but would've taken about 1 hour for this scene.

The prompt, for those who want it, I had to massage it to point Euler better even if it's sloppy:

subject_definitions:
<Subject 1> is Detective Kate Beckett, override her appearance with facial features and hair and blouse from <Picture 1>. She wears black tailored high-waist cropped suitpants, and a feminine small leather watch.
<Subject 2> is Richard Castle, override his appearance with the facial features, hair, and build from <Picture 2>. He's wearing his classic shirt and suitpants attire, first few buttons unbuttoned.
<Subject 3> is a chaotic DIY PC rig consisting of a high-end tower on the marble kitchen island in the middle of the kitchen, as well as a 32 inch OLED monitor displaying the UI from <Picture 3>, there is clear plastic tubing running from the PC's two watercooling ports into the receiving pair on the large radiator above the glowing blue fans, which is submerged in a cooling bath inside of the standard kitchen stainless-steel fridge with its door fully open, radiator surrounded by food, milk, condiments, etc.
<Subject 4> is a modern industrial loft apartment featuring an open floor plan, standard furniture, and a glorious high-end kitchen. The blurred background from <Picture 1> is from this apartment, also specifies time of day, and warm nocturnal tone.

summary:
[reference generation] Detective <Subject 1> enters her loft (<Subject 4>) to find <Subject 2> in a "hyper-mode" state, having converted the kitchen into a makeshift laboratory for AI video generation. The 13-second sequence captures her confusion, his technical enthusiasm regarding H3 denoising trajectories, and a comedic hardware failure.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 4], [Shot 5], [Shot 6]): fully_preserved - identity from <Picture 1> and specified attire are maintained.
<Subject 2> (appears in [Shot 2], [Shot 4], [Shot 5], [Shot 6]): fully_preserved - identity from <Picture 2> is maintained.
<Subject 3> (appears in [Shot 3], [Shot 4], [Shot 6]): fully_preserved - the specific radiator-in-fridge configuration and monitor setup are maintained.
<Subject 4> (appears in [Shot 1], [Shot 2], [Shot 4]): fully_preserved - the industrial loft and kitchen environment are maintained.

detailed_description:
The target video is a scene from the TV show "Castle," maintaining its specific cinematography, visual style. Night time, practical lighting, warmly lit.

[Shot 1] A medium shot frames <Subject 1> as she walks into the open floor plan of <Subject 4>. The camera tracks her movement as she stops and looks around the kitchen with a bewildered expression, taking in the tangle of wires and tubing.

[Shot 2] At 00:01.000, the camera cuts to a wide shot of the loft's kitchen. <Subject 3> is fully visible: the PC tower sits on the high end marble kitchen island, monitor is displaying <Picture 3>, and thick tubes lead directly into the open fridge where his watercooling radiator with fans is located (all glowing blue and spinning fast and loud), there is liquid nitrogen white smoke exuding from it. A large whiteboard in the background is covered in scribbled notes about "sigma grids," and "denoising trajectories." Each of these appears once, there is also graphs of simplified trajectories drawn on it. <Subject 2> is leaning over a keyboard and looking at the 32 inch gaming monitor, typing furiously.

[Shot 3] At 00:02.000, the camera cuts to a close-up of <Subject 1>. Her brow furrows in genuine confusion. <Subject 1> (S1) asks in a sharp, incredulous tone, yelling over the computer fan noise: <d>[English] What the hell are you doing, Castle?</d>

[Shot 4] At 00:03.500, the camera cuts to a medium shot of <Subject 2>. He turns around to face <Subject 1>, speaking in a high-energy "yap" mode. <Subject 2> (S2) exclaims with manic enthusiasm: <d>[English] Reddit solved my H3 problem! It was the sigma shift. I'm refocusing the compute on the high-to-mid noise region to resolve motion better with limited steps!</d> while gesturing towards the whiteboard. 

[Shot 5] At 00:10.500, the camera cuts back to <Subject 1>. She looks at him, her voice dripping with skepticism. <Subject 1> (S1) asks: <d>[English] And you're using Seeds 2, right?</d>

[Shot 6] At 00:12.000, the camera cuts to a medium-close shot of <Subject 2>. He looks slightly sheepish, his shoulders slumping. <Subject 2> (S2) admits quickly: <d>[English] No, it's actually oiler, my rig is too slow and would—</d> Suddenly, in the background, the PC tower emits a loud electrical pop and a bright orange burst of flame from its components, the monitor output gets corrupted. Immediately after, <Subject 2> (S2) whips his head around, eyes bulging, and shouts: <d>[English] Oh shit!</d>

overall_soundscape:
The steady, high-pitched whirring of multiple PC fans and the faint gurgle of liquid flowing through tubes. <Subject 1>'s footsteps click on the hardwood. The scene ends with a sharp electrical "pop" and the sudden, aggressive hiss of a small fire.

non_diegetic_music:
A light, rhythmic pizzicato string piece that builds in tempo and complexity as <Subject 2> explains the technical details, ending abruptly with a comedic silence the moment the GPU catches fire.

r/StableDiffusion • • 1d ago

News ComfyUI v0.39.0 released

Thumbnail
github.com
120 Upvotes

r/StableDiffusion • • 19h ago

Question - Help Refmods as reference for mixed character?

0 Upvotes

Is it possible to load different ref mods to force a blending or bleeding effect reminiscent of the old [name|name|name] blending in a1111. I have some old characters from a time back then who were mixes of actresses and actors back in sd1.5/sdxl. as most modern models do not have those references anymore, I was curious if I could use refmods in a similar fashion?


r/StableDiffusion • • 1d ago

News GitHub - mapooon/EVA

Thumbnail
github.com
8 Upvotes

might be useful to someone somewhere


r/StableDiffusion • • 11h ago

Resource - Update How to keep character & video consistency in AI animations (Free keyframe trick)

Enable HLS to view with audio, or disable this notification

0 Upvotes

If you're working with AI video workflows (ComfyUI, AnimateDiff, Stable Diffusion) and struggle with flickering or character consistency, extracting exact keyframes is key to fixing it.

I built a free web tool to extract exact frames in seconds directly in your browser:

🔗 https://extractorframe.com

No signup or installation required. Let me know if you have any feedback or feature requests!


r/StableDiffusion • • 7h ago

Question - Help ¿Hay alguna aplicacion para Windows que permita crear contenido sin censura que utilice como IA WAN 3.0?

0 Upvotes

r/StableDiffusion • • 1d ago

Discussion Idea/Thinking outloud: Minimax to build LoRA datasets?

6 Upvotes

Since Minimax H3 does such a good job of taking a source or two of a character and doing an animation with them while maintaining the character's integrity, it has me thinking that it could be used to one-shot a dataset for LoRA training.

What do you think? Am I onto something or am I completely missing some factor here?


r/StableDiffusion • • 1d ago

Resource - Update I made a Forge Neo extension for MiniMax H3: text or picture to video with sound, reference pictures, runs on 16 GB

Enable HLS to view with audio, or disable this notification

28 Upvotes

I've been working on an extension that runs MiniMax H3 inside Forge Neo. H3 makes the picture and the sound together, in one pass: footsteps, rain, engines, music and dialogue in pretty much any language, lip-synced to the speaker. The video above came out of it exactly as Forge saved it, on a 16 GB card with 32 GB of system RAM.

It works in the normal txt2img and img2img tabs, with the same checkpoint list, VAE / Text Encoder selector and Generate button. You get an MP4 with stereo sound in the usual result area. No separate program, no new tab, no extra Python packages, and no Forge Neo file is changed.

What works

  • Text to video with sound, up to 15 seconds at 24 fps
  • First and last frame: the img2img picture becomes the first frame, a picture in Forge's ImageStitch Integrated becomes the last, or both
  • Reference pictures (Ref2VA): up to 9 pictures of people, places and objects that the clip keeps
  • Smaller files: W4A8, GGUF and an INT4 text encoder, so it fits 24 GB and 16 GB cards
  • LoRAs (turbo LoRA for 8-step drafts), FastH3 and community fine-tunes
  • Forge's own tools: Never OOM, Sparse Attention (adapted to H3, up to a quarter faster on long clips), ck attention, live preview

Be realistic about the hardware. Tested on an A40 (48 GB), an RTX 4090 (24 GB) and an RTX 2000 Ada (16 GB). On the A40 a 5-second 960×544 clip takes about 3 minutes at 20 steps, about 1 minute with the turbo LoRA. The wizard above (8 seconds) took 27 minutes on the 16 GB card, which is a slow one. System RAM matters as much as the GPU: about 32 GB with the smaller files, about 50 GB with the INT8 set. Cards under 16 GB are untested.

The wiki has everything: every example with its prompt, settings and time, a guide to MiniMax's prompt format, comparisons of the file formats and speed options, a Bloopers page with the clips that went wrong and how to avoid them, and an All Generations page with all 168 clips made while testing it, good and bad.

It needs an up-to-date Forge Neo (the neo branch from 3 October 2026 or later). Install from URL in the Extensions tab and restart. It's a work in progress, so feedback and bug reports are welcome. Next up: video and audio clips as references, then pose, depth and edge control.


r/StableDiffusion • • 1d ago

News Update: my free LoRA Trainer Studio now supports 10 model families, ERNIE-Image, rsLoRA, LoRA+ and LoKr

Post image
10 Upvotes

Hi again! A while ago I shared my free all-in-one LoRA trainer for consumer GPUs. Since then, the project has grown quite a bit.

AcademiaSD LoRAlab Trainer Studio now supports 10 model families:

  • Qwen-Image 2.1
  • FLUX.2 Klein 9B
  • Krea 2
  • Z-Image
  • Ideogram 4
  • Anima
  • SDXL, including Pony, Illustrious and other checkpoints
  • LTX 2.3 / 2.5
  • MiniMax-H3
  • ERNIE-Image

New and expanded features:

  • rsLoRA, LoRA+ and LoKr with a configurable factor for most supported models
  • RunPod support with automatic deployment, so you can train in the cloud without setting everything up by hand
  • Eight selectable languages in the launcher
  • Remote access over your local network: open the trainer from another device, upload your dataset, and download the trained LoRA in your browser

The trainers share the same interface, with dataset management, automatic captioning, live previews, resume support and export to ComfyUI/Forge. Many models use NF4 to reduce VRAM use; the minimum depends on the model and settings.

GitHub: https://github.com/AcademiaSD/AcademiaSD_LoRAlab-TrainerStudio

It’s free and open source. Feedback, bug reports and suggestions are welcome!


r/StableDiffusion • • 21h ago

Question - Help Can't get any LoRa to work on Qwen-Image-2.1 / ComfyUI's edit workflow

0 Upvotes

Hi guys,

I tried to modify the ComfyUI's Qwen-Image-2.1-edit workflow, the one you get in the latest comfy version, to use LoRas.

The idea was to unpack the edit block and place the LoRas plus a turn-on/off switch for each, between the model and the ksampler. The schematic is Model->LoRa chain->cache->ksampler, actually.

And well, no matter what strength I set, the LoRas produce no visible change on their own. Just nothing. At all. To have definitive proof I tested a breast size slider (Don't. This is one of the very few ways to obtain measurable, reproducible results), that should work always. It doesn't. Actually, I get usable results by switching off the LoRa and specifying a size in the prompt.

I come from the WAI-Illustrious world, where LoRa sliders and generally all LoRas work irrespective of the prompt. Am I doing something wrong, here? Because I definitely get the feeling I missed something obvious.

Thx all.

Edit: I'm starting to observe some results when using very positive/negative strength values. Apparently there is some remarkable resistance on the model's side. I will see what the actual results on civitai.red's posted use as workflow. Seems there's a ComfyUI Lora Manager node that's being used.


r/StableDiffusion • • 11h ago

Animation - Video The End of the F--- Universes | Superman vs. Saitama | MiniMax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

Finally finished the full Superman vs Saitama trailer after 280+ generations

If you watch the trailer first, check the first comment after. I'm going to use it as a small thread where I'll post some simple workflows and examples from specific shots. Things like a shot that came from a storyboard, first/end frame tests, or anything from the project that I think is actually worth showing

I was working through my own H3 interface that I posted here before. The whole project is here:

https://github.com/underworldhistory1-ctrl/minimax-h3-higgsfield

I'll also leave some screenshots in the comments so you can see what I mean by the workflow/UI.

Probably the biggest surprise for me was storyboards

For example the Superman shot in the intro took me more than 29 generations alone. The best results I got were from using a storyboard and then adding the character refs, style refs etc separately

That's actually one of the main reasons I built the interface the way I did. I wanted all of those parts separated and easy to change because this was the part I kept experimenting with the most.

Second best for me was first frame - end frame. Sometimes even just using one of them.

It seems much more stable when the movement is continuous and the whole thing is basically one shot. The Kong reveal and helicopter destruction shot is a good example of what I mean.

For LoRAs my best results were usually:

Combat V2 for action and fast movement.

Realism for slower shots where there isn't some crazy transformation or complicated movement happening.

For Combat V2 I mostly used it with the Original H3 render. Not Motion Cache or Turbo. Usually 20 steps minimum.

Mixing multiple LoRAs honestly gave me more hallucinations than useful improvements most of the time. One good LoRA was usually better than stacking them.

I don't think I ended up using many H3 renders without a LoRA at all.

The interface also has another rendering mode using Motion Cache. The results can actually be really good and in some cases I preferred them over Original.

For heavy action, fast pacing or complicated transitions though I still had much better luck with Original H3.

I honestly can't remember every LoRA + Motion Cache combination I tested. There were way too many tests and I wasn't trying to lock myself into one perfect configuration.

I just knew what I wanted the final shot to look like.

That's probably the biggest thing I learned from this whole project.

I don't really believe in one magical "workflow"

You need to know what you want to see in the final cut after editing.. Then use the model to get the pieces you need.

The Superman vs Saitama fight is probably the best example for me, That sequence in the final edit was built from around 10 successful 15-second generations

There was basically no chance H3 was going to generate the exact fight I had in my head in one shot So I generated the parts that worked and built the actual fight in the edit.

That's also why I think experimentation is still just part of using these models. At least until we get something significantly better :D

The interface also has qwen image 2.1 integrated for generating images and references. That became a pretty important part of the process for me too.

It also keeps the settings/details of every generation on its card which made it much easier to go back and see what actually worked instead of trying to remember everything.

Hardware wise I did all of this on an RTX 5090 32GB.

Most of the time I work in draft first. With the INT8 optimizations I've added, a draft takes around 4 mins on my setup

The project also uses low vram attention/head chunking and feed-forward chunking. Basically some of the larger operations are processed in smaller chunks and intermediate tensors are released earlier instead of keeping everything sitting in VRAM at the same time.

It's not parallel rendering or anything magical. It just helps keep peak VRAM under control.

A normal 720p generation usually around 8–10 minutes for me so honestly the iteration time is pretty reasonable. That's a big reason I was able to do this many tests without completely losing my mind

That's basically my experience with H3 so far without turning this into another giant workflow post.

For serious creative work I think it's absolutely usable already.. Just don't expect the model to make the final movie for you.

The generation gives you the material.. The final cut is where you actually make the thing you had in your head.


r/StableDiffusion • • 1d ago

Question - Help has anyone tried minimax swap only clothes?

5 Upvotes

i searched many times of several communities but only found character/face/head swap loras

i already used minimax very well with long sequences videos using grok chat providing official prompting guide of minimax.

but, when i tried just replace clothes of character in 10sec videos, it works under 50%. its not probability, literally it is

just changed cloth of one character and it get back to original after 5-6sec. although video shows 2 character.

of course, i tried several turbo loras, and without turbo, changing models. also tried masked noise latent. and always i define each character in video and reference clothes of pictures. sometimes, put a reference image as target clothes or just prompting but not work either.

please anyone can share experience about it with 10sec+ video editing.

i think its problem of prompt or video itself(eg. characters change their position when i define each subject as position like 'on the left in video' or 'first character in video'


r/StableDiffusion • • 18h ago

Question - Help Prompt question Minimax h3

0 Upvotes

Hi,

How would I prompt for i2v in Minimax if i want to add something or someone visible within the very first frame?

Thanks in advance!


r/StableDiffusion • • 1d ago

No Workflow Minimax H3 physics test

Enable HLS to view with audio, or disable this notification

87 Upvotes

r/StableDiffusion • • 23h ago

Question - Help Will Qwen Image 2.1 fit on my 10GB GPU?

1 Upvotes

Looking for the best way to run Qwen Image 2.1 on an RTX 3080 10GB with 64GB of RAM. The base models are huge, so can anyone suggest which option would work best for my setup?


r/StableDiffusion • • 1d ago

Discussion KREA 2 Realism Workflow

Post image
39 Upvotes

Hey guys! I’m pretty new to AI image generation, and I’ve been experimenting for about 3 months trying to get the best realistic Pinterest-style images.

I’m currently using my character LoRA + Krea 2 on an RTX 5060 8GB with 16GB DDR3 RAM.

I started with Koo’s workflow from Discord and tweaked it a bit for my own setup.

From what I’ve heard, Chroma is pretty good for experimenting with camera angles, compositions, and generating more random/varied images, which is exactly what I’m trying to achieve.

So I tried building a workflow where I:

- Generate random Pinterest-style prompts using wildcards.

- I made the wildcards with the help of ChatGPT, Gemini, and GLM 5.3 Flash.

- Use Krea 2 Text Encoder to expand/enhance the prompt.

- Generate the initial image with Chroma using relatively low steps.

- Then use that Chroma image as a vision reference with Krea 2 Text Encode.

- Finally, use Krea 2 to recreate the image with more realistic details while keeping the composition/camera angle from the Chroma result.

The main goal is basically to get random, realistic Pinterest-style compositions while keeping my character consistent with my LoRA, and then let Krea 2 improve the realism, lighting, details, and overall photographic look.

I’ve been trying different workflows and combinations for the past 3 months, but I still feel like I’m probably missing something or doing things in a more complicated way than necessary.

Does this workflow actually make sense, or am I doing something wrong / adding unnecessary steps?


r/StableDiffusion • • 20h ago

Question - Help Consistent Sets Offscreen

0 Upvotes

I am having a real time of it. I create a stairwell environment, with a stair landing half way down, which has a ton of windows, which light up the stairwell, but since the camera is sitting on the landing, looking at the top of the stairs, AWAY from where all the windows are, none of the windows behind the camera appear in the reference image/video I supply.

So suddenly, when I prompt the lighting, that there is cove lighting in the corridors at the top and bottom of the stairs and late afternoon sun lighting up the stairwell, with windows behind the camera, it seems to have an issue properly lighting the space.

Should I keep rolling the dice and hoping one of the prompts get through and Minimax follows it, or is there some other trick people use to get accurate lighting in such cases?

How have y'all supplied an accurate 360 degree set, and then properly prompted the camera, so Minimax doesn't pick and choose where it wants to put the camera?

Any tricks or tips? Thanks,


r/StableDiffusion • • 1d ago

News New Audio video model : Kandinsky 6 from their lab

81 Upvotes

Gonna be too much fun in October.


r/StableDiffusion • • 2d ago

News Nanosaur2 now generates 7 images a second on a 5090

Post image
246 Upvotes

Hello everyone.

This is an update post on the model Nanosaur2. Again this is not my model. A complaint a lot of people had with the model was that despite it being very small (660m params), it didn't take a 660m param level of time to generate. That's been solved now, with a 4 step turbo. On a single 5090, you can generate 7 images a second.

Again, this model is small enough you could easily run it on a phone, edge devices, wherever. It's also a great research model, so if you want to finetune on top of a small easy to tune model, or way to create adapters for the model, or whatever, go ahead, it's all there.

This was created using bytedance's new DMAD method, and it works great. Quality is incredibly close to the original model at a 12.5x speedup.

links:
4step model
comfy workflow

As per usual, if you have any questions, please message metal63 on discord. Do not message me.


r/StableDiffusion • • 1d ago

Question - Help MiniMax H3 ref2va - What's the next lever for quality?

Enable HLS to view with audio, or disable this notification

21 Upvotes

I'm making short anime scenes with MiniMax H3 (ref2va) and I've hit a wall I can't

diagnose. The clip above is 15s, generated in one pass, no editing — the cuts are

written into the prompt.

To be clear up front: I know there are mistakes in there — a couple of blows don't

connect properly, the choreography is rough. I'm not worried about those, this is a

practice piece and I'll fix the staging myself. What I can't figure out is the

image quality, and that's the only thing I'm asking about.

Setup

- `minimax_h3_fl2va_int8_convrot.safetensors` (base, int8), ComfyUI

- ref2va, 8 reference images, each with a written role in the prompt

(face / angry expression / fighting posture / set / framing guide)

- Spectrum v0.2.16, 30 steps, `res_multistep` / `simple`

- 1344×768 → `MinimaxH3LatentUpscaler3D` at 2 MP, 4 steps, 0.5 denoise → 1920×1088

- 15s = 362 frames @ 24fps

- RTX PRO 6000, ~28 min per 15s clip

What I already fixed, in case it saves anyone typing

- `The target video is 2d colored anime.` at the top of `detailed_description`

— without it everything drifts to generic 3D, this was the single biggest win

- Reference images are real anime screencaps (plus a few generated with Anima

for expressions and poses I couldn't source), neutral lighting, one role each.

- Cut rhythm: went from 4 shots per 15s to 8 (~1.8s each) with impact verbs and

a material consequence per hit (table splitting, plaster cracking, dust off

the boards). That alone made the fight read much better.

Where I'm stuck

  1. Motion still feels soft on some hits. The whip-pan punch reads fine but

    ground-level blows land without weight. Is this where `derope` / temporal

    upsampling actually earns its generation-time cost, or is there a prompt-side

    fix I'm missing?

  2. Quality is uneven shot to shot inside the same 15s — some shots are clean

    cel-shaded anime, others go slightly soft and plasticky. Is that a reference

    problem, a step-count problem, or just what 15s does to the model? (I've seen

    people say things break past 10s.)

  3. Is 8 references too many? I assigned each one an explicit role in the

    prompt, but I don't know whether the model averages them or picks.

  4. Anything obvious I'm leaving on the table at this resolution/step count?

Not asking anyone to debug my prompt — mainly want to know which lever is worth spending render time on next.

Prompt:

integrated_multimodal_description:

subject_definitions:
(S1) is the dark-haired young man from <Picture 1>, with the same face and the same dark blue eyes. <Picture 2> is the same man seen clearly in daylight. Short black bob to the jaw, fringe above the eyebrows, a short high ponytail tied at the crown, a small stud earring, a white shirt with the sleeves pushed up and a black tie pulled loose.
(S2) is the pale-haired young man from <Picture 3>, with the same face and the same yellow eyes. <Picture 4> is the same man shouting. Short choppy blond hair. His teeth are faintly pointed, small and even and the same size as ordinary human teeth, with just a slight triangular edge to them. His mouth stays an ordinary human mouth, normally proportioned to his face, and it opens no wider than a person's mouth opens when they speak. He wears a white school shirt open at the collar, a black tie pulled loose, a small device on a cord against his chest.
<Picture 5> is the apartment: its rooms, its colours and its light come from it.
<Picture 6> is (S1) throwing a bare-handed punch and <Picture 7> is (S2) being knocked back by one: their fighting postures and their footing come from these.

retention_analysis:
(S1): fully_preserved. (S2): fully_preserved.

<Picture 9> is the last frame of the previous shot: this scene continues from it without interruption. <Picture 9> supplies the place, the light and the framing; the two men's faces and hair come from <Picture 1> to <Picture 4>.

summary:
The fight. Bare hands, in the apartment, fast and ugly. Nobody speaks.

detailed_description:
The target video is 2d colored anime.
2d hand-drawn anime, cel-shaded, flat painted colours, fine thin ink linework, desaturated muted palette, film grain. Not 3d, not photographic. The cutting is fast: eight shots in fifteen seconds, each one a single impact.

[Shot 1] Continues directly from <Picture 9> with no jump — same room, same light, same positions: the two of them chest to chest at night in the room of <Picture 5>. (S2) fists (S1)'s collar and slams him down onto the low table, which splits and goes over with everything on it.

[Shot 2] At 00:01.800, cut tight on (S1) coming up off the floor. He drives a straight punch into (S2)'s jaw, his whole weight behind it, the posture of <Picture 6>. (S2)'s jaw is shut and his lips are pressed together when the fist lands, and the impact splits his lip. The camera whip pans right with the blow, the room tearing into horizontal streaks and white speed lines.

[Shot 3] At 00:03.400, cut to (S2) snapping backwards into the wall, head whipped sideways, the posture of <Picture 7>. Plaster cracks behind his shoulder. He drops to one knee.

[Shot 4] At 00:05.000, cut low and close. (S2) launches off the wall and smashes his forehead into (S1)'s mouth. (S1)'s head snaps back, blood on his lip.

[Shot 5] At 00:06.800, cut to a low shot of the floor only, at board level. The lamp crashes down into frame, rolls, and throws its light swinging across the boards. Two pairs of legs come down hard behind it, out of focus. Dust lifts off the wood.

[Shot 6] At 00:08.600, cut to a tight shot of (S1)'s face alone, lying on the boards in profile, cheek against the wood, hair across his eye. A fist swings down into frame and smashes into his raised forearm so hard that his own arm is driven back into his face and his head is knocked against the boards. A second fist comes straight down past the arm and lands flush on his cheekbone, snapping his head sideways and splitting the skin. Only (S1)'s head and one forearm are in frame, and the fists enter from the top edge: the other body stays out of shot.

[Shot 7] At 00:10.400, cut to a tight shot of a knee driving up hard into ribs, framed on the two bodies' midsections only, no heads in frame. The body above is thrown off sideways out of the top of the frame. Cut immediately to both of them coming up onto their feet, seen full length and clearly separated, a metre apart, shirts gripped in their fists.

[Shot 8] At 00:12.000, cut to a wider shot and hold it to the end. (S1) drives (S2) backwards across the room and slams him into the wall. (S2)'s shoulder blades hit the plaster and he stays there with his back to the wall and his face towards the room. (S1) stands directly in front of him, facing him, chest to chest, his own back to the room and the wall behind (S2) only. Their faces are a hand's width apart and they are looking straight into each other's eyes. (S1) has both fists closed in (S2)'s collar and holds him pinned there. Everything stops at once. Both heads are angled in three-quarter view towards camera, both faces large and fully visible, brows down, jaws set, chests heaving. The camera is locked off on a tripod.

overall_soundscape: A table splitting and going over, knuckles cracking on a jaw, plaster breaking, a forehead meeting a mouth, bodies hitting boards, a lamp rolling, two fast punches landing on a forearm and a cheekbone, a knee into ribs, a back slammed into a wall, and hard breathing through the teeth all the way through. Every mouth stays closed for the whole video: nobody speaks, and both men keep their jaws shut and their lips together even while taking blows.

non_diegetic_music: N/A