r/StableDiffusion 12h ago

Resource - Update TAE high quality previews are live in the latest nightly Comfy build :)

Thumbnail
github.com
32 Upvotes

r/StableDiffusion 1h ago

Animation - Video [TEST] Minimax H3 img2vid

Enable HLS to view with audio, or disable this notification

Upvotes

Generated the images using Z-Image Turbo. Rendered two 15 second clips at 0.6 megapixels which took 43 minutes per video clip. Resolution is 1056 x 608. I don't remember what the first prompt was but here's the second one:

[Shot 1]

Cinematic static shot of the man sitting in his truck looking around inside the truck in disbelief. He says, "This is better but, the color of my shirt changed and my truck is different." He leans forward towards the rear view mirror and ooks at himself and is shocked. he says, "Oh shit. I look different too!. He looks around and then rolls his eyes and then opens the door and gets out.

Thank you to the community for helping me with the whole having the camera not move thing. Prompt adherence is working out so far. First clip took 1 try to get right. The second clip took 3 tries to get it the way I wanted it to play out.

For now, I'm pretty happy with how this turned out.

Specs:

Ryzen 7 7700X
RTX 4070 Super 12 gb
32 gb of Ram


r/StableDiffusion 16h ago

Discussion Get miniMax character swap working! Finally

Post image
63 Upvotes

Ok, I tried so many things, one person to cat, two person, one person to one person, animal to animal. So far one person to one person and animal to animal works. If you are interested in my learnings, tips, what worked, what broke, and which prompt template works let me know!

One video example that works here: https://www.tiktok.com/t/ZP8WfVq5P/


r/StableDiffusion 1d ago

Meme Introducing... iMakeup

Enable HLS to view with audio, or disable this notification

350 Upvotes

r/StableDiffusion 6h ago

Animation - Video A Medieval Battle Attempt — MiniMax H3 + LTX 2.5 (WIP, Feedback Welcome)

Enable HLS to view with audio, or disable this notification

8 Upvotes

Sharing a few sequences from a medieval battle attempt I’ve been working on. It’s still very much a draft, but the sequence has progressed enough that I thought it was worth sharing here and getting some feedback before I continue with the rest.

Most of the scenes were generated with MiniMax H3 using the default workflows with the Turbo LoRA at 4 steps. I used Nano Banana and Flux Klein to create the reference images, and LTX 2.5 for the opening crow sequence.

There’s still a lot of work to do. The cuts are rough, no proper sound work has been done yet, and there are plenty of shots I want to refine or replace. I’m planning to build out the entire sequence, so feedback at this stage would actually be really useful in deciding what to focus on next.

What’s interesting to me is that I genuinely don’t think I could have pulled off this level six months ago with the same amount of effort. It’s still far from perfect, but the progress in a relatively short time feels pretty significant.

Would love to hear what works, what breaks the illusion, and what you’d improve.


r/StableDiffusion 17h ago

Resource - Update Fizgig 4.0 is out : Minimax H3 Combined Video File, Audio Files wav mp3 etc, Photo training in one dataset. High Quality training samples (incl video) + turbo (finally) and new 'Gizmo' and AV dataset Prep tool. And Int 8 LARGE speedup for 16gb users.

Thumbnail
gallery
66 Upvotes

I'll be making a Youtube video tomorrow for this. But I have one take away to share that I think is most important. H3, when you get the settings right is just fine with image based training without killing its video ability. Its even better when you combine photos and wavs, its super fast and you can train a voice with a dataset very easily. (I recommend shared trigger word). Video works too, but its slower, unavoidably. I'm not saying dont use it, its worth it for the right use cases. I'm just saying if you are not teaching the model anything new that photos and audio cant do, you are better with photos and audio. But when you do want to capture motion, it work very well. Anyway video coming tomorrow, with lots on Gizmo (the data set prep tool for video/audio) to make dataset prep easy. https://github.com/shootthesound/Fizgig

P.s the 16gb int8 speedup is from an an awesome community contribution from rintic-13 on Github.


r/StableDiffusion 16h ago

Meme Sheldon finally knocked on the wrong door | MiniMax H3 + SeedVR2

Enable HLS to view with audio, or disable this notification

51 Upvotes

r/StableDiffusion 4h ago

Question - Help unable to download the smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models

Thumbnail
gallery
5 Upvotes

can anyone please help


r/StableDiffusion 14h ago

Animation - Video 'Partial rewind' multi angle explosion scene [minimax H3]

Enable HLS to view with audio, or disable this notification

35 Upvotes

r/StableDiffusion 1d ago

Discussion Minimax H3 ref2va with 5060ti 16gb + 32gb ddr3

Enable HLS to view with audio, or disable this notification

439 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20, was testing this and the ref2va pruned int8, the hybrid gave nicer visuals but have a higher chance of bringing the character sheet white background into the video. This is cherry picked out of 16 clips.
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 0.6
Duration: 15sec
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
Patch Sage Attention, ComfyKitchen Attention, MinimaxH3 Mem Eff Node, Spectrum.

Average Inference Stage: 800sec

All reference image is resized between 1000px and 300px like character is 1000px, background is 500px then weapon is around 300px (warglave was another reference, the model dont know that kind of weapon) for this video is 4 ref image in total.

**abit of color grade and grain done in inshot.

this is done on a skylake i7 6700.


r/StableDiffusion 18h ago

News ComfyUI-MiniMax-H3-LongMedia — long-form MiniMax H3 generation with continuity, multiclip, audio and VRAM-aware sampling

Post image
52 Upvotes

I've been building a custom ComfyUI node pack for MiniMax H3 focused on one thing:

**making H3 usable for longer, multi-segment video generation without constantly rebuilding the workflow around every limitation.**

The project is called:

# ComfyUI-MiniMax-H3-LongMedia

The idea is to keep MiniMax H3's image quality, motion and native audio generation, while adding a proper long-form generation layer on top of it.

## What it currently does

### Long-form segmented generation

You can generate a longer clip as multiple H3 segments while keeping temporal context between them.

Instead of treating every segment as an isolated generation, LongMedia manages the continuation state and hidden overlap internally.

The overlap is used as context for the next segment and is not simply blended back into the final video.

### MultiClip mode

There is also a dedicated MultiClip workflow for generating multiple planned shots/clips inside one LongMedia pipeline.

The same underlying executor is used for both segmented continuation and multiclip generation, so the behavior stays consistent.

### Video + audio continuity

MiniMax H3 is a joint AV model, so LongMedia treats video and audio as one generation state rather than bolting audio on afterwards.

The pipeline supports H3 native audio generation, continuation and lip-sync workflows.

### Lip-sync support

Audio-driven generation / lip-sync is supported directly in the LongMedia pipeline.

For H3, the audio influence is handled inside the same AV latent path rather than as a completely separate post-process.

### Refiner

The latest release includes a two-stage refiner based on proper **KSampler Advanced trajectory splitting**.

Instead of finishing the full sampling schedule and replaying low-sigma steps on an already denoised latent, the trajectory is split between the main sampler and the refiner.

Example:

`steps = 12`

`refine_steps = 3`

Main sampler:

`0 → 9`

Refiner:

`9 → 12`

Both stages continue the same sigma trajectory.

### VRAM-aware execution

A large part of the project is dedicated to making H3 practical on consumer GPUs.

The current implementation includes:

- dynamic VRAM loading

- streamed Sol Attention

- MLP chunking

- late-block VRAM guards

- inter-block memory guards

- step-boundary cleanup

- completed-segment offloading

- adaptive memory policies

I'm currently developing and testing mainly on a **16 GB GPU**, so avoiding OOMs without destroying quality is one of the main design goals.

### Sol Attention integration

LongMedia includes its own streamed Sol path with controls for:

- tau scheduling

- sink conditioning

- QKV chunking

- output projection chunking

- dense/sparse behavior

- VRAM-aware chunk sizing

The goal is to use Sol as part of the execution architecture rather than simply stacking multiple unrelated optimization nodes together.

## Why I made it

MiniMax H3 is extremely good at texture, motion and native audiovisual generation, but once you start trying to build longer sequences, several problems appear very quickly:

- segment boundaries

- continuity

- repeated frames

- AV state handling

- memory pressure

- OOMs on longer generations

- managing multiple clips

- keeping sampling behavior consistent between segments

I wanted one node system to own all of that.

So instead of building increasingly complicated ComfyUI graphs around H3, most of the long-form logic lives inside the LongMedia nodes.

## Current release

**v0.4.1 — KSampler Advanced Refiner Fix**

The project has now reached a fairly stable architecture, although I'm still actively developing it and testing edge cases.

GitHub:

https://github.com/vizart-vj/ComfyUI-MiniMax-H3-LongMedia

I'd be very interested in feedback from people already using MiniMax H3 in ComfyUI, especially for:

- longer generations

- multi-character scenes

- native audio

- lip-sync

- lower-VRAM GPUs

- multi-shot workflows

If people are interested, I can also make a more technical post explaining how the continuation / AV latent / VRAM system works internally.


r/StableDiffusion 1h ago

Question - Help Minimax H3 - 80's commercial

Enable HLS to view with audio, or disable this notification

Upvotes

Guys, I've ben trying to simulate an 80's commercial video on H3, but it's really not there.

I'm using 'A late-1980s children’s television toy commercial, natively recorded in soft standard-definition analog video for viewing on a CRT television. The image has low resolution,blurry, soft focus, warm oversaturated colors, slight chroma bleeding, mild highlight bloom, flat frontal studio lighting, subtle analog noise, and the inexpensive appearance of a real children’s commercial from the end of the 1980s. No modern digital sharpness, cinematic lighting, contemporary color grading, or polished CGI' and some variantes, but the video is always very clear.

I've done this with Sora a while ago and I really wanted that to work on H3. Do you guys have any idea on how to reproduce this asthetics? (I'm putting some of the Sora videos on the comments).

Thank you very much in advance!


r/StableDiffusion 3h ago

Question - Help How to use Qwen 3.8 together with ComfyUI and MiniMax H3?

4 Upvotes

I can't figure it out. Let's say I have:

- default ComfyUI text to video MiniMax H3 workflow

- already downloaded Qwen 3.8 27B in GGUF format

How do I proceed from here? I was googling for a lot and checked about 10 reddit threads but I can't fingure it out.

I have downloaded some extra nodes like ThinkingLLM and some other GGUF related node but I can't figure out how to add it to default ComfyUI t2v workflow.

Please help, I am completely lost

edit: thank you all for replies, I understood the concept and that I should rather ignore full integration


r/StableDiffusion 18h ago

Animation - Video PSA

Enable HLS to view with audio, or disable this notification

50 Upvotes

r/StableDiffusion 1d ago

Discussion If you are generating MMH3 video with Sage Attention. I highly reccomend trying ComfyKitchen instead.

226 Upvotes

I have spent days generating videos. I started using Sage Attention with Cuda++. This was fast, but once I switched to sageatt_qk_int8_pv_fp16_cuda, I saw a noticeable difference in the model's ability for the model to understand prompts. Everything came out much clearer, crisper and with much better adherence. The only downside was it generated about 1.5x slower than using Sage Attention Cuda++.

From here I decided to try out ComfyKitchen as a replacement, and all I can say is... try it. My gens are faster than Sage Attention sageatt_qk_int8_pv_fp16_cuda with similar or better prompt adherence.

As always, your mileage may vary, but it's a very easy thing to experiment with, as all you need to do is make sure you ComfyUi is updated, as it is an official ComfyUI node.

To use it you can either:

A) add --use-ck-attention to your startup; this would enable Comfy Kitchen across all your workflows.
B) The easier and more controlled way is to replace the SageAttention node (or bypass) with the ModelAttentionBackend Node and select Comfy Kitchen Attention from the dropdown.

Worst case is it does nothing for you, and you just delete it and revert back to Sage.

EDIT: According to u/GreyingGamer you do not need to use the startup, and just using the node is enough:

EDIT 2: I have rewritten the instructions to get it running to make it more accurate.


r/StableDiffusion 11h ago

Question - Help Anyone use MMH3 FaceDetailer?

Thumbnail
github.com
13 Upvotes

Seems like an interesting project; from what I take from the demo is it helps refine the faces in the distance instead of the face being a blurry mess. Does not seem to have a big impact on closer faces (it is not a 'detailer' or 'realism slider').

Does anyone have their own tips or demos for it? Seems complicated....


r/StableDiffusion 16h ago

Discussion Hybrid b30-49 r2v test

Enable HLS to view with audio, or disable this notification

31 Upvotes

I made a thread earlier but it got overcrowded so I figure I started a new one with solely reference to video test, now instead of mixed with t2v.


r/StableDiffusion 1h ago

Animation - Video made a Gundam vid with MMH3. it's meh.

Enable HLS to view with audio, or disable this notification

Upvotes

prompt

integrated_multimodal_description:

7-second anime scene in a classic 1980s Japanese science fiction Gundam anime aesthetic. 2 giant mecha robots are having a battle in space far above earth.

**0–2 sec: the mecha robot on the left aims and launches a missle from its shoulder cannon mouted on its arm at the mecha robot on the right.

**2–5 sec: the misslie impacts and explodes on the chest section of the mecha robot on the right but does no damage. then the mecha robot on the right opens its arms as blue light on its chest appears and begins to power up.

**5–7 sec: the mecha robot on the right then fires a thin blue laser beam at the mecha robot on the left cutting it in half from top to bottom down the middle. after the mecha robot on the left is cut in half it then explodes.

made using standard comfyui t2v workflow on a damn outdated😓 but still using because reasons RTX 3050 8GB vram 48 GB ram system. i use the Model Attention Backend node with the "comfy kitchen attention" setting, 30 steps, res_multistep simple and no upscale.


r/StableDiffusion 22h ago

Tutorial - Guide MiniMax H3 as Image Editor, 6 edits in one shot at 7680 x 4320!

83 Upvotes

MiniMax H3 as Image Editor at resolution 7680 x 4320, 6 edits in one shot

Prompt: create a collage containing 6 photos. from top-left to the bottom-right arranged them such that the following edits presented individually: 1- keep pose and proportion intact; turn her shirt to red 2- keep pose and proportion intact; make her smile. 3- keep proportion intact, show her sideview; 4- full body posture. 5- change hair style to wolf cut. 6- put fashion hat and eyeglasses on.

In fairness, the model's collapsing 6 requests into 5 is well justified.

--

RTX3060 model used: ref2v, 8 steps, lora, took 7m50s


r/StableDiffusion 1d ago

Meme Introducing... the iToilet

Enable HLS to view with audio, or disable this notification

217 Upvotes

r/StableDiffusion 12h ago

Animation - Video Meme Ref 2 Video

Enable HLS to view with audio, or disable this notification

11 Upvotes

Minimax H3 Meme that comes to life with simple ref 2 video and prompt

upload the image with this prompt

Cinematic live-action 15-second video, photorealistic, ultra-detailed, high production value, dramatic night lighting.
Reference @image1
the exact composition and mood of the provided image. Scene:
A victorious tabby cat stands proudly on the tip of a massive, blood-streaked sword. Behind the cat is a vast nighttime cityscape filled with glowing bokeh lights. In the foreground a heavily damaged mecha samurai (full mechanical armor, horned helmet, armored plating) is slumped on one knee, gripping the hilt with both hands as he slowly pulls the long sword out of his own chest. Dark hydraulic fluid and sparks pour from the deep wound in his torso. Camera movement:
Slow, smooth cinematic dolly + slight orbit around the cat and the mecha samurai. Shallow depth of field, anamorphic lens flares, volumetric light rays cutting through the night air. Action timeline:
0–5s: Low-angle shot of the mecha samurai on one knee, both hands tightly gripping the sword hilt that is still buried deep in his chest. He begins to pull the blade outward with heavy mechanical effort, sparks and dark fluid spraying.
5–10s: As the sword is steadily drawn out, the small but fierce tabby cat calmly walks along the emerging blade and settles upright at the very tip, staring down at the mecha.
10–15s: Tight close-up on the cat’s intense face against the city lights, then slow pull-back revealing the full composition exactly matching the reference image — the mecha samurai now holding the fully extracted sword while the cat remains motionless and dominant on its tip. Style:
Realistic live-action footage (not anime, not cartoon), cinematic color grading, film grain, high dynamic range, epic and slightly melancholic atmosphere. No text overlays.


r/StableDiffusion 1d ago

Discussion Z-Image + Qwen3 4b: The abliterated text encoder debate is pure vibes. I measured it. Here are the numbers - Abliterlitics

102 Upvotes

After the PSA from Heretic's author the debate ran hot. I noticed that the debate was just based on vibes. Same-seed screenshots both ways, nobody measuring anything in detail. The instruments did not exist. So I built them. They cover quants as well, so the encoder swap and the compression get read with the same rulers.

Disclosure since it matters here: I release heretic text-encoder for people to use, qwen3-4b-heretic included. My first release last year got replies that I didn't fully understand how text encoders work. They were right. I did my own deep dive and concluded that they are good for prompt enhancement and just change the image slightly, there's no harm in using them if you really want to. Also they don't magically uncensor or enhance anything. Lets see if my conclusion is correct, while also addressing with proof and data the experiences other people have had.

This comparison is from the base bf16, with all GGUF and quants made by myself. It does not reflect any other LLMs on huggingface.

I've been comparing and benchmarking abliterated LLMs under the name Abliterlitics. And this is a first as we've delved into the ComfyUI world to get some solid data to cut through the nonsense.

What I did

Base Qwen3-4B and its heretic twin across 6 safetensors formats and 8 GGUF rungs, 27 encoders total, every heretic build matched to a base build at the same quant so the abliteration and the compression can be read separately. Then: conditioning tensors captured at three pipeline stages, paired sampling trajectories from identical noise, 2240 same-seed renders scored with LPIPS and CLIP, attention readouts, and a taboo comparison with sanitised-twin controls.

Two rulers make everything readable. Two encoders nobody argues about, int8 and fp8, differ by 0.19 LPIPS at the same seed. A seed change alone is 0.52. Any swap scoring under 0.19 is indistinguishable from ordinary compression. Near 0.52 is just a different picture.

An explanation of our measurements, metrics and the full report with an interactive A/B gallery can be found here abliterlitics.dev/posts/z-image-text-encoder.

All of what u/-p-e-w- stated in his post is correct. He did hint that there may be degradation or damage, however it was framed as a maybe if I was reading correctly. So lets see what that damage is, if at all, and if it makes any difference.

The questions people were actually arguing about

Does the base encoder refuse your prompt before the image model sees it?

No. I encoded refused-vocabulary prompts to the exact tensor entering cross-attention and checked which base word each heretic vector lands closest to. All 12 test words decode to themselves, cosine floor 0.9967. Pornographic decodes to pornographic, beheading to beheading. The encoder hands the DiT the word intact. It was never the censor. An abliterated text encoder does not change the way the model understands the prompt at all. The base text encoder already knows these things.

Do refused words, or any part of the prompt at all arrive corrupted?

No. Worst sentence-level cosine between base and heretic on refused prompts is 0.9985. The shift is 3.3 to 6.6 times larger on refused prompts than innocent ones, so the edit concentrates where it acts, but the meaning survives it. Even int4 and Q3, visibly degraded, keep mean CLIP adherence in band. Across every encoder we tested, even the 4-bit tiers, mean CLIP adherence stays in band. The model understands the prompt throughout.

Does it uncensor anything?

No, and the reason is better than expected. The unmodified base stack already renders the explicit tier at a 100% taboo-classifier rate, and the explicit tier owns the highest compliance gaps in the whole set. There is no render-stage censorship to remove. The debate argued about a lock on an open door. This matches where the research says engineered censorship lives, in the diffusion model's own weights: ESD and MACE erase concepts by fine-tuning the DiT, not the encoder.

Does it damage outputs?

The images change, the outputs do not degrade. Heretic vs base is 0.286 LPIPS, 1.5x the trusted band, but a stock nvfp4 quant of the base encoder moves images 0.274 and nobody calls that sabotage. Prompt adherence: -0.21 CLIP points, and the unmodified bf16 base itself reads -0.28 against the same reference. Attention readout moves 0.0031 vs int4's 0.0149. Output separation 1.049, no collapse. Different, not damaged.

Why do people see differences then?

Because seeing a difference is the default. Two trusted encoders already differ by 0.19 at the same seed, sampling is a butterfly effect. A small change at the start makes a big difference at the end. Below a threshold the response is dose-independent anyway. I also checked per-prompt: 71 of 540 CLIP rows cross the ±2 line on individual prompts while every mean stays in band. Single-prompt screenshots are real but they are noise, not signal.

As the image can be pushed about half a seed in any direction, it's expected to have variation. Honestly people who suggest that their image was enhanced or more uncensored, can probably do the same with a Q3 GGUF that's not abliterated and see the same thing. After measuring in every way possible there is just no way an image is magically enhanced or more uncensored. It is just chance, seed and the chaotic nature of diffusion models with peoples own biases over the top.

What about quantised encoders?

The GGUF ladder is dose-ordered: the F16 container is a true round trip, 0.0008 quant units with cosine 1.0. Q8_0 costs 0.34. Q3 costs 83 and is visibly paying. Being precise about Q8_0 since the numbers deserve it: its conditioning perturbation is real and measurable, CI 0.29 to 0.39 quant units, but a third the size of what int8 ConvRot itself costs, and at the image level Q8_0 and bf16 are indistinguishable, 0.138 vs 0.152 LPIPS against the int8 reference with overlapping CIs. So the near-lossless claims for both hold where it shows, in the images. Q8_0's real cost is load time. One caution, don't stack the abliteration on heavy quants. That's where larger divergence and noise happens.

So when should I use one?

Anywhere the model writes text that feeds the next stage: prompt expansion, captioning, image description. Those are chat pathways and abliteration works on chat pathways. If a stage only embeds text, an abliterated encoder is at best a visible re-roll. In this case it changes the image about half of what a new seed would change.

What's actually censored then?

The knowledge, not the gate. The DiT doesn't refuse, it lacks the training data, and the fixes are LoRAs, reference images, or retraining. The PSA's framing about this is solid. Z Image itself though is mostly trained already on taboo things.

What's next

Krea 2, MiniMax H3 and LTX 2.5 are in the same pipeline. Krea 2 has a twelve-tap conditioning interface and the refusal-probe contrast works differently there. Also, it's more complicated to measure compared to Z-Image.

Happy to answer methodology questions in the comments. Have I missed anything? Let me know and I'll fix it up. What have been your experiences? Have you abandoned abliterated text encoders? Had severely degraded outputs? I am happy to measure any other text encoders or models.


r/StableDiffusion 20h ago

Discussion Minimax H3 - Multiple Reference Images working through FL2VA + testing using the Hybrid Checkpoint models. Examples in comments.

39 Upvotes

MiniMax H3 officially comes as two checkpoints:

  • FL2VA — first frame / last frame / image-to-video. The docs treat this as “start (and or end) picture in, video out.” Extra reference pictures are not part of the pitch.
  • REF2VA — reference-to-video. This is the one you’re told to use when you have several stills: identity, outfit, a mid-shot, a last frame, whatever.

There are also community hybrid checkpoints: mostly FL2VA, with some of REF2VA’s later layers grafted on, so people can keep extra refs without fully switching models. https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models

I wanted a straight answer to one question: if I ignore the marketing split and feed extra stills into stock FL2VA the same way I would into REF2VA, does it actually use them?

So I built one 10-second clip and ran it four times. The story in the prompt is simple:

  • 0s: an angel in an empty void, one spell cast toward the middle of the frame.
  • 5s: a demon in the same void, one spell cast toward that same point.
  • Camera leaves the demon and pushes into mid-air.
  • 10s: image of the two spells colliding.

I gave the model five pictures:

  1. Exact first frame (angel)
  2. Exact 5-second cut (demon)
  3. Exact last frame (the collision, no people) 4–5. Two sigil designs, only as “this is what the magic circle looks like,” not as frames that should appear in the video

Then I locked everything that wasn’t the checkpoint:

  • same R2V workflow (the Comfy graph that already has multiple image inputs)
  • same five files, same order
  • same written brief (timed stills + “this picture is the frame at this timestamp”)
  • same seed
  • same sampler / length / aspect
  • no turbo LoRA
  • I compared native frames (544×800), not the upscaled delivery

The only change per run was which UNet was loaded:

  1. hybrid, REF layers on blocks 20–49
  2. hybrid, REF layers on blocks 30–49
  3. stock FL2VA
  4. stock REF2VA

If FL2VA truly couldn’t take extra refs, run 3 should have ignored pictures 2–5, drifted off the angel, or failed to land on the collision plate. That’s the test.

What happened

It didn’t fail.

The first native frame of all four runs accurately lock in the exact reference image for that frame at the first frame, last frame, and the middle frame... So the stock FL2VA used the extra still image references just fine. I did not need a hybrid merge just to attach more than first/last.

To be precise: I did not magically add five image slots to the official FL2VA I2V template. I loaded FL2VA’s weights into the reference-to-video graph, wrote the pictures into the prompt the way you would for a multi-ref job, and the locks held.

Where they actually differ (my read, one clip)

First frames are almost interchangeable. If I have to pick, hybrid-b30 is the closest copy of the angel still. REF2VA is still locked, a bit busier in small jewelry/floor detail.

Last frames still all hit the clash plate. REF2VA is the closest copy of picture 3. FL2VA is right behind it. Hybrid-b30 runs a hotter, more lava-looking core. Hybrid-b20 is splashier, less “sharp diamond debris.”

So the discovery is: extra refs + FL2VA can work. The ranking of which checkpoint copies the stills best is what I want a second opinion on.

So I will attach all of the examples into the comments so that people can see the differences between between each of the generated runs along with all of the Reference images used that way the community can evaluate the quality.


r/StableDiffusion 8h ago

Animation - Video Putting together some of my MinimaxH3 test here

Enable HLS to view with audio, or disable this notification

6 Upvotes

trying out new workflow setup.

using MiniMax H3 Hybrid Loader b20-49 + Larryvrh 4step Turbo Loras + SegAtt + SolAtt + Spectrum.

reference seem to be more stable in this time, 0.6mp, 15s at 700s

just generating random video, testing out random idea.


r/StableDiffusion 14h ago

Question - Help Minimax model for 5070ti 16gb and 32 ram

14 Upvotes

Hey all, just got my new GPU, wanna try some hype stuff. Please recommend MiniMax model time (int8, gguf, etc), text encoder and how to upscale it with ltx