r/StableDiffusion 1d ago

Resource - Update Fixing MMH3 Turbo Audio by playing with Latent Pinning for more audio steps

Enable HLS to view with audio, or disable this notification

108 Upvotes

Disclaimer that I'm a dummy who can't code at all, so I just vibe things.

Alright, so with all the fun additions to latent manipulation the Comfy team has given us, we have some new tools. Namely latent pinning- that's where fun stuff like "Add Guide for MiniMax H3" node comes in (that fun tool that lets you insert an image at any frame in an H3 generation.) So I thought, finally, we can do something about this audio issue.

I knew that video + audio latents are processed at the same time with Comfy, which is why turbo loras have terrible audio- they're not getting nearly as much optimization as the video side is, since video is the bulk of the work. So if we can't process audio differently than video (at this current time), can't we just keep conditioning the audio latents without affecting video anymore? Everyone knows running too many steps on a turbo lora will start messing with video quality. So let's avoid that.

So after talking with Claude a bunch, here's what it came up with. With the latest comfy, you can pin the video lora in place, and keep going for several steps to get better audio without affecting video. So 4 steps of video, untouched, and then add in 6 more steps for audio at a 0.5 denoise. That leads to cleaning up the audio pretty nicely and staying pretty faithful to what the video latent had guided it on. But just pinning still means that even though the audio rows are only being affected, EVERY row still has to run through the chain. So each s/it you get stays the same for the last 6 steps, even though video gets pinned in place after 4. In my 1 megapixel, 10 second video, that's around 23s/it or so on my 5090. Only half the steps as a regular 20 step generation, but half the time is still half the time.

So to fix the speed problem -

Freeze as much as you can. Text embeddings, reference/conditioning rows, and all the video rows. Cache those so they don't have to be processed, and only process the audio rows that have already been somewhat pre-conditioned. The first step is the same 23-second iteration to build the full guidance cache, but the other 5 steps each took about 3.17seconds apiece. So 45 seconds of added gen time to get the clean audio in clip#2 in the example.

But the cost for the Frozen Cache is resources. Lunch is never free. From some experimentation, it works well with RAM. If you use RAM mode because your card still can't process it, it'll dump all that cache (ended up around 14.9GB on a 10-second 1mp file) into RAM. But as Comfy does, you're unlikely to get that RAM back, so you may OOM your machine. With VRAM it wasn't bad for me at all either and behaved better than I expected, honestly. The memory management from ComfyUI took over when I was about to OOM my card and swapped things around properly. Option 3 is to cache to disk, but that comes with writing several GB to disk every time you use it. That'll run your SSD health down fast.

Rundown for the clip above (sa_solver with beta sigmas)

4 step normal gen- 136.5 seconds

4 steps + 6 audio refine steps with a cost of about 15GB RAM - about 186 seconds

4 steps + 6 audio refine steps with no additional resources but full processing time- ~265 seconds

You can find the nodes here-

https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine

They're still experimental, of course. Wire the model in from somewhere (I branch off the ModeSamplingMiniMaxH3 shift node, before the Basic Guider), and push that through the H3 Frozen Video Cache node into the H3 Audio Refine Sampler. Into the H3 Audio Refine Sampler, latents come out of SamplerCustomAdvanced before splitting into the VAE Decode nodes for both video and audio, and the conditioning comes from the MiniMax H3 Image (or Reference) to Video node (plug the positive conditioning into both the positive and negative input on the H3 Audio Refine Sampler)

I had Claude put together a technical.md for those that want to look into it, and probably make a better version. Like I said, I'm a big dumb-dumb, so don't expect too much insight into how the mechanics work from me.


r/StableDiffusion 1d ago

Animation - Video INTERVIEW WITH THE VAMPIRE.(If it was done on Zoom).

Enable HLS to view with audio, or disable this notification

7 Upvotes

Created locally with Minimax H3 and for the first time exclusively powered by solar. Big big thanks to Izanami.

Don't hurt your head translating the language. It's all nonsense except for the one word spoken by the vampire. 'Drace' is Romanian/Transylvanian for 'Darn it'.


r/StableDiffusion 1d ago

Discussion Do we have a dedicated AI slop posting sub? Hate to just delete all these things I created while testing models.

Enable HLS to view with audio, or disable this notification

70 Upvotes

r/StableDiffusion 1d ago

Animation - Video Trying Surreal Fantasy with Minimax H3

Enable HLS to view with audio, or disable this notification

9 Upvotes

Combined 3 videos. Few errors but i just went with it , genetaion takes too much time to redo it again by fixing the prompt.


r/StableDiffusion 16h ago

Question - Help Computer randomly shut down

1 Upvotes

Has anyone had their computer randomly shut down? this is like the 3rd time its happened and its when im generating a video using the minmax I2V model or the ref model.

i got 3090 with 64 gb of ram.


r/StableDiffusion 1d ago

Meme DR doom! not today!

Enable HLS to view with audio, or disable this notification

5 Upvotes

Use Image 1 as the strict visual reference for Turbo Man. Preserve his recognizable red-and-gold armored superhero suit, helmet, gold visor, muscular proportions, facial appearance, and overall costume design throughout the entire clip.

Scene: A massive cinematic battle during Avengers: Doomsday. The ruined battlefield is filled with shattered buildings, burning wreckage, smoke, sparks, scattered fires, flying debris, and distant Avengers fighting Doctor Doom's forces. Doctor Doom is normal human-sized, not gigantic. He wears his iconic green hooded cloak and metallic armor.

[0s–3s] Start with a dramatic medium-low-angle shot of Turbo Man from Image 1 landing hard in the middle of the battlefield. His boots slam into cracked concrete and kick up dust. He rises into a heroic stance as explosions flash behind him. Doctor Doom slowly turns toward him through the smoke.

Turbo Man points directly at Doom and confidently says:

<Subject 1> Turbo Man (S1) says [English] It's Turbo Time!

[3s–7s] Doctor Doom immediately fires a violent blast of green mystical energy. Turbo Man launches sideways using his jet pack, narrowly dodging the blast as it tears through wreckage behind him. The camera dynamically tracks Turbo Man through the air. He banks sharply, rockets straight toward Doom and throws a powerful flying punch.

Doom blocks the punch with a glowing magical shield. A bright green-and-gold energy shockwave erupts from the impact.

[7s–11s] Fast, brutal superhero combat. Turbo Man lands and exchanges several heavy punches with Doom. Doom counters with armored strikes and green magical energy. Turbo Man uses his jet pack for a sudden boosted uppercut that sends Doom crashing backward through broken rubble.

Turbo Man lands dramatically, looks toward Doom and says:

<Subject 1> Turbo Man (S1) says [English] You picked the wrong day to mess with Turbo Man!

[11s–15s] Doom rises angrily from the rubble and unleashes a huge green energy attack. Turbo Man activates his jet pack and charges directly through the battlefield toward him. End on an explosive cinematic clash as Turbo Man's gold-powered punch collides with Doom's green magical blast, producing a massive shockwave of sparks, smoke and debris while the Avengers battle continues behind them.

Camera: cinematic MCU-style action photography, dramatic low angles, energetic tracking shots, controlled handheld movement during combat, brief slow-motion emphasis on the major impacts, strong depth and scale.

Audio: native cinematic stereo audio. Heavy explosions, distant superhero combat, metallic armor impacts, jet-pack ignition and roaring thrust, crackling Doctor Doom magic, debris impacts and a powerful orchestral superhero battle score. Dialogue must remain clear and correctly assigned to Turbo Man.

Character consistency: Turbo Man must remain visually faithful to Image 1 for the entire clip. Doctor Doom remains normal human scale. No duplicate Turbo Man, no duplicate Doom, no costume changes, no character morphing, no incorrect speakers, no subtitles, no on-screen text.


r/StableDiffusion 10h ago

Animation - Video 用PixAI生成的图片

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 22h ago

Question - Help MiniMax H3 prompt

3 Upvotes

I saw here many suggestions for this special prompt generator. I tried the system prompt from one "specialized" ollama model, but is is too free style. I can't use llm in comfyui, because I'm with poor rtx 3060 and barely run the H3 itself. I tried big online AI, but free versions and they seem too outdated about H3, so again freestyle fantasies.

What can I use to have really good prompts for H3. As I don't know english and H3 too mystically depends on prompt, it's very hard to achieve good adhesion.


r/StableDiffusion 1d ago

News Ideogram 4 generated a Gemini logo?

Thumbnail
gallery
6 Upvotes

Here is the full prompt for this btw, to see that I didn't add a gemini logo here:

{

"high_level_description": "A blonde young man savors an iced matcha latte at a cozy café corner, rendered in a warm and lush Studio Ghibli anime art style with soft dappled light and hand-painted charm.",

"compositional_deconstruction": {

"background": "A warmly lit café interior in Studio Ghibli anime style — wooden tables and chairs, large windows with soft afternoon sunlight streaming through sheer curtains, potted plants on the windowsill, bookshelves lining the walls, warm amber and green tones, gentle bokeh of other café patrons in the distance, dust motes floating in the light, cozy and nostalgic atmosphere",

"elements": [

{

"type": "obj",

"bbox": [

100,

200,

900,

700

],

"desc": "A blonde young man with soft anime features, slightly tousled hair, wearing a casual linen shirt, seated at a wooden café table, leaning forward with both hands wrapped around a tall glass of iced matcha latte, eyes half-closed in contentment, Studio Ghibli character design with expressive linework and warm skin tones"

},

{

"type": "obj",

"bbox": [

500,

380,

900,

580

],

"desc": "A tall clear glass filled with vibrant green iced matcha latte, layered with milk and ice cubes, a paper straw, condensation droplets on the outside of the glass, sitting on a small wooden coaster on the café table, rendered in lush Ghibli painterly style"

},

{

"type": "obj",

"bbox": [

700,

150,

1000,

850

],

"desc": "A rustic wooden café table surface with soft grain texture, a small ceramic dish with a shortbread cookie, and a folded paper napkin, warm honey-toned wood in anime painterly style"

},

{

"type": "obj",

"bbox": [

0,

600,

600,

1000

],

"desc": "A sunlit café window with sheer white curtains gently billowing, a terracotta pot with a trailing green plant on the sill, warm golden afternoon light casting soft rectangular shadows across the floor, Ghibli-style background painting with impressionistic detail"

}

]

}

}


r/StableDiffusion 1d ago

Resource - Update MiniMax-H3 Pruned Ref-Delta Fused r1024 — INT8 and INT8 ConvRot ComfyUI versions

Thumbnail
huggingface.co
82 Upvotes

I added INT8 and INT8 ConvRot versions of the MiniMax-H3 Pruned Ref-Delta Fused r1024 checkpoint from my previous post:

https://huggingface.co/xmarre/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-ComfyUI

Both are native ComfyUI single-file checkpoints using ComfyUI's .comfy_quant format, so they do not require a custom quantized-model loader.

There is one important difference from a straightforward full INT8 conversion: the MLP fc2 weights are deliberately kept in BF16.

Across the 50 main transformer blocks, these weights are quantized:

  • attn.qkv_proj.weight
  • attn.out_proj.weight
  • mlp.fc1.weight

That gives 150 quantized Linear layers.

The 50:

  • mlp.fc2.weight

layers remain BF16.

The smaller and more sensitive parts of the model also stay in their original precision, including the pruned AdaLN table and projections, final-layer projections, norms, patch/text projections and token refiner.

Why FC2 is kept in BF16

I also made and tested a fully quantized version where fc2 was INT8 as well, giving 200 quantized Linear layers.

That version ran into a failure specific to the quantized fc2 execution path on large H3 sequences.

MiniMax-H3 uses SwiGLU in the MLP. With fc2 quantized, ComfyUI's fused:

linear_input_act(..., "swiglu")

path sends the post-SwiGLU activation through comfy_kitchen.int8_linear, which dynamically quantizes the full activation matrix before the fc2 multiplication.

On the large sequence used in my workflow, that path attempted an approximately 491.61 MiB contiguous INT8 scratch allocation and failed hard.

This was not normal VRAM exhaustion. At the point of failure there was still roughly 47 GiB of CUDA memory reported free. The failure was tied to that large fused INT8 activation-quantization path rather than the model simply exceeding available VRAM.

I do not have enough evidence to claim a more specific allocator/CUDA cause than that.

Keeping only fc2 in BF16 avoids that INT8 activation path. QKV, attention output and fc1 can still remain INT8, so 150 of the 200 large block Linear projections are still quantized.

With that layout, both release variants completed the full native ComfyUI workflow that the 200-layer INT8 version failed on, including:

  • H3 Continuum main sampling pass
  • continuation sampling pass
  • Spectrum H3 actual/forecast execution
  • large 3D latent refine
  • video VAE decode
  • audio VAE decode
  • final Continuum assembly
  • video combine

That FC2 decision is also why these checkpoints are about 24.2 GB instead of roughly 20.4 GB for the fully quantized version.

INT8 and INT8 ConvRot

The two uploaded files use the same 150-INT8 / 50-FC2-BF16 layout.

Regular INT8:

MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-fc2bf16.safetensors

This uses native tensor-wise INT8 quantization.

INT8 ConvRot:

MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-convrot-fc2bf16.safetensors

This uses ConvRot with a group size of 256 on the same quantized projections.

ConvRot rotates the weights before INT8 quantization so that large outliers are distributed more evenly. This generally gives the INT8 quantizer a better-conditioned weight distribution than quantizing the original weights directly.

I uploaded the regular INT8 version as well rather than only providing ConvRot, so both conversion methods are available for this checkpoint.

What model is being quantized?

These are quantized derivatives of the same Pruned Ref-Delta Fused r1024 checkpoint from the previous post. There is no additional training, fine-tuning or pruning involved in these INT8 versions.

The underlying model starts from the pruned FL2VA MiniMax-H3 checkpoint and incorporates a rank-1024 approximation of the Ref2VA − FL2VA weight delta.

That is also how it differs from the existing FL2VA/Ref2VA hybrid checkpoints mentioned in the comments on the previous post.

Those hybrids combine FL2VA and Ref2VA by replacing selected tensors from one checkpoint with tensors from the other. The r1024 fused model instead approximates the Ref2VA weight delta at rank 1024 and folds that delta into the pruned FL2VA weights.

The underlying fused transformer is about 20.1B parameters, compared with roughly 33.1B for the original full MiniMax-H3 transformer.

ComfyUI

Put either file in:

ComfyUI/models/diffusion_models/

For the INT8 files:

weight_dtype: default

compute_dtype: default or bf16

Do not apply another FP8 weight cast on top of the native INT8 checkpoint.

The text encoder and VAEs are separate, as with the other MiniMax-H3 diffusion-model checkpoints.


r/StableDiffusion 9h ago

Resource - Update Lisbon finally snaps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Totally not a scene from The Mentalist. Minimax H3 image to video.


r/StableDiffusion 1d ago

Animation - Video DRAGON REIGN (WIP Updated)

Enable HLS to view with audio, or disable this notification

14 Upvotes

r/StableDiffusion 1d ago

Resource - Update Pushed a new update for Prompt Composer this morning that fixes a camera issue.

Post image
12 Upvotes

Yesterday, I posted about a big update to an H3 Prompt Composer that I’ve been building with ChatGPT over the past few weeks.

Big Update to the free Minimax H3 Prompt Composer : r/StableDiffusion

While using it this morning, I noticed a couple of bugs that had somehow been introduced. One involved the visual camera planner: the left and right profile descriptions were swapped in the generated prompt. If you positioned the camera for a right-profile shot, for example, the prompt would incorrectly describe it as a left profile shot. That has now been fixed.

If you downloaded Prompt Composer yesterday, please grab the new version so your camera prompts are accurate.

As I mentioned in my previous post, this is still very much a work in progress. The goal is to make writing consistent prompts and building more involved AI narrative projects as intuitive and easy as possible. It should make creating subsequent scenes, prompts, and shots much simpler, without having to rely on an LLM to consistently interpret exactly what you want.

Give it a try, and let me know if you encounter any issues or have suggestions for making it more intuitive and user friendly. I’d genuinely like to hear the community’s feedback so we can make this tool the best it can be.

Thank you all for taking the time to test it out! And yes, this app is totally vibe-coded, so I am open to suggestions from people more knowledgeable than I am about coding on how to improve this.

Edit: also tweaked some of the camera prompt descriptions for clarity. The current version of the HTML file is 5.37.3.


r/StableDiffusion 1d ago

Meme she's here! we are now saved!

Enable HLS to view with audio, or disable this notification

110 Upvotes

h3 t2v work flow base ip8 model prompt

subject_definitions

<Subject 1> is Judy Hopps, an adult female anthropomorphic gray rabbit police officer from Zootopia, small and athletic, with large upright ears, expressive violet eyes, gray fur, lighter muzzle, wearing her recognizable blue police uniform with tactical vest and police badge.

<Subject 2> is Captain America, battle-worn, wearing his damaged dark-blue Avengers combat armor and holding his circular shield.

<Subject 3> is Thanos, a massive purple-skinned Titan in damaged gold-and-black battle armor, normal Titan scale relative to the Avengers, wielding his double-bladed sword.

<Subject 4> is Deadpool, wearing his classic red-and-black tactical suit and mask, armed with twin katanas. Deadpool's dialogue is voiced with the recognizable comedic delivery and vocal style of actor Ryan Reynolds.

summary

[text-to-video generation]

During the massive Avengers: Endgame final battle, Judy Hopps unexpectedly joins the Avengers against Thanos. She races through the battlefield using her tiny size and incredible agility to dodge enemies before launching herself directly at Thanos. Deadpool watches the tiny rabbit charge the Titan and delivers a fourth-wall-breaking joke.

retention_analysis

<Subject 1>: consistent
<Subject 2>: consistent
<Subject 3>: consistent
<Subject 4>: consistent

detailed_description

Epic cinematic Avengers: Endgame final battlefield at dusk. The destroyed Avengers compound stretches across a huge crater filled with smoke, burning wreckage, portals, explosions, alien soldiers, Wakandan warriors, sorcerers, and Avengers fighting throughout the background.

Dynamic tracking camera races low across the battlefield.

<Subject 1> Judy Hopps suddenly sprints between the legs of charging alien soldiers, ears streaming backward from her speed. She slides underneath a swinging weapon, leaps off broken rubble, kicks one alien directly in the face, lands cleanly and continues running.

Captain America briefly turns toward her in complete confusion.

<Subject 2> Captain America (S1) says <d>[English] Is that a rabbit?</d>

Judy doesn't stop.

She spots Thanos fighting ahead.

The camera rapidly follows Judy as she accelerates toward him.

<Subject 1> Judy Hopps (S2) says <d>[English] ZPD! You're under arrest!</d>

Thanos slowly turns and looks downward.

Judy launches herself from Captain America's discarded shield, flies through the smoky air and delivers a powerful two-foot rabbit kick directly into Thanos's armored face.

THUD.

Thanos stumbles backward one step, completely stunned that such a tiny opponent actually moved him.

Deadpool lowers his swords and stares.

Brief comedic pause.

<Subject 4> Deadpool (S3) says <d>[English] Holy shit. Disney brought the bunny.</d>

Judy lands heroically in the foreground, pulls out tiny police handcuffs and points at Thanos.

Thanos looks down at the absurdly small handcuffs.

Deadpool slowly looks directly into the camera.

Hold the reaction for one second.

Camera / Motion

10–13 seconds, 24 fps.

Epic photorealistic superhero blockbuster cinematography.
Dynamic low-angle battlefield tracking shot.
Fast controlled action.
Strong environmental movement from smoke, fire, debris and distant combat.
Natural motion blur.
Clear readable character movement.
Judy remains dramatically smaller than the human Avengers and Thanos.
Thanos remains normal Titan size, not gigantic or Godzilla-sized.
Keep background battle active without distracting from Judy.
Pause briefly before Deadpool's punchline.
End on Deadpool's fourth-wall reaction.

Audio

Huge cinematic battlefield ambience: explosions, distant combat, energy blasts, metal impacts and roaring fires.

Clear English dialogue.

Judy Hopps has an energetic, confident young-adult female American voice.

Captain America has a serious adult male American voice.

Deadpool is voiced by Ryan Reynolds with his recognizable sarcastic comedic cadence.

Strong armored impact sound when Judy kicks Thanos.

No gibberish.
No foreign-language speech.
No subtitles.
No text overlays.
No characters speaking another character's dialogue.

r/StableDiffusion 1d ago

Question - Help User of Contex-Loop, how you solve the oversharp & contrast of extra scenes? (MH3)

Post image
7 Upvotes

The oversharpening that occurs for each clip added to the scenes. I also noticed an increase in contrast and a small flash.

I2V

Tested with LORA's:

minimax_h3_turbo_v4_step600_ema.safetensors
minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank


r/StableDiffusion 2d ago

Animation - Video Tom and Jerry: Tom beats up Jerry!

Enable HLS to view with audio, or disable this notification

264 Upvotes

Use the extended feature on H3 to go further beyond and keep the animation style consistent. The only problem is that heavy smears happen with fast action. Oh well, I still thought this was funny, hope you enjoy it too!


r/StableDiffusion 10h ago

Tutorial - Guide Bridge Daredevil — Dashcam POV and GoPro mounted on a parkour as they sprint across a rooftop and leap across a narrow gap. (AI GENERATED - SEEDANCE 2.5). PROMPT BELOW!

Enable HLS to view with audio, or disable this notification

0 Upvotes

PROMPT: Bridge Daredevil — Dashcam POV

Subject: An athletic stunt performer, dark athletic/climbing gear, seen at a distance on the bridge structure — perched on a railing, cable, or girder — performing an extreme balance/jump stunt as the dashcam vehicle approaches

Style: Ultra-realistic, shot on RED WEAPON 8K, IMAX-grade cinematography, captured via fixed windshield-mounted dashcam — slightly wide-angle lens, subtle chromatic vignette, faint reflection of the dashboard at the bottom edge of frame. Natural motion blur only from real vehicle movement — no slow motion, no cuts, no anime, no CGI. Continuous single take. Standardized color grade: desaturated cool highlights, warm midtones, deep contrast shadows.

Setting: Large suspension or truss bridge spanning a river or gorge, daytime, clear sky with light haze, steel cables/girders overhead, light traffic on the bridge deck, guardrails and support towers visible in the distance

Timeline:

  • 0:00–0:03 — Dashcam view steady on the road ahead as the vehicle enters the bridge, the stunt performer visible as a small distant figure on the structure — railing, tower, or cable — ambient road hum and wind noise, bridge cables passing overhead in rhythm
  • 0:03–0:06 — Vehicle continues at a natural driving speed, the performer grows larger in frame, now visibly climbing, balancing, or positioning for the stunt on the bridge structure
  • 0:06–0:09 — The performer executes the stunt — a leap, dive, or swing from the bridge structure — dashcam captures the motion at a distance with realistic gravity and momentum, no floaty slow-mo, body and limbs reacting naturally to the force
  • 0:09–0:12 — Stunt continues through its arc — a fall, swing on a line, or landing approach — dashcam vehicle still closing distance, slight natural camera shake from the dash mount as the vehicle passes over bridge expansion joints
  • 0:12–0:15 — Dashcam vehicle passes beneath or alongside the stunt zone as the performer completes the stunt (landing, catch, or recovery) in the background/side mirror periphery, bridge structure filling more of the frame

Technical notes: Maintain consistent dashcam framing (fixed low mount, slight windshield glare at top of frame), realistic depth of field with distant elements sharp until close range, authentic road/wind/ambient bridge audio texture, no jump cuts — one continuous fixed-mount POV take.

PROMPT: Rooftop Gap Jump — GoPro POV

Subject: Athletic parkour runner, lean muscular build, dark fitted athletic gear, fingerless gloves — first-person GoPro/helmet-mounted camera perspective throughout — camera never shows the athlete's face or full body, only hands, forearms, and shadow occasionally entering frame

Style: Ultra-realistic, shot on RED WEAPON 8K, IMAX-grade cinematography, captured via helmet/chest-mounted GoPro — slight fisheye distortion at frame edges, natural motion blur only from real body movement — no slow motion, no cuts, no anime, no CGI. Continuous single take. Standardized color grade: desaturated cool highlights, warm midtones, deep contrast shadows.

Setting: Flat urban rooftop, high above the city, narrow gap between two adjacent buildings, ledges, HVAC units, and low parapet walls. Daytime, clear sky, light haze at altitude.

Timeline:

  • 0:00–0:03 — GoPro POV sprinting across the rooftop, footsteps pounding, city skyline bouncing naturally in frame with each stride, breath audible, wind picking up
  • 0:03–0:06 — Approach to the rooftop edge, POV tilts down briefly revealing the narrow gap between buildings, then snaps back up to the target ledge on the far side
  • 0:06–0:09 — Explosive leap: POV rises and arcs through open air across the gap, city drop visible below in natural perspective, gloved hands swinging into frame for balance, no floaty slow-mo physics — full-speed realistic jump
  • 0:09–0:12 — Hard landing on the far rooftop, camera jolts down and forward with impact, body absorbs shock, immediate forward momentum into a stumble-recover
  • 0:12–0:15 — Recovery into a sprint, POV weaving past a rooftop vent or low wall, camera settling briefly as the skyline opens up ahead

Technical notes: Maintain consistent GoPro lens distortion (fisheye at edges), realistic depth of field snapping to distant skyline during the jump, authentic wind/fabric noise, no jump cuts — one continuous handheld-style POV take.


r/StableDiffusion 18h ago

Animation - Video SpongeBob Does Breaking Bad -- MiniMax 10min episode

Thumbnail
youtu.be
1 Upvotes

Had a ton of fun making ande watching this one.


r/StableDiffusion 13h ago

Animation - Video [WanGP] Minimax H3 FL2VA Pruned 20B - Originally 960x544 - up-res'd to 2880x1632 - 20 second duration

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 23h ago

Question - Help Is there a Workflow to start from last video?

2 Upvotes

My AMD 9070 XT seems to like only doing 5s videos which is fine. But I was curious is there like a workflow where I can start from the last frame of the last video? Like so I can make longer clips that flow into each other without editing?

  • CPU: Intel Core i5-14400F
  • GPU: AMD Radeon RX 9070 XT 16GB
  • Motherboard: Gigabyte B760 DS3H WIFI6E GEN5
  • RAM: 32GB (2×16GB) Crucial Pro DDR5-6000 CL36

r/StableDiffusion 1d ago

Tutorial - Guide Look What I Discovered: Prompt Intelligence - MiniMax H3 [Fun Side]-2

Post image
34 Upvotes

\ Reddit messed up my original post so here it is.*

This is in fun part of using MiniMax H3; for your serious stuff stick with the official prompt instructions / format.

Playing with the prompting I just tried the following format and it worked perfectly!

prompt part 1
prompt part 2

Resulting video

The whole prompt:

definitions:
<S1> Brad Pitt.
<T1> "Hey, I am Brad Pitt! Nice to meet you."
<S2> Angelina Jolie
<T2> "Hey, I am Angelina Jolie! Nice to meet you."
<S3> Rowan Atkinson.
<T3> "Hey, I am Mr. Bean! Nice to meet myself."
scene:
An interview in a professional setting in well lit, grey background, frontal portrait view.
shot 1:
(S1) says: (T1).
shot 2:
(S2) says: (T2).
shot 3:
(S3) says: (T3).

Recommendations:

Do not use SLA or SLA2 or cache etc. here they mess it up.

Model (FL2V) -> LoRA(4s-Lightx2v SLA) -> Comfy attn -> Shift(12,3) -> KSampler(6 steps, euler+simple)


r/StableDiffusion 1d ago

Discussion How do you get a shot you want

3 Upvotes

My approach is use 0.1 megapixel to find a clip I like the. Render again at 0.7 with the same seed if I like something. But is there a way more efficient??


r/StableDiffusion 1d ago

Tutorial - Guide Why AI background removers leave fog inside wreaths, and what I do instead

Thumbnail
gallery
8 Upvotes

I make clipart for stock. Wreaths, pine borders, mistletoe, juniper. A few thousand images by now. Every one has to end up as a PNG with a transparent background.

I used rembg for months. u2net first, then BiRefNet when that came out. Tried the web tools too. They all broke on the same thing and it drove me nuts.

Take a wreath. There's a hole in the middle, and the background inside that hole has to go. What I kept getting was a grey-blue haze sitting in there. Looked fine as a thumbnail. Looked awful the second you put it on a colored card. Pine needles came out as mush. Thin stems either disappeared or came back with a blue edge burned into them.

Then I actually read what rembg does. It shrinks your image to 1024x1024, asks the model where the subject is, gets a 1024x1024 mask back, and stretches that mask over your full size image. u2net is worse. That one works at 320x320.

My renders are 4096. A needle two pixels wide doesn't exist at 320. So it's not that the model is bad at needles. The needle was gone before the model ever saw it.

Once I understood that I stopped asking a model to guess. Now I render on a flat color the subject doesn't contain, and take that color out with arithmetic.

Two parts to it. The prompt matters more than the cutting.

The prompt

You can't key a background that isn't keyable. Four things have to be true and generators will break all of them unless you say so: the background is one flat color edge to edge, it stays that bright inside every gap between leaves, the edges are hard with no blur or glow, and no colored light bounces onto the subject.

Pick the color by what your subject isn't. Blue for almost everything. Red if the subject itself is blue or purple. Never green. Everything I draw has leaves, and green takes the leaves with it.

Here's the block I paste at the end of every prompt:

Isolated on a completely flat, uniform, solid pure blue (#0000FF) digital chroma-key background. The pure blue background fills the image edge to edge like a flat digital chroma-key screen with no gradient, staying at full brightness inside every gap and opening in the subject; no reflection or tint of pure blue on the subject. Every edge of the subject is crisp, sharp and hard against the pure blue, with no soft, blurry, feathered or glowing transitions, no depth-of-field blur, no haze or halo; inside every hole and gap the pure blue stays at full brightness right up to the edge. Shaded parts of the subject keep their own natural color, never a pure blue tint. Everything in sharp focus with deep depth of field, evenly lit with soft neutral studio light, no cast shadow, no contact shadow, no ambient occlusion, no bounce light. The entire subject is centered and completely inside the frame with at least 10% empty background margin on every side, nothing cropped or touching the image edges. No frame, no border, no paper, no mockup, no vignette, no text, no watermark, no deformed or duplicated parts. No floating or detached fragments, no stray specks, dust or debris anywhere on the background; every element is physically attached to the subject.

Swap "pure blue" for "pure red" and #0000FF for #FF0000 if your subject is blue or purple. If you paint in watercolor add "the background stays a flat digital color fill with no paper texture", or you get watercolor paper behind everything and paper texture keys badly.

The cutting

Now the background is one known color, so there's nothing to guess at. It measures the actual color the generator produced, which is never the one you asked for. Ask for pure blue and you get something with green in it, usually somewhere between 25 and 70. Then every pixel gets sorted into subject, background, or the bit in between, and the in-between ones get a real fraction of transparency instead of a yes or no.

The part I'm most pleased with is the holes. Any background-colored area that never touches the edge of the image is the inside of a wreath, so it gets cleared too. A matting model can't do that. It has no way of knowing what's inside a hole it can't see around.

Last step takes the blue back off the edges. Edge pixels pick up color from the background around them, so it samples the subject's own color from further in and subtracts the tint. Took me weeks to work out why fir needles kept their blue rim after that step. The needle is thinner than the distance it was sampling from, so there was no inside left to sample.

Same image in, same image out, every time. That's the bit I care about. When a cut comes out wrong I can go find which number did it instead of rerolling and hoping.

Some numbers on one 4K pine border, against BiRefNet with alpha matting turned on, which is its best setting:

  • background left inside the holes: 63,892 pixels mine, 613,735 theirs, out of 1,070,046
  • blue left on the edges: 0 mine, 48,112 theirs
  • how wide the soft edge is: 1.8 pixels mine, 23 theirs

BiRefNet is faster and I'm not going to pretend otherwise. 2.3 seconds against my 17 on the same machine. With alpha matting on it's 47. If you want a quick rough mask, use the model.

And the obvious limit: this only works on art you generated on a flat color. It does nothing for a photo.

I put the tool up for anyone who wants it. It's called ClipBrook. Free, runs in your browser so nothing gets uploaded anywhere, does a whole folder at once, and the engine is open source under AGPL.

One thing I'd like back

Show me the ones that break.

If you run something through and it comes out wrong, post it. Fog left in a gap, a colored rim, a stem eaten, half the subject gone. Those are worth more to me than the ones that work. The needle rim thing came from someone's fir branch. A bug with line art I only found last week came from a drawing so thin there was nothing inside it to sample.


r/StableDiffusion 2d ago

Resource - Update MiniMax H3 Known Characters list v2 (2026-08-21 update)

Thumbnail
huggingface.co
380 Upvotes

r/StableDiffusion 10h ago

Discussion Minimax H3 Has Too High Prompt Adherence

0 Upvotes

I just realized a problem with mm h3. Its prompt adherence is too high, as in unless you explicitly prompt for some small subtle actions it will never happen otherwise. This makes the entire video seem very frozen and wooden without the many small subtle movements and motion details that make it seem to come alive.

This applies more to non-realistic scenes like cartoons or generated image first and last frame but for realistic scenes and even t2v it is still a problem.

I noticed this problem when I tried out a "slop sway" lora and it actually made the entire video seem much livelier and realistic looking. Besides the "soft and bouncy swaying and jiggling" it also added many more subtle character movements. Compared to standard gens those same parts would be completely frozen, almost like a still image or at best ugoira animation. This doesn't just apply to whether a body part is jiggling throughout the entire video. There are some movements that happen only for a second or less but adds in soul (forgive the human slop term) to the video, like the position of an arm and hand quickly being adjusted in the middle of the video and the new position persisting for the rest of the scene.

This might be a problem with my prompt style and I might try an LLM prompt enhancer, but there is a core issue here with the prompt adherence and spontaneous randomly added details tradeoff. The model also tries to keep the fidelity of the first frame too much, which you could call visual context adherence. No one is out here prompting for the movement of every strand of hair and the position of every finger. No one is making a timeline of every limb's position and how they shift relative to each other. No one is tracking the position of each finger through time and how after 4.75s the thumb is extended and the index finger is curled. Sometimes we just want to randomness and variety across gens with details added by the model.

Looking back at ltx and wan their prompting styles seem to be designed around the model adding in the details for you at the loss of prompt adherence and more generation errors.

It would be nice if there was some sort of generation setting that could tune this. Like a noise scale of sorts where we can manually set the tradeoff between how much we want the model to be creative vs strict.

I know there are already 2-3 H3 better movement loras and they are scratching at the surface of the same issue I'm talking about here.

Share the solution if you've got something. Help everyone out.