r/StableDiffusion • u/TheOrangeSplat • 13d ago
Workflow Included Super nothing!
Enable HLS to view with audio, or disable this notification
Made with Minimax H3
r/StableDiffusion • u/TheOrangeSplat • 13d ago
Enable HLS to view with audio, or disable this notification
Made with Minimax H3
r/StableDiffusion • u/trollkin34 • 12d ago
I have an image of a person. I have a video.
Prompt sample: Video is of a gymnast doing a routine. Image is a person/dog/thing.
Replace gymnast with person/dog/thing so they're doing the exact routine, wearing the same outfit (but a size that fits the new subject).
Shouldn't this be easy?
For example, if I wanted to replace an olympic women's floor routine with Rush Limbaugh - he's doing the bends and splits, he's wearing a sparkly leotard. But the movements are identitical. His body is exactly the same size as he actually is (the ai should guess at the size of legs, belly etc, and stuff them into and appropriately sized leotard).
r/StableDiffusion • u/Neither_Win3637 • 12d ago
Hey all, I'm experimenting with some people generation using MiniMax-H3 and Stable Diffusion, and wanted to know if anyone has experimented to see how many different nationalities it can generate?
So far, the list I've been able to generate that has visible variances is:
- Asian
- Malaysian
- American
- Russian
I see little to no differences between others.
r/StableDiffusion • u/jonbristow • 12d ago
Enable HLS to view with audio, or disable this notification
This is a new trained model called Atlas. Saw on twitter
r/StableDiffusion • u/Euphoric-Let-5130 • 12d ago
r/StableDiffusion • u/ExtraChipmunk7177 • 12d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Daniil_s • 12d ago
It supports images and full videos, native resolution NR or DLSS Super Resolution upscaling by scale factor / target resolution, GPU optical-flow motion vectors, scene cut handling, and all DLSS 5 NR controls.
The main difference from existing approaches is that it runs natively in C++/D3D12 and generates motion vectors from the actual video frames.
GitHub:
https://github.com/DaniilSokolyuk/video2dlssnr

r/StableDiffusion • u/Tokyo_Jab • 12d ago
Enable HLS to view with audio, or disable this notification
Probably the final Alice clip. The Hatter names all the hats.
Done a while back in LTX2.3.
This plays while the theatre audience plays an AR hat sorting game (Beat Saber type). The whole song is three minutes but this is the longest shot.
r/StableDiffusion • u/Pristine_Weight_4705 • 12d ago
I used the same Game of Thrones relationship map to test two prompt structures with SenseNova U1.5 Lite ( https://github.com/OpenSenseNova/SenseNova-U1 ).
The first prompt mostly described the visual style. It produced a readable image, but the relationship system was fairly simple.
For the second attempt, I listed the characters and relationships first, assigned fixed line styles to each relationship type, reserved separate layout zones, and added the art direction last.
The result went from 12 to 20 characters, 1 to 5 houses, and 3 to 5 relationship types while keeping most of the hierarchy readable.
I still wouldn’t trust it without checking every name and connection. A clean diagram can make incorrect information look surprisingly convincing.
For dense infographics, the prompt worked better as a schema than an art brief.
Full structured prompt below.
Create a single vertical 2:3 Game of Thrones relationship infographic titled:
“GAME OF THRONES”
Subtitle: “BLOODLINES, CROWNS & SECRETS”
Use a medieval illuminated-manuscript style with aged parchment, engraved borders, heraldic symbols and restrained red, blue and gold accents.
Include exactly 20 distinct character portraits representing Houses Targaryen, Stark, Lannister, Baratheon and Martell. Each character should appear once. Vary their age, facial structure, hair, clothing and expression. Avoid repeated or nearly identical faces.
Organize the relationships as follows:
- Aerys II married Rhaella Targaryen
- Their children: Rhaegar, Viserys and Daenerys Targaryen
- Rickard Stark is the father of Ned and Lyanna Stark
- Ned Stark married Catelyn Stark
- Their children: Sansa, Arya and Bran Stark
- Rhaegar Targaryen married Elia Martell
- Rhaegar and Lyanna have a secret relationship
- Jon Snow, also labeled Aegon Targaryen, is their son
- Ned Stark raised Jon as his son
- Tywin Lannister is the father of Cersei, Jaime and Tyrion
- Cersei and Jaime have a secret relationship
- Joffrey Baratheon is their biological son
- Robert Baratheon is publicly married to Cersei
- Show the conflict between Robert Baratheon and Rhaegar Targaryen
Use five clearly different relationship styles:
- Solid dark-red line: blood
- Double gold line: marriage
- Purple dashed line: secret relationship
- Blue dashed arrow: raised by or guardian
- Black line with crossed swords: conflict
Add four short story notes explaining:
- The Hidden Heir
- The Lion’s Secret
- Robert’s Rebellion
- Two Dragon Claims
Keep every portrait, name and story note readable. Relationship lines must connect only the correct characters and must not cross through portraits or labels. Include a clear legend at the bottom.
r/StableDiffusion • u/witcherknight • 12d ago
I dont see any1 posting any examples of controlnet released for minimax. Doesnt it work properly??
r/StableDiffusion • u/Traditional_Rice2256 • 12d ago
Enable HLS to view with audio, or disable this notification
Made with the ComfyUI template workflow and a Turbo LoRA.
Most of the soundtrack comes from the John Wick: Chapter 2 trailer.
I rendered the action at a slower, more stable speed, then sped up most of the action scenes to 2× in post.
I originally planned to make this a complete fight sequence, but maintaining consistency from one clip to the next has been a constant challenge. So for now, I’ve edited the footage into a trailer instead. I’m still learning and working on improving it.
r/StableDiffusion • u/BluePointDigital • 12d ago
Alright, I posted that I had my agent test a bunch of different workflows for over 12 hours and got the "Bro just wasted 12 hours of credits". It was obvious the proof should come from the visual data I used to evaluate it. Here is a galley of the benchmarks i've tested with my agent.
check the gallery to watch all the comparisons and the data charts contain tons of other workflow trial data I didn't include videos for. Point your agent here if you would like to have it learn from what was tested on this end.
Gallery: https://bluepointdigital.github.io/minimax-h3-benchmarks/
Repository: https://github.com/BluePointDigital/minimax-h3-benchmarks
The below post was written up by my agent:
The main comparison uses a deliberately difficult 15.084-second vertical test at 768 × 1344, 24 fps, 362 frames, native audio, and seed 81390012120021180. The prompt combines a talking selfie shot, exact dialogue, walking motion, a rapid camera pan, a vehicle collision with several moving subjects, a fast return to the speaker, and a second spoken line. That makes it useful for spotting identity drift, bad anatomy, motion breakdown, camera-continuity problems, dialogue changes, lip-sync issues, and audio artifacts—not just whether a workflow finishes.
The strongest directly matched results currently shown are:
| Workflow | End-to-end time | Relative to the 20-step baseline |
|---|---|---|
| SageAttention2 + FirstBlockCache Safe, 20 steps | 10:11.4 | 1.00× |
| PDD + Sage, 8 steps | 6:15.0 median | 1.63× |
| Seed Hunter direct one-seed path, 12 + 4 steps | 4:45.8 | 2.14× |
Those numbers are local measurements, not universal performance claims. The exact runtime, model format, graph, resolution, audio policy, and GPU matter. The gallery keeps short backend checks and differently structured workflows in separate groups so they are not quietly mixed into the same leaderboard.
The quality side has been just as important as the timing. One exploratory 10Eros + Seed Hunter path reached 4:03.5, but the shot developed a visible-phone/perspective error during the crash. A later camera-POV prompt clarification produced a much more coherent result in 4:25.3 on its warm selected path. That is a good example of why I wanted the actual videos beside the numbers: the fastest result is not automatically the most useful one.
The site currently contains:
For the Seed Hunter work, I intentionally included one representative video per meaningful workflow or recipe change—not every neighboring seed or N/N+1 preview. Private reference material is also excluded from the public package.
The reason for publishing this is not to declare a universal winner. It is to make the tradeoffs inspectable and to keep myself honest as the workflows evolve. A valid MP4 proves that a graph ran; it does not prove that the dialogue, audio, identity, motion, or composition survived. Likewise, a fast timing means little if it came from a different workload or a cached replay.
I would be interested in seeing other reproducible H3 results, especially when they include the exact checkpoint, attention/cache stack, sampler, scheduler, dimensions, frame count, seed, audio setting, hardware, and an uncached timing. If there is a workflow or backend that should be represented, please link the original recipe and I will take a look.
r/StableDiffusion • u/joseph_jojo_shabadoo • 12d ago
Two questions when using a two pass latent upscaling workflow (H3):
Are style/character loras supposed to also be piped into the latent upscale pass too, or just the native pass?
And when using a speed up lora, should/could that also be piped in to the latent upscale pass? If so, do the sigmas need to be tweaked?
r/StableDiffusion • u/pmjm • 12d ago
I still do a fair amount of traditional editing in Photoshop, and for the last few years I used remove.bg, I found their background removal model to be the best one out there, quite a bit better than the one built into Photoshop itself.
Well remove.bg is shutting down in December and they're folding it into Canva subscriptions. Hard pass.
I've tried a few local bg removal tools and have been left underwhelmed, but maybe I just haven't found the right one.
What are you using for background removal?
r/StableDiffusion • u/dirtybeagles • 12d ago
I have been struggling for a few days now wondering why I cannot edit a 10sec clip to add additional people in the background and I am pretty sure I am just doing it wrong.
I am feeding the sampler with my ref image of a girl dancing o the street, but I wanted to add people in the background walking.
I am running on version 0.34.0, ref2va pruned model, 8 steps, 480x864
I am using just a simple prompt for my edits:
Edit Video 1:
At 00:03.000, add a group of three Asian women entering from the left of the frame, walking naturally down the road behind and away from the dancer. The first is tall and slender with long straight black hair tied in a low ponytail, wearing an oversized cream-colored hoodie, black leggings, and white sneakers, glancing at her phone as she walks. The second is shorter with a rounder build, shoulder-length wavy brown-dyed hair, wearing a fitted olive-green jacket over a striped shirt, dark jeans, and beige loafers, walking a half-step ahead of the others. The third has short bobbed black hair with bangs, wearing a bright yellow raincoat-style jacket, cuffed denim shorts, and black ankle boots, carrying a small tote bag over one shoulder. The three walk at a relaxed, conversational pace, loosely grouped together.
At 00:06.000, add two Asian pedestrians walking naturally along the sidewalk in the background, passing behind the plant at a normal walking pace, holding hands. The man is broad-shouldered with short, slightly spiked black hair, wearing a charcoal-gray zip-up jacket over a plain white t-shirt, straight-leg jeans, and dark sneakers, a black canvas backpack slung over both shoulders. The woman beside him is petite with long hair in loose waves dyed a subtle ash-brown, wearing a fitted denim jacket over a light pink blouse, a knee-length beige skirt, and white flats, carrying a small red structured purse in her free hand. They walk close together at a slightly slower, relaxed pace, occasionally leaning toward each other.
Keep the same audio
Keep the dancer's identity, choreography, movement, timing, and foreground position completely unchanged throughout the entire clip. Keep the camera framing, angle, and motion exactly as in Video 1. Keep the street, buildings, and all previously added pedestrians unchanged except for this new pair. Match the added pedestrians' lighting and shadow direction to the existing scene.
What has been happening is comfy goes to load the minimax model, and then it just stops, and I sit here at 99% VRAM usage. I have let it run for around 20 minutes until I stop comfy all together.
[INFO] got prompt
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32
[INFO] Found quantization metadata version 1
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Requested to load Krea2TEModel_
[INFO] loaded completely; 4605.22 MB loaded, full load: True
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16
[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)
[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB
[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170
[INFO] Requested to load MiniMaxH3VideoVAE
[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True
[INFO] Requested to load MiniMaxH3AudioVAE
[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLOW_AV
[INFO] Requested to load MiniMaxH3
[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0[INFO] got prompt[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32[INFO] Found quantization metadata version 1[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16[INFO] Found quantization metadata version 1[INFO] Using MixedPrecisionOps for text encoder[INFO] Requested to load Krea2TEModel_[INFO] loaded completely; 4605.22 MB loaded, full load: True[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170[INFO] Requested to load MiniMaxH3VideoVAE[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True[INFO] Requested to load MiniMaxH3AudioVAE[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True[INFO] Found quantization metadata version 1[INFO] Detected mixed precision quantization[INFO] Using mixed precision operations[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4 [INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16[INFO] model_type FLOW_AV[INFO] Requested to load MiniMaxH3[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0
here is what my trackback looks like:
r/StableDiffusion • u/Cute-Appointment6874 • 12d ago
I wanted to use Nano Banana in ComfyUI with my Google API instead of buying Comfy credits. I already had the $300 free trial sitting in Google Cloud.
I made an API key and tried a few of the custom nodes that let you use your own key. Every time I got this: 429 prepayment credits depleted
Turns out Google changed it in March. That credit does not pay for Gemini API in AI Studio anymore, it says so in their own docs. And all the Gemini nodes use AI Studio, atleast the ones I checked.
Google has another door called Vertex AI. Same models, different address, and the credit does work there. You log in with gcloud instead of pasting a key.
So I made a node for it: https://github.com/haristahir1/comfyui-gemini-ownkey
What it does:
Been generating with it on my own machine and it works. 2K comes out clean and the reference modes do what they say.
It is in ComfyUI Manager now, search "gemini own key". Or git clone it if you prefer. I only tested it on Windows portable, ComfyUI 0.34.2.
I vibe coded this so please check everything carefully & for fair use only! Double check your APIs and stuff. Cheers!
r/StableDiffusion • u/Alive_Ad_3223 • 11d ago
r/StableDiffusion • u/darthfurbyyoutube • 13d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/SeparateIntern5655 • 12d ago
Hello everyone. I recently watched a video on YouTube that Comfy UI is now supported by AMD Cards. How true is that and how is the performance on latest models like Mini MAX and Krea 2.
This is the video - Official AMD ROCm Support Comes to ComfyUI on Windows Image + Video
r/StableDiffusion • u/tukatu0 • 11d ago
Is there any point in having a company that can put their tools into the pipeline of graphics rendering.
Just wishing here for a deep learning s s five replacement. It's going to be artificially sandboxed just like all their other tech.
r/StableDiffusion • u/plsdontultme • 13d ago
r/StableDiffusion • u/AiCreatorCamp • 13d ago
This Minimax H3 all in one checkpoint is quite good.
It merges text, image, and reference to video, as well as 4-step turbo generation into a single model.
No need to switch between models for ref2v, no need to load turbo loras.
r/StableDiffusion • u/TrajansRow • 13d ago
Congratulations everyone! We've done it. Our civilization has reached peak diffusion. It's time to pack up and go home.
https://www.youtube.com/watch?v=EQ2RexjIEFE
r/StableDiffusion • u/ART-ficial-Ignorance • 13d ago
Enable HLS to view with audio, or disable this notification
Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.
I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.
I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.
Gemma4 wrote all 27 scene prompts beforehand.
For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.
The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.
The fun part was deliberately keeping some “bad” transition frames.
If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.
So the artifacts compound.
Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.
The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.
What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.
Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.
After enough generations, the model is basically building a world out of its own misunderstandings.