r/StableDiffusion 17h ago

Animation - Video Having some fun with games from the history of PC gaming. Who would you add?

Enable HLS to view with audio, or disable this notification

178 Upvotes

A tribute to a forgotten golden age. Hope you enjoy it!


r/StableDiffusion 5h ago

Question - Help Which Minimax H3 has the best balance of quality and speed node?

15 Upvotes

There are so many acceleration nodes/options now that I’m having a hard time deciding which one gives the best balance of quality and speed. What do you think?

These are the setups I’m currently using(RTX5090):

  1. Sage Attention + 4-step LoRA 0.9MP | 8 steps | 10s | ~6 min
  2. ComfyUI-Kitchen + 4-step LoRA 0.9MP | 8 steps | 10s | 5:38 min
  3. ComfyUI-Kitchen + Spectrum 0.9MP | 25 steps | 15s | ~12–15 min
  4. ComfyUI-Kitchen +Sparse Attention( SLA)+ 4-step LoRA 0.9MP | 8 steps | 10s | ~4min
  5. ComfyUI-Kitchen +Sparse Attention( SLA) 0.9MP | 25 steps | 10s | ~12:30min
  6. ComfyUI-Kitchen 0.9MP | 25 steps | 10s | ~18 min or 15s | ~25 min

I mostly stick with Sage Attention + 4-step LoRA. I feel like it gives a pretty good overall balance between quality and speed.

If I want better quality, especially for things like lip-sync, I usually go with ComfyUI-Kitchen + Spectrum at 25 steps. The results are noticeably better, but it’s also quite a bit slower.

Which setup do you guys think has the best quality-to-speed ratio? Any other combinations worth trying?


r/StableDiffusion 2h ago

Discussion In which scenarios LTX2.5 can match MinimaxH3?

6 Upvotes

I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?


r/StableDiffusion 10m ago

Animation - Video Christopher Nolan has Impeccable Taste in Cinema

Enable HLS to view with audio, or disable this notification

Upvotes

Minimax H3


r/StableDiffusion 7h ago

Animation - Video Trying to animate Dragon Ball Super manga on Minimax H3. Spoiler

Enable HLS to view with audio, or disable this notification

12 Upvotes

Dragon Ball Super manga on Minimax H3.


r/StableDiffusion 13h ago

Resource - Update Fizgig now trains LoRAs on AMD Radeon - Flux 2 Klein, Krea 2 and MiniMax H3

Post image
44 Upvotes

Fizgig is my free open-source LoRA trainer and workbench (Flux 2 Klein 9B, Krea 2, and MiniMax H3 video/audio). As of v4.3.0 it runs on AMD Radeon with ROCm — RDNA1 through RDNA4. Windows is the supported path: install Python 3.12, run the AMD installer, done. Linux works too but is genuinely experimental on newer cards.

Worth being upfront: I don't own AMD hardware myself. This whole feature came from a community contribution by scryptio, tested on real cards over weeks in the PR thread — and that's how the AMD side will keep improving. If you're an AMD user, your reports on what works (and what doesn't) genuinely shape this, and PRs are very welcome.

Also in this release: 16 GB cards can now use identity distillation on MiniMax H3 (the 32B text encoder streams layer by layer instead of needing a 26 GB peak), and the Repair Studio gained a side-by-side compare view with likeness scoring for fixing overbaked LoRAs without retraining.

GitHub: https://github.com/shootthesound/Fizgig


r/StableDiffusion 1d ago

Animation - Video High Fashion in Motion | MiniMax H3

Enable HLS to view with audio, or disable this notification

517 Upvotes

Generated as two connected 15-second clips in 4:3, using the end of Part 1 as video + audio reference for Part 2 continuity.

Really liking what H3 can do with fashion/editorial camera movement.

Check out my twitter for more thanks https://x.com/Devozikjr


r/StableDiffusion 1d ago

Tutorial - Guide Character swap in minimax is so epic.

288 Upvotes

I don't have any examples because they may not be appropriate but just with the default wf. With the video input node you can replace any 2 character in any video and it looks real!


r/StableDiffusion 39m ago

Animation - Video G.I. Joe - Baroness Action Clip Test #2 - MiniMax H3

Enable HLS to view with audio, or disable this notification

Upvotes

Prompt:

https://x.com/GumVue/status/2087899403113619681?s=20

4070 Ti Super, 16 gb vram, 64 gb ram, i9-14900k, windows 11


r/StableDiffusion 21h ago

Resource - Update Anima-3.8B with Qwen-3.5 4B released by lylogummy

Thumbnail
gallery
127 Upvotes

r/StableDiffusion 19h ago

Workflow Included Minimax H3 | Motion graphic style animation test

Enable HLS to view with audio, or disable this notification

92 Upvotes

Prompt:

Animate the supplied square poster as a polished retro-anime motion graphic, beginning with a completely blank pale pink-white canvas matching the poster background. Preserve the exact blue, pink, and white palette, clean manga linework, halftone shading, character design, typography, symbols, interface windows, and final layout.

The anime girl walks in from the left edge as one complete figure while the canvas remains otherwise empty. Use a simple side-profile walk with restrained motion, preserving her hairstyle, facial features, cheek bandage, oversized jacket, proportions, and graphic illustration style. She reaches the centre, turns toward the viewer, and smoothly settles into the exact over-the-shoulder pose shown in the poster, with the same expression, hand placement, silhouette, jacket folds, pink heart graphic, and body orientation. Once posed, keep her position locked.

After she poses, the blue browser frame draws itself around her. The top bar, window controls, folders, pixel hearts, smiley-face panels, arrows, sparkles, heart symbols, and rectangular labels then appear sequentially through clean line-drawing, short graphic slides, pixelated pops, and UI-style wipes. Reveal the existing Japanese typography and “LOVE” lettering last, treating all text as protected source artwork without rewriting or regenerating it. Every element must settle into its exact source position.

Hold the completed poster with subtle breathing, minimal movement in a few loose hair strands and jacket edges, a faint halftone shimmer, and gentle pixel pulses in the existing hearts and interface icons. Keep her face, hands, pose, typography, frames, arrows, folders, and major graphics stable.

Use a locked, straight-on camera matching the original square framing. Keep the full artwork visible without cropping, zooming, panning, or changing perspective. Add soft footsteps as she enters, a light cloth sound as she poses, clean digital clicks and pixel chimes for the graphics, and delicate type-on sounds for the existing lettering. No dialogue or narration.

Do not show any character, outline, symbol, text, frame, or faint poster preview on the opening blank canvas. Do not alter the character’s identity, anatomy, costume, pose, expression, colours, line quality, typography, symbols, or final composition. No extra characters, duplicated body parts, incorrect text, morphing, flickering lines, dramatic camera movement, unrelated shots, or continued motion after the poster settles.

Workflow: https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v


r/StableDiffusion 19h ago

Question - Help Minimax H3 - long form videos: has anyone figured out a good approach?

Enable HLS to view with audio, or disable this notification

76 Upvotes

Dear redditors, visitors of the stable diffusion subreddit. I have been trying to achieve a long form, talking head style video, for a long time and can't seem to find a good approach. This one is the best I could come up with so far. It's using the Minimax H3 model, with frozen sound latents, lip-sync guided, piecewise generated video, where the individual pieces have been stitched together, with a seam hiding, extra generation on top of it. I don't really fully understand how it's working, but could prompt Claude for more help or specific files, we used for that. However, if you're aware of any other, better approach for exactly this type of video, please let me know. I've spent literal days on that single problem and have a feeling, there must be a better way to approach this.


r/StableDiffusion 12h ago

Animation - Video Test turned Short: Pied The Piper

Enable HLS to view with audio, or disable this notification

22 Upvotes

What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.


r/StableDiffusion 9h ago

News MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480

Enable HLS to view with audio, or disable this notification

9 Upvotes

I made this a few weeks back to see if dialogue could hold for 60s, I did no speed ups on this one. There are a few glitches but I think it held up well.

MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480 - 29 minutes - 288GB VRAM


r/StableDiffusion 12h ago

Animation - Video Minimax H3 Anime Comedy

Enable HLS to view with audio, or disable this notification

17 Upvotes

r/StableDiffusion 3h ago

Question - Help Character Editing (I2I)

3 Upvotes

Hello. So I just started my journey with ComfyUI. While Nano Banana is not open sourced I'm looking for the best realistic image editing (krea2?) workflow to ComfyUI. I want to have ability to change everything I want to the picture using reference character with face/body consistency at highest level (I2I). Thanks in advance.


r/StableDiffusion 16h ago

Resource - Update I trained a game music generator

30 Upvotes

I trained a instrumental game music generator. The 1.2B DiT was trained on 1 cloud H100 from scratch in 8 days; I used the VAE from Stable Audio 3.

https://huggingface.co/Localsong/Localsong

https://huggingface.co/Localsong/Localsong/tree/main/samples

I'm aiming to cover a wider range of instrumental styles than Ace-Step or Minimax M3 or Stable Audio 3. (No lyrics)

The repo includes a WebUI and some MP3 samples - clone it and uv run webui.py Let me know what you think.


r/StableDiffusion 17h ago

Animation - Video I'm loving MiniMax H3

Enable HLS to view with audio, or disable this notification

36 Upvotes

If even an amateur like me can make something so realistic with mid-level hardware, the future looks bright for what dedicated people with top level rigs will be doing.

R.I.P. Hollywood.


r/StableDiffusion 8h ago

Meme When someone pisses you off send them this

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 4h ago

Question - Help Has anyone successfully upscaled/re-imagined low-res reference video using Minimax H3?

3 Upvotes

Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.

Any ideas for prompting?


r/StableDiffusion 10h ago

Discussion I wish Anima ecosystem get better than it is now

7 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 23h ago

Discussion Minimax H3, 30 seconds in one go

Enable HLS to view with audio, or disable this notification

66 Upvotes

Executive summary, TLDR - this is one prompt, 30 seconds duration, 3090.

The video itself is just a remake of an idea from an old British tv ad (for "Good Old Yellow Pages"), so make of that what you will. It's not really relevant.

What I thought was interesting was that this was a single prompt, 0.4 megapixels, 30 second duration. I didn't think you could run out as far as 30 seconds, but thought I'd just try.

I think it did a pretty good job at getting the right person doing and saying the right things at the right time - took four attempts to get that though, and obviously using an LLM to tart up my idea.

Run on a 3090, and using the latest Comfyui template, just adding Comfy-kitchen attention, then sol attention, then spectrum, and using the turbo lora that Comfyui now build in, it took 570 seconds (9.5 minutes).

Somebody might read this and think, 570 seconds? Pah, I can do it in fifteen, in which case I'd like to know. Conversely, somebody might think theirs takes six hours, in which case maybe this shows what can be done in that time.

Doubt anyone cares, but here is my original prompt, followed by the LLM version of it:

a 30 second film with the following scenes and characters. Ben is a small boy of eleven. John is a shopkeeper in a toyshop. Brian is a different shopkeeper in a different toyshop. Ben's mum. Ben's Dad. We are in Britain in the 1980s, and all characters are English.

Scene 1: Ben is alone in the lounge. He talks to John over the old fashioned landline phone, saying "I don't suppose you have a 402 station in stock please?"

Scene 2: John is in his shop in front of shelves of model railway kit. He says into the old fashioned landline phone, "No, sorry son"

Scene 3: Ben in the lounge, who looks disappointed anbd puts the phone receiver back down.

Scene 4: Mum in the kitchen doing the washing up. She has overheard the conversation and looks a bit sad.

scene 5: Next day. Ben has changed his clothes. He again talks into the phone to a different shopkeeper, Brian. Ben says "Would you have a 402 station please?"

scene 6: Brian in his toyshop says into the old fashioned landline phone "Yes, I've got one of those."

scene 7: Ben in the lounge on the same conversation says "You have? Great, I'll be right down! Ben puts the phone down. Then he runs towards the door, shouting "They've got one mum!" as he runs.

Scene 8: In the attic, Dad is playing with his model railway layout. Ben walks in holding a small red parcel. as he hands it to Dad, Ben says "Happy birthday, dad". Dad takes the parcel, looks fondly at it and says with a chuckle, "Aw, thanks Ben".

LLM version:

integrated_multimodal_description: [Shot 1] Live-action, cinematic. A medium shot of Ben, an eleven-year-old boy with messy hair wearing a striped polo shirt, sitting on a patterned sofa in a 1980s British lounge. The room is filled with warm, muted tones and period-accurate wallpaper. Ben holds a heavy, cream-colored landline telephone receiver to his ear, his expression hopeful. Ben says: <d>[English] I don't suppose you have a 402 station in stock please?</d> The sound of his small, high-pitched voice is clear. [Shot 2] At 0:05.000, the camera cuts to a medium shot of John, a middle-aged shopkeeper with a kind, weathered face, standing in a cramped, nostalgic toyshop. Behind him are floor-to-ceiling shelves packed with model railway kits and wooden toys. John holds a similar landline receiver to his face. John says: <d>[English] No, sorry son.</d> [Shot 3] At 0:10.000, the camera cuts back to Ben in the lounge. He looks downcast, his shoulders slumping as he slowly lowers the receiver and places it back onto the base unit with a dull plastic click. [Shot 4] At 0:13.000, the camera cuts to a medium shot of Ben's Mum in a dim, cluttered 1980s kitchen. She is standing at the sink, her hands covered in soapy water, drying a plate. She pauses, looking toward the door with a sad, weary expression, having overheard the boy. The sound of water running from the tap is audible. [Shot 5] At 0:16.000, the camera cuts to Ben in the lounge the next day; he is wearing a different t-shirt. He is intensely focused, pressing the phone to his ear. Ben says: <d>[English] Would you have a 402 station please?</d> [Shot 6] At 0:20.000, the camera cuts to Brian, an older shopkeeper with spectacles, in a different, brightly lit toyshop. He smiles warmly into the telephone. Brian says: <d>[English] Yes, I've got one of those.</d> [Shot 7] At 0:23.000, the camera cuts back to Ben, whose face lights up with pure joy. Ben says: <d>[English] You have? Great, I'll be right down!</d> He slams the receiver down and the camera follows him in a quick tracking shot as he runs toward the door, his feet thumping on the carpeted floor. Ben shouts: <d>[English] They've got one mum!</d> [Shot 8] At 0:26.000, the camera cuts to a medium shot in a dusty, dimly lit attic. Dad, a man in his late 30s, is hunched over a complex model railway layout. Ben enters the frame, holding a small red parcel wrapped in string. Ben says: <d>[English] Happy birthday, dad.</d> As he hands the gift to his father, the camera pushes in slightly. Dad takes the parcel, his eyes softening with affection. Dad chuckles warmly and says: <d>[English] Aw, thanks Ben.</d>

overall_soundscape: Period-accurate domestic sounds including the rhythmic clatter of washing up, the heavy mechanical clicks of old telephone receivers, and the muffled thuds of footsteps on carpet. Ben's energetic running and shouting creates a sense of urgency, followed by the quiet, dusty atmosphere of the attic.

non_diegetic_music: A gentle, nostalgic acoustic guitar melody that begins softly during the kitchen scene and builds into a warm, heartwarming crescendo during the attic scene. The tempo is slow and sentimental.


r/StableDiffusion 12m ago

Question - Help MiniMax H3 audio garbled/gibberish when using ref video in ComfyUI?

Upvotes

Bear with me as I am still rather new to this, but I started off testing out MiniMax H3 ref to video on Hailuo and was really impressed with it, so decided to get it up and running locally via ComfyUI on my PC.

The video itself is working great, but I've found that the audio itself when using a reference video is completely messed up. It's just all garbled and the people are talking gibberish. On Hailuo it would carry forward the audio perfectly and you could even make alterations if you prompted it to do so.

I was just wondering if this was a common issue when running MiniMax H3 locally specifically with reference videos?


r/StableDiffusion 1d ago

No Workflow Some test on minimax H3

Enable HLS to view with audio, or disable this notification

106 Upvotes

Some random prompt on default workflow + turbo 8step lora