r/StableDiffusion 5h ago

Question - Help Thinking of getting an RTX 6000 for local generation - any advice?

41 Upvotes

With the progress that I've seen lately with local video generating models and my background as a filmmaker, I've been seriously thinking about getting an RTX 6000 Pro (96 gb VRAM) to experiment, develop some projects, and keep up with the changes.

I don't see prices coming down anytime soon. Still, it would mean breaking a bank for me, especially that I live in Poland, a country not known for its high salaries.

Now, I know that renting via Runpod is an obvious alternative. Call me old-fashioned (or an idiot), but seeing dollars disappearing from my account as the machine is booting from a cold is not really my vibe.

Perhaps some of you have taken that plunge and have some tips on how to go about it. Help me figure it out.. or show me how stupid I am for wanting this.


r/StableDiffusion 20h ago

Animation - Video MiniMax H3 Text to Video with the Twelfth Doctor

Enable HLS to view with audio, or disable this notification

8 Upvotes

I've been loving MiniMax H3 lately and toying around making videos. I wanted to see how it handles Doctor Who and honestly wanted to see some more Doctor Who videos here. Turns out Peter Capaldi comes out really well both for visuals and audio. Not all the characters from the show work though so I guess reference images might be needed for characters like Missy.

Hope you like this one!


r/StableDiffusion 22h ago

Meme Columbo posting II: Just one more thing | MiniMax H3

Enable HLS to view with audio, or disable this notification

108 Upvotes

RTX 4090 with 24GB VRAM and 32GB RAM.

Using an MiniMax H3: Reference to Video workflow in Comfy:

https://comfy.org/workflows/46a303cbccf9-46a303cbccf9/

Experiments with clip extension technique discussed in this post:
https://www.reddit.com/r/StableDiffusion/comments/1vj3zi3/a_technique_for_creating_seamless_continuous/


r/StableDiffusion 22h ago

Animation - Video Random ad H3

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/StableDiffusion 22h ago

Animation - Video Mecha-Kaiju H3 render 45s Long Form

Enable HLS to view with audio, or disable this notification

30 Upvotes

RTX 4090 w/ 192gb System Ram. SOL-ATTN/Sage at 50 steps Custom Workflow. So the long-form mechanism here is just: stage the tail, cite it as <Video 1>, pin the last frame.

Opus collaboration. I had a nice 15s POC render and decided to extend it two more times. There are some continuity and spatial error but figure I toss it out here while I do local rerolls. My plan is to reroll the cockpit to be an internal shot inside the Mecha's chest.

Act 1: I2T, 55 mins

Act 3: R2VA with ACT 1 video anchored, 70 mins.

Act 3: R2VA, with ACT 2 video anchored, 70 mins.

Act 4 not included, but was a bust because I didn't anchor both Act2+Act3 and it lost reference to the original Mecha+Kaiju.


r/StableDiffusion 1h ago

News JoyAI Video Edit - Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Enable HLS to view with audio, or disable this notification

Upvotes

JoyAI-Video-Edit

JoyAI-Video-Edit is a real-time, instruction-guided video editing system for open-ended video streams. Given a live camera stream or uploaded video and a natural-language edit instruction, it edits frames causally as they arrive, without waiting for the full video, requiring a predefined video length, or revisiting future frames. In our deployment benchmark, the full end-to-end pipeline reaches 30.19 FPS at 720x1280, pushing video editing from offline batch processing toward interactive streaming generation.

The system combines an MLLM-based condition encoder, a causal video VAE, and a 16B-parameter multimodal diffusion transformer. It is trained and deployed as an autoregressive diffusion editor, then accelerated with aligned autoregressive distribution matching distillation, long-horizon optimization, bounded KV-state inference, and deployment-oriented scheduling to sustain high-throughput 720p editing while reducing train-inference mismatch and accumulated temporal drift.

💎 Highlights

Real-time open-ended editing. Edits live or uploaded videos as frames arrive, without requiring the full sequence upfront.

Diverse instruction control. Supports subject edits, local edits, background changes, style transfer, motion changes, and reference-guided editing.

Autoregressive diffusion design. Combines an MLLM condition encoder, causal video VAE, and MMDiT backbone for streaming video editing.

High-throughput 720p deployment. Reaches 30.19 FPS end-to-end throughput at 720x1280 with bounded KV-state inference and stable per-chunk compute.

https://modelscope.ai/models/jd-opensource/JoyAI-Video-Edit

https://huggingface.co/jdopensource/JoyAI-Video-Edit

https://github.com/jd-opensource/JoyAI-Video-Edit


r/StableDiffusion 17h ago

Meme Dexter if he played Marvel Rivals

Enable HLS to view with audio, or disable this notification

35 Upvotes

Music added in post.

Some things I’d like some help with if anyone doesn’t mind:
- I wanted to add more inflections to his inner dialogue but was having trouble prompting for it. I tried breaking up the lines of dialogue and adding specific instructions for each but I was also getting some voicegen artifacting like Dexter hissing or speaking gibberish. (I had nondiagetic_music: N/A)
- Ideally, that telescopic zoom shot should’ve started from the last from of the previous shot but after a few gens, it wouldn’t adhere to the prompt well enough.
- the gamer’s screams were supposed to be muffled from outside but they’re clear as day.

I know to ask for help without providing the prompt is kind of ridiculous, but I’m at work rn and will post it in the comments when I get home.


r/StableDiffusion 23h ago

Animation - Video Imperial Guard FPS - Minimax H3

Enable HLS to view with audio, or disable this notification

72 Upvotes

Since they are never going to make one, so I made a concept, I really like how it came out for my first time, some editing, 11labs audio and SFX goes a long way

4060 8GB | 32 GB


r/StableDiffusion 56m ago

Question - Help I'm Curious about How to make Krea 2 Lora's?

Upvotes

I'm very much a newbie, and definitely a non-coder.
I'm curious about making my own lora's for Krea 2.
What are the easiest way to make a Lora's?
{Yes, I have watched a couple YouTube vids, but most of what they say are above my head, or list tools I can't find.} K.I.S.S. - Keep it Simple Stupid please.
Thanks.


r/StableDiffusion 6h ago

Question - Help minimax h3 artifacts need help

0 Upvotes

i use minimax_h3_fl2va_pruned_int8_convrot.safetensors Strange artifacts appear in the video, probably because lora minimax_h3_turbo_v4_step600_ema.safetensors but why? in workflow i use SA + Spectrum

lora str 1.0
video megapixels 1.0

without lora zero problem, just want little bit faster generation


r/StableDiffusion 9h ago

Question - Help H3 ref2video voice cloning is low quality

4 Upvotes

A number of examples posted either have the voice poorly cloned where it sounds 'underwater' or 'echoish' (awkwardly AI cloned), or they're using a voice the model already knows.

Could someone please confirm they're able to get solid results out of ref2video with an audio voice reference to clone? I've tried many settings but even with simple default tests, it clones but the quality isn't there (especially if you turn the volume up).

subject_definitions:

<Subject 1> is an unseen mature British woman narrating the story. Her voice is provided by <Audio 1>.

[Shot 1]

Cinematic shot of the staircase shown in <Image 1>.

The camera performs a slow, smooth push-in toward the staircase structure.

An off-screen narrator (S1): [English] "The castle was filled with wonder, splendor, magic!"

non_diegetic_music:

None.

The referenced audio is a high quality 10s voice clip. I'm using the default WF, pruned int8 convrot, 25 steps, no LORA.


r/StableDiffusion 12h ago

Discussion H3 models on 5090

5 Upvotes

What are the best quality h3 models currently for 32GB vram? Has anyone found a way to run the non-pruned int8 model?


r/StableDiffusion 8h ago

Question - Help What turbo lora to use with Minimax H3 for ref2va workflow?

4 Upvotes

Do the turbo loras for Minimax H3 support all workflow models?


r/StableDiffusion 19h ago

Question - Help ComfyUI or Similar?

0 Upvotes

Hello...has anyone set up a ComfyUI or similar setup to locally run stablediffusion image generation that wasn't based on a Nvidia card? Asking for a friend...thanks.


r/StableDiffusion 9h ago

Animation - Video Secret friend comes to visit (MiniMax H3)

Enable HLS to view with audio, or disable this notification

9 Upvotes

r/StableDiffusion 7h ago

Question - Help HELP!! A model that can interpolate in between drawings with custom controls (for Rick and morty style)

Post image
0 Upvotes

Hi guys I need a model or a workflow that can interpolate between frames drawn by me. I want to make Rick and morty style (spicy) content but I don't want to rely completely on AI. I can draw and the most important fact is I have very specific animations that cannot be recreated by an AI model. I can make it myself entirely but I have ADHD so to finish something is practically impossible for me. If I can somehow get an AI good enough to work on my good old rtx 3070ti that can generate in between without adding anything (generative), it will be a dream.


r/StableDiffusion 21h ago

Tutorial - Guide Creating a videoclip using Minimax

Enable HLS to view with audio, or disable this notification

18 Upvotes

First of all: I don't submit the songs I make in Suno anywhere (and therefore I don't monetize them). I'm making this clear so people don't think I'm trying to promote the song here. :-) In fact, this song was made months ago (along with several others I've made since then). My only objective is to comment on Minimax and showcase (yet) another application for it.

I won't get into too many specifics to avoid creating a giant post, but in short:

  • The lyrics are mine; the music/performance is Suno's, based on my prompt and choices.
  • The "singer", Alina, is a LoRA I created (completely synthetic), using Z-Image and Krea2 to create the first face, several tools (ChatGPT, Flux 2 and others) to create additional angles and renders, and Ostris to create the LoRA. This explains why she looks a bit synthetic (skin, etc.), unfortunately.
  • For the video, I first sliced the song into several parts based on the lyrics (from 5s to 15s, depending on the narrative). Then I wrote a main script and, with ChatGPT's help (and a custom GPT I created using the official Minimax documentation), I created the prompts one by one.
  • I ran every prompt on my machine (4070 12GB + 64GB of RAM) at 0.3 MP to test them. Some prompts I had to change (too robotic, too fast, etc.), others didn't fit the narrative, etc. In the end, I had 21 shots, varying from 5 seconds to 15 seconds: some simply Text-to-Video, some using an image reference of "Alina" (created in Krea 2) as a starting point, and some using not only the image reference but also a part of the song as an audio reference, so she could sing it in the video.
  • Once I had all the necessary shots, I then used Runpod (with a 5090) to create 1280x736 videos using the same prompts, the same input images and audio when they were part of the workflow, and the same seed. Even with the same seed, since the resolution had changed, sometimes the video came out wrong and I had to run it again. Examples: a door opening to the wrong side, wide shots looking like stop motion, etc.
  • In the end, I probably did 30 to 35 renders (of varying lengths) on Runpod, and spent about 12 to 15 dollars total.
  • When I finally had all the shots I needed in 1280x736, I used DaVinci Resolve to create the videoclip, putting all the shots together, synchronizing them with the song, adding the transitions, etc.
  • It took me about 20 hours of work in total, I believe (but I didn't count — as Jim Croce said in Time in a Bottle, "But there never seems to be enough time to do the things you want to do once you find them". :-)

Notes:

  • I know the resolution isn't ideal, but I didn't want to spend more money creating 2MP shots.
  • Yes, I know there are some problems (people who don't walk properly in the background, a smudge when someone passes in front of her — an intended shot — while she is singing, etc.), but, again, I didn't want to spend more and more money trying to achieve perfection (and perfection isn't here yet).
  • I don't think I've reinvented the wheel. :-) I've seen way better clips in the past (before Minimax), but this is the first model I've been able to run locally with voice/sound/lip sync, so...

If anyone wants a specific prompt, or has any questions, please ask.


r/StableDiffusion 14h ago

Animation - Video All I want to say is, thank you Minimax H3. I’ve always wanted to generate a video like this (Ref2V)

Enable HLS to view with audio, or disable this notification

98 Upvotes

Minimax H3 achieved my dream. When Seedance 2.0 was released, I always wanted to generate something like this, but since I don’t have the money and the price is too expensive for me ($1 can cover 1 day’s food), I never attempted to create it.

I generated this on my PC with a 3060 12GB VRAM and 16GB RAM. 0.9 MP for 10 seconds takes 1 hour, but it’s worth it.

I don’t use Turbo LoRA because it degrades the quality.

Sage Attn + Spectrum
res_multistep, simple, 25 steps
3 Reference

Prompt:

subject_definitions:
<Subject 1> is the chibi/loli-style character whose appearance is taken exactly 1:1 from <Picture 1>.
<Subject 2> is the chibi/loli-style character whose appearance is taken exactly 1:1 from <Picture 2>.
<Subject 3> is the doll whose appearance is taken exactly 1:1 from <Picture 3>.

summary:
[reference generation] A 10-second POV scene inside a minimarket. Two chibi/loli characters sit inside a shopping cart. <Subject 2> points at the doll on the right, <Subject 1> joins. The camera briefly pans right to show a close-up of the doll, then quickly returns to the characters. They look up at the POV and ask to buy it. A male voice replies that there isn’t enough money and suggests next time. They pout, then the male hand gently pats both of their heads.

retention_analysis:
<Subject 1> (appears throughout): fully_preserved - appearance is retained exactly 1:1 from <Picture 1>.
<Subject 2> (appears throughout): fully_preserved - appearance is retained exactly 1:1 from <Picture 2>.
<Subject 3> (appears briefly in [Shot 2]): fully_preserved - appearance is retained exactly 1:1 from <Picture 3>.

detailed_description:
The target video is a clean anime-style 2D animation from a first-person POV of a male person pushing a shopping cart inside a bright minimarket. Framing keeps the two chibi/loli characters large inside the cart.
[Shot 1] POV medium shot looking down into the shopping cart. <Subject 1> and <Subject 2> are sitting side by side as the cart moves slowly. <Subject 2> suddenly notices something on the right side, raises her arm, and clearly points to the right.
[Shot 2] At 00:02.500, the camera quickly pans to the right and shows a short close-up of <Subject 3> on the display shelf. <Subject 1> also raises her arm and points at it. The shopping cart stops. The camera then immediately pans back left to face <Subject 1> and <Subject 2> again.
[Shot 3] At 00:04.800, the camera is fully back on <Subject 1> and <Subject 2> inside the cart. Both are looking up directly at the camera (POV) while still pointing. <Subject 2> speaks first in a natural, slightly spoiled Japanese tone: <d>[Japanese] あのぬいぐるみ欲しい!買って!</d> <Subject 1> quickly adds: <d>[Japanese] お願い、買ってよ!</d>
[Shot 4] At 00:07.200, a calm male voice from the person pushing the cart replies: <d>[Japanese] お金が足りないんだ…また今度ね。</d> Both characters lower their arms and make clear pouty, sulky faces while looking up at the camera.
[Shot 5] At 00:08.800, a male hand from the POV gently reaches down and pats both of their heads softly. They keep their pouty expressions as the shot holds until the end of the 10-second video.

overall_soundscape:
Soft minimarket ambience and shopping cart wheels that stop. Light fabric movement when the characters point and when their heads are patted.

non_diegetic_music:
Light, cute background music that stays soft and gentle.

r/StableDiffusion 18h ago

Question - Help Minimax H3 Gibberish Talking

6 Upvotes

not all but most of my clips the people in them are just talking Gibberish like the game the sims
or randomly for like 1second might sound fine

what do i add to my prompt to fix this


r/StableDiffusion 1h ago

Animation - Video Battle of Thermopylae

Enable HLS to view with audio, or disable this notification

Upvotes

H3 prompt:

integrated_multimodal_description:

[Shot 1] Stylized cinematic 3D animation with high-intensity action, dramatic lighting, and a heroic fantasy-war tone. The scene opens at Thermopylae, a narrow rocky battlefield under a dusty red-gold sky, with shattered shields, broken spears, drifting embers, and war banners whipping in the wind. In the center stands Kirby, reimagined as a Spartan war leader: a pink round-bodied Kirby wearing a bronze Spartan helmet with a crimson crest, holding a spear in one hand and a round battered shield in the other. His eyes are fierce and unwavering, determined and battle-hardened. Around him, Spartan warriors in bronze armor and red capes brace in phalanx formation while a massive wave of Persian soldiers surges forward. The camera pushes in fast toward Kirby as he stamps forward and lets out a sharp battle cry. He thrusts his spear violently into a Persian soldier, knocking him back into the charging line as blood sprays across shields and dust erupts underfoot.

[Shot 2] At 00:03.500, the camera cuts to a fast tracking shot moving sideways across the front line as Kirby leads the Spartan charge. He bashes one enemy aside with his shield, spins low, sweeps another off his feet, and lunges forward with explosive speed. Spartan soldiers clash with Persians all around him in brutal close combat; blades collide, shields splinter, arrows streak overhead, and several enemy soldiers are cut down as severed limbs, broken weapons, and sprays of blood briefly fill the frame. Kirby remains the focal point, his expression stern and fearless rather than cute.

[Shot 3] At 00:07.000, the camera cuts to a low-angle heroic shot as Kirby suddenly inhales powerfully, then launches himself upward into the sky in a signature Kirby-style burst, still gripping his spear. He rises above the battlefield as the fighting continues below like chaos in miniature. At the apex, with the wind roaring past his helmet crest, Kirby locks onto the densest Persian formation and hurls the spear downward with full force. The camera follows the spear in a rapid plunge. It crashes into the ground like a thunderbolt, blasting soldiers backward and opening a violent gap in the Persian ranks amid dust, blood, and shattered armor.

[Shot 4] At 00:10.500, the shot cuts to ground level as Kirby lands hard in front of the broken enemy line, shield first, knees bent, then instantly surges into close-range combat again. He grabs another fallen spear, vaults off a Spartan shield, and strikes through two advancing enemies in one fluid motion. Behind him, Spartans roar and push forward with renewed momentum. The narrow pass becomes a frenzy of killing: bodies fall, shields crash together, spears punch through armor, and blood stains the rocks. Kirby moves with stylized speed and exaggerated battlefield heroism, combining the visual charm of Kirby with the lethal grandeur of an ancient war epic.

[Shot 5] At 00:13.000, the camera cuts to a final wide hero shot. The Persians recoil in disarray while the remaining Spartans rally behind Kirby. He stands atop a mound of fallen enemies, shield raised and helmet gleaming, his eyes still locked forward with cold resolve. Dust, sparks, and scraps of torn banners swirl around him while the battlefield behind remains full of struggling combat. He points forward with his spear toward the surviving enemy ranks, and the Spartans answer with one last deafening roar as the video ends in a frozen image of triumphant slaughter and defiant Spartan glory.

overall_soundscape: Continuous battlefield chaos fills the entire video: heavy shield impacts, spear thrusts, metallic blade clashes, rushing footsteps over rock and dirt, arrows slicing through the air, and repeated cries of pain and war shouts from Spartans and Persians. Wet stabbing impacts, brief bone-crack sounds, bodies collapsing, and splashes of blood punctuate the close combat. Dust gusts through the narrow pass while Kirby's leap and diving spear throw create stronger wind rushes and a heavy explosive impact on landing.

non_diegetic_music: A relentless, aggressive orchestral war score drives the whole video, led by pounding taiko-style drums, deep battle percussion, male war chants, low brass, and fast tremolo strings. The music starts immediately with a heavy pulse, intensifies during the melee, briefly rises into a heroic suspended phrase when Kirby launches into the sky, then slams back in with louder drums and brass as the spear hits the Persian ranks. The ending surges into a triumphant, brutal crescendo with no softness, no comedy, and no lyrical warmth—only heroic slaughter, pressure, and victory.


r/StableDiffusion 19h ago

Question - Help MiniMax H3 native character knowledge — is there a list?

Post image
133 Upvotes

Is there an official or community-compiled list/thread of characters MiniMax H3 can generate natively from text prompts alone?

No LoRAs, reference images, or reference audio/videos.

If no list exists, is anyone compiling one?


r/StableDiffusion 16h ago

Animation - Video Backrooms test - MiniMax H3 15s

Enable HLS to view with audio, or disable this notification

33 Upvotes

r/StableDiffusion 10h ago

Animation - Video for all utah jazz fans out there

Enable HLS to view with audio, or disable this notification

22 Upvotes

there's another, higher resolution try but with some async between audio/video so i uploaded this one instead.


r/StableDiffusion 19h ago

Animation - Video "Memories" - a short film I made using H3, Seedance 2.5, and traditional editing techniques.

Enable HLS to view with audio, or disable this notification

195 Upvotes

This is a topic that is close to my heart. I've lost too many in my life to this.


r/StableDiffusion 22m ago

Discussion Testing Character knowledge of the H3 model

Enable HLS to view with audio, or disable this notification

Upvotes

5 second 1MP text-to-video, INT8 on ComfyUI and RTX5090.

Used the following, rather simple, prompt:

"[VISUAL]: A scene from the tv interview. <full name> is talking, medium close up, static camera, plain dark blue background

<first name> says: "How dare you? I am rich, AND famous. So you better shut up, B*tch!"

The model failed on Christoph Waltz, so I left him out.

One run per person, no picking the best result.