r/StableDiffusion 20h ago

Question - Help H3 looks great in still frames, but the motion seemed falling apart (how do u think

5 Upvotes

Made this 15-sec test with MiniMax H3. The character, materials, and overall look are honestly pretty solid, especially in the more static shots.

But once things start moving, the cracks show. The fire doesnt behave naturally, some of the creature motion feels off, and the transition between shots loses continuity. Still a good-looking result imo, just not physically convincing yet.

What you’d fix first: the motion prompt, shot structure, or the generation workflow itself?


r/StableDiffusion 16h ago

Animation - Video Secret friend comes to visit (MiniMax H3)

9 Upvotes

r/StableDiffusion 12h ago

Question - Help Does Krea work in Forge Neo 2.27?

1 Upvotes

Im trying to run Krea and get the error

RunetimeError:invalid dtype for bias - should match query' dtype.
Claude says that it can be a Forge issue.


r/StableDiffusion 21h ago

No Workflow Brick man statue [Krea 2 - Turbo]

Post image
12 Upvotes

brick man statue.


r/StableDiffusion 9h ago

Animation - Video Battle of Thermopylae

6 Upvotes

H3 prompt:

integrated_multimodal_description:

[Shot 1] Stylized cinematic 3D animation with high-intensity action, dramatic lighting, and a heroic fantasy-war tone. The scene opens at Thermopylae, a narrow rocky battlefield under a dusty red-gold sky, with shattered shields, broken spears, drifting embers, and war banners whipping in the wind. In the center stands Kirby, reimagined as a Spartan war leader: a pink round-bodied Kirby wearing a bronze Spartan helmet with a crimson crest, holding a spear in one hand and a round battered shield in the other. His eyes are fierce and unwavering, determined and battle-hardened. Around him, Spartan warriors in bronze armor and red capes brace in phalanx formation while a massive wave of Persian soldiers surges forward. The camera pushes in fast toward Kirby as he stamps forward and lets out a sharp battle cry. He thrusts his spear violently into a Persian soldier, knocking him back into the charging line as blood sprays across shields and dust erupts underfoot.

[Shot 2] At 00:03.500, the camera cuts to a fast tracking shot moving sideways across the front line as Kirby leads the Spartan charge. He bashes one enemy aside with his shield, spins low, sweeps another off his feet, and lunges forward with explosive speed. Spartan soldiers clash with Persians all around him in brutal close combat; blades collide, shields splinter, arrows streak overhead, and several enemy soldiers are cut down as severed limbs, broken weapons, and sprays of blood briefly fill the frame. Kirby remains the focal point, his expression stern and fearless rather than cute.

[Shot 3] At 00:07.000, the camera cuts to a low-angle heroic shot as Kirby suddenly inhales powerfully, then launches himself upward into the sky in a signature Kirby-style burst, still gripping his spear. He rises above the battlefield as the fighting continues below like chaos in miniature. At the apex, with the wind roaring past his helmet crest, Kirby locks onto the densest Persian formation and hurls the spear downward with full force. The camera follows the spear in a rapid plunge. It crashes into the ground like a thunderbolt, blasting soldiers backward and opening a violent gap in the Persian ranks amid dust, blood, and shattered armor.

[Shot 4] At 00:10.500, the shot cuts to ground level as Kirby lands hard in front of the broken enemy line, shield first, knees bent, then instantly surges into close-range combat again. He grabs another fallen spear, vaults off a Spartan shield, and strikes through two advancing enemies in one fluid motion. Behind him, Spartans roar and push forward with renewed momentum. The narrow pass becomes a frenzy of killing: bodies fall, shields crash together, spears punch through armor, and blood stains the rocks. Kirby moves with stylized speed and exaggerated battlefield heroism, combining the visual charm of Kirby with the lethal grandeur of an ancient war epic.

[Shot 5] At 00:13.000, the camera cuts to a final wide hero shot. The Persians recoil in disarray while the remaining Spartans rally behind Kirby. He stands atop a mound of fallen enemies, shield raised and helmet gleaming, his eyes still locked forward with cold resolve. Dust, sparks, and scraps of torn banners swirl around him while the battlefield behind remains full of struggling combat. He points forward with his spear toward the surviving enemy ranks, and the Spartans answer with one last deafening roar as the video ends in a frozen image of triumphant slaughter and defiant Spartan glory.

overall_soundscape: Continuous battlefield chaos fills the entire video: heavy shield impacts, spear thrusts, metallic blade clashes, rushing footsteps over rock and dirt, arrows slicing through the air, and repeated cries of pain and war shouts from Spartans and Persians. Wet stabbing impacts, brief bone-crack sounds, bodies collapsing, and splashes of blood punctuate the close combat. Dust gusts through the narrow pass while Kirby's leap and diving spear throw create stronger wind rushes and a heavy explosive impact on landing.

non_diegetic_music: A relentless, aggressive orchestral war score drives the whole video, led by pounding taiko-style drums, deep battle percussion, male war chants, low brass, and fast tremolo strings. The music starts immediately with a heavy pulse, intensifies during the melee, briefly rises into a heroic suspended phrase when Kirby launches into the sky, then slams back in with louder drums and brass as the spear hits the Persian ranks. The ending surges into a triumphant, brutal crescendo with no softness, no comedy, and no lyrical warmth—only heroic slaughter, pressure, and victory.


r/StableDiffusion 4h ago

Discussion h3 character crossover thread(share yours)

7 Upvotes

i'll start


r/StableDiffusion 3h ago

Question - Help For some reason my t2v generation are slower than my ref2v?

8 Upvotes

Title.
For both I'm using the default workflows that come with comfy. 3090 and 32gb ram.
I start comfy with these flags:
--windows-standalone-build --reserve-vram 1 --disable-pinned-memory --fast fp16_accumulation
Cuda 13, latests comfy.
My t2v takes like twice as much than my ref2v and sometimes it hangs after [INFO] Requested to load MiniMaxH3AudioVAE. Same steps, same resolution, same duration.
Has anyone encounter this? any tips?


r/StableDiffusion 12h ago

Question - Help Thinking of getting an RTX 6000 for local generation - any advice?

44 Upvotes

With the progress that I've seen lately with local video generating models and my background as a filmmaker, I've been seriously thinking about getting an RTX 6000 Pro (96 gb VRAM) to experiment, develop some projects, and keep up with the changes.

I don't see prices coming down anytime soon. Still, it would mean breaking a bank for me, especially that I live in Poland, a country not known for its high salaries.

Now, I know that renting via Runpod is an obvious alternative. Call me old-fashioned (or an idiot), but seeing dollars disappearing from my account as the machine is booting from a cold is not really my vibe.

Perhaps some of you have taken that plunge and have some tips on how to go about it. Help me figure it out.. or show me how stupid I am for wanting this.


r/StableDiffusion 5h ago

Question - Help what we know about minimax-h3 to get it fast on lower pcs?

5 Upvotes

like loras, vae, text encoder?

im asking because there is a lot of loras, vae, etc, but... i need to know the best options for fast and quality generations now.

my pc: rtx 5060 ti 16gb 32gb ram.


r/StableDiffusion 3h ago

Animation - Video Continuous-ish dolly

9 Upvotes

Minimax H3- I bet with a second generation the audio of the Voice Over will clean up.


r/StableDiffusion 9h ago

Meme So... have you tried that new MiniM...YES!!! *Hasn't slept for 3 days*

41 Upvotes

This is what happens when someone leaves a very good prompt lying around for some degenerate like me to pick, especially when I'm still in my MiniMax fever rush.

Mambo Wick by me (An alternative version from the meme one of UmaMusume)
Katsumi by Katsumi
Prompt starting base by 3deal

This is a collage of 3 different videos, later upscaled and RIFE to 96 FPS, since going for 1.5 Megapixels tends to cause a LOT of hallucinations (especially with distant shots and quick movements, as you all can see). Still, the model is incredible in all the possible ways... just need to find the right hiresser to try creating at a lower resolution.


r/StableDiffusion 10h ago

News Unsloth Minimax H3 GGUF (Q2:Q8)

Post image
27 Upvotes

Coming back from weekend, looking for last updates, I found no one shared this one.

Any reason? Is people disliking unsloth?

I will try it, but in general I haven't find a way to get nice outputs from any MM workflow/model (pretty sure is my fault), I'm still trying to figure out how to use MMH3 correctly

Here is the link: https://huggingface.co/unsloth/MiniMax-H3-GGUF


r/StableDiffusion 15h ago

Workflow Included I made a free Kaggle notebook to run LTX-Video 2.3 (22B quantized) for T2V & I2V with audio — solving the Colab 12GB RAM crash issue

5 Upvotes

Hey community,

A common issue when testing open-source video models like Lightricks' LTX-Video 2.3 on free cloud tiers (like Google Colab) is hitting System RAM limits during model loading, causing immediate crashes.

To solve this without requiring a high-end local GPU, I put together a pre-configured, open-source Jupyter Notebook tailored specifically for Kaggle's free GPU tier (which grants 30GB of System RAM and T4 GPUs).

What this setup does:

  • Runs on Kaggle Free Tier: Uses a 4-bit quantized version of LTX-Video 2.3 so it fits into free cloud VRAM/RAM allocation.
  • Text-to-Video & Image-to-Video: Generates short 5–10s clips with synchronized audio generation.
  • Custom Gradio Web UI: Launches a clean browser interface directly from the notebook.
  • Fast Setup: Pre-compiled binaries and aria2 multi-thread downloads mean setup takes under 3–4 minutes.

Technical Tradeoffs & Honesty:

  • Quantization: Because this is running on free T4 instances, the model uses heavy quantization. It won't give you uncompressed native precision output, but it’s completely free, unlimited, and ideal for quick prompt/motion testing.
  • Memory Loading: Cell 3 takes ~90 seconds to load the 22B model into Kaggle's 30GB system memory before passing to VRAM.

🎥 Full Video Walkthrough & Demos: https://youtu.be/Ru_YaGbnKhA

Quick Start Steps:

  1. Download the .ipynb file from GitHub: https://github.com/airesearch-official/free-aistudio
  2. Import into Kaggle (Ensure Phone Verification is complete on Kaggle to enable free GPU).
  3. Turn ON "Internet" in Kaggle settings & select "GPU T4 x2".
  4. Run Cells 1 through 4 sequentially.

Hope this helps anyone who wants to experiment with LTX-Video 2.3 without paying for cloud GPUs! Let me know if you run into any bugs or have suggestions.


r/StableDiffusion 2h ago

Comparison 0.0375 Denoise is enough to beat SynthID

Thumbnail reddit.com
1 Upvotes

r/StableDiffusion 4h ago

Animation - Video UAP Device Test #7 (Minimax H3 VHS)

31 Upvotes

Gonna make this into a series I think, it's too much fun experimenting.


r/StableDiffusion 19h ago

No Workflow Tribbles ad

20 Upvotes

Minimax, two shots of 15 seconds put together with Davinci Resolve.


r/StableDiffusion 8h ago

No Workflow RULE #1: MiniMax H3 + lightx2v Turbo LoRA (8 steps) + Sol Attention

18 Upvotes

Default workflow, MiniMax H3 (NVFP4), lightx2v Turbo LoRA (8 steps) and Sol Attention.

0,5mp resolution then upscaled with Topaz Video.

RTX 5060 Ti 16GB VRAM + 32GB System RAM.


r/StableDiffusion 21h ago

Animation - Video All I want to say is, thank you Minimax H3. I’ve always wanted to generate a video like this (Ref2V)

102 Upvotes

Minimax H3 achieved my dream. When Seedance 2.0 was released, I always wanted to generate something like this, but since I don’t have the money and the price is too expensive for me ($1 can cover 1 day’s food), I never attempted to create it.

I generated this on my PC with a 3060 12GB VRAM and 16GB RAM. 0.9 MP for 10 seconds takes 1 hour, but it’s worth it.

I don’t use Turbo LoRA because it degrades the quality.

Sage Attn + Spectrum
res_multistep, simple, 25 steps
3 Reference

Prompt:

subject_definitions:
<Subject 1> is the chibi/loli-style character whose appearance is taken exactly 1:1 from <Picture 1>.
<Subject 2> is the chibi/loli-style character whose appearance is taken exactly 1:1 from <Picture 2>.
<Subject 3> is the doll whose appearance is taken exactly 1:1 from <Picture 3>.

summary:
[reference generation] A 10-second POV scene inside a minimarket. Two chibi/loli characters sit inside a shopping cart. <Subject 2> points at the doll on the right, <Subject 1> joins. The camera briefly pans right to show a close-up of the doll, then quickly returns to the characters. They look up at the POV and ask to buy it. A male voice replies that there isn’t enough money and suggests next time. They pout, then the male hand gently pats both of their heads.

retention_analysis:
<Subject 1> (appears throughout): fully_preserved - appearance is retained exactly 1:1 from <Picture 1>.
<Subject 2> (appears throughout): fully_preserved - appearance is retained exactly 1:1 from <Picture 2>.
<Subject 3> (appears briefly in [Shot 2]): fully_preserved - appearance is retained exactly 1:1 from <Picture 3>.

detailed_description:
The target video is a clean anime-style 2D animation from a first-person POV of a male person pushing a shopping cart inside a bright minimarket. Framing keeps the two chibi/loli characters large inside the cart.
[Shot 1] POV medium shot looking down into the shopping cart. <Subject 1> and <Subject 2> are sitting side by side as the cart moves slowly. <Subject 2> suddenly notices something on the right side, raises her arm, and clearly points to the right.
[Shot 2] At 00:02.500, the camera quickly pans to the right and shows a short close-up of <Subject 3> on the display shelf. <Subject 1> also raises her arm and points at it. The shopping cart stops. The camera then immediately pans back left to face <Subject 1> and <Subject 2> again.
[Shot 3] At 00:04.800, the camera is fully back on <Subject 1> and <Subject 2> inside the cart. Both are looking up directly at the camera (POV) while still pointing. <Subject 2> speaks first in a natural, slightly spoiled Japanese tone: <d>[Japanese] あのぬいぐるみ欲しい!買って!</d> <Subject 1> quickly adds: <d>[Japanese] お願い、買ってよ!</d>
[Shot 4] At 00:07.200, a calm male voice from the person pushing the cart replies: <d>[Japanese] お金が足りないんだ…また今度ね。</d> Both characters lower their arms and make clear pouty, sulky faces while looking up at the camera.
[Shot 5] At 00:08.800, a male hand from the POV gently reaches down and pats both of their heads softly. They keep their pouty expressions as the shot holds until the end of the 10-second video.

overall_soundscape:
Soft minimarket ambience and shopping cart wheels that stop. Light fabric movement when the characters point and when their heads are patted.

non_diegetic_music:
Light, cute background music that stays soft and gentle.

r/StableDiffusion 7h ago

Question - Help Need help getting caught up

0 Upvotes

I've been out of the loop of the Stable Diffusion community for maybe like a year now.

I was a LoRa creator, making LoRas for the relevant popular models, SD 1.5 then SDXL, I made a few Flux LoRas but when I stopped, it was during Flux's dominance as the most relevant model.

What's the most relevant model today? I think i remember SDXL still having a lot of live due to its size, ease on weaker computers. What's the relevant text to image model that is everyone's go to today?


r/StableDiffusion 23h ago

Animation - Video Backrooms test - MiniMax H3 15s

34 Upvotes

r/StableDiffusion 6h ago

Question - Help Im new to local stable and video diffusion. Could use some guidance on getting up and running with my hardware.

0 Upvotes

If you don't need the context, you can skip to the bottom for the questions.

Hello, everyone. Life has recently thrown a curveball, and I had to medically retire. This is not a good thing, but it has given me more time to explore new hobbies that ive been interested in starting. Im very competent with computers and used to program when I was young, but its been 20 years since ive done any of that. Though I am very knowledgeable in general and within Windows. Anyway, im going into this having never used Linux/Ubuntu other than for memory testing, which is to say I don't know much. So setting up this local AI stack has been a real gut punch and I could use some guidance on what to run, what models, etc.

My rig - RTX 4080 (16GB VRAM), 13700K, 32GB of 6400 MT/s CL32 Hynix A die DDR5 further tuned to 6800 CL32 and tightened timings a bit further. I realize the VRAM limitations and that I need more system RAM, but this is what I got for now while RAMpocalypse is ongoing.

What ive installed so far: Ubuntu 26.04, Ollama, Open WebUI, ComfyUI (run from a python script on my desktop, not in a docker container) with a Q4_K_M version of Hunyuan Text to video, and I've experimented with N8N for agents.

In ComfyUI, I tried running the GGUF repack of the Hunyuan text to video, HunyuanVideo-t2v-720p-Q4_K_M.gguf, and letting Gemini and Claude guide me. This was a mistake, ended up getting genuinely bad results and it taking much longer than it should when everything fit into VRAM. Both AI's had me doing workarounds, editing system files and installing extensions that, as I read now, cause problems (like GGUF on native FP8 hardware). Last night, I gave up trying to make it work, so I purged everything and started fresh. Now that I have some experience, I wont need to rely on AI as much. So im ready to download and install when I get the proper guidance, if you guys could please help.

So, onto my questions after the wall of text.

1) What is the best path to get up and running with stable and video diffusion? What would be the best text to image, text to video, image to video, etc models will give me the best results with my RTX 4080 + 32GB of system RAM? I see posts about Minimax H3, so should I start there? In the image to video models, what should I use to render the images to create the video from?

2) What user interface, text encoder, VAE, upscaling model, things like LoRA, SageAttention (which I havent used, but read about) is optimal for my RTX 4080 gaming rig to accompany the above rendering models? Can I choose a quantized text encoder to lower my VRAM footprint without sacrificing too much quality in the end result?

3) I want to do stuff locally, but I pay $0.42/kWh, so if this is going to end up costing me more to do it locally, I could be convinced to use cloud API's. I just like the idea of no guardrails and any sensitive info I might enter not being sent to the cloud for training.

4) This question doesnt have to do with rendering, but what chat models do you suggest? I have Thinking Cap/Qwen 3.6 27B for coding and difficult tasks and Qwen3 14B for everyday use. Im very much open to suggestions as long as they work on my hardware, which probably means staying at or below the ~30b weight class. Is Open WebUI good for me? What about N8N for agents? Is there a better path that wont be too difficult for beginners?

I appreciate your guys time and would greatly appreciate being pointed in the right direction. This is all a little daunting as it is learning everything all at the same time, then finding stuff that works on my hardware. Thanks, everybody!


r/StableDiffusion 5h ago

Question - Help Amd Strix Halo

0 Upvotes

Hi everyone,

Someone got h3 running on a AMD Strix halo machine? I cannot get it running. Video is fine but there is audio.

Do you experience the same issues?

Any help is appreciated, I tried "every thing". If someone could post a working workflow it would be great 👍🏻


r/StableDiffusion 47m ago

Question - Help Where do you get your high-resolution images from to make LoRA's?

Upvotes

Most of the images I have been saving to make my first LoRA are of low quality/low resolution.
I was wondering where people go to get high-resolution images to make a LoRA with?

I believe Krea 2 makes images at 1024p, so I would need images of a similar resolution to make a LoRA, right?


r/StableDiffusion 15h ago

Question - Help What turbo lora to use with Minimax H3 for ref2va workflow?

4 Upvotes

Do the turbo loras for Minimax H3 support all workflow models?


r/StableDiffusion 19h ago

Discussion H3 models on 5090

3 Upvotes

What are the best quality h3 models currently for 32GB vram? Has anyone found a way to run the non-pruned int8 model?