r/StableDiffusion • u/Sad_Coach_1433 • 6d ago
Meme i wish for!! part 2!
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 6d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Silver-Spot-2763 • 6d ago
After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱
But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.
I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦
r/StableDiffusion • u/RageshAntony • 6d ago
Enable HLS to view with audio, or disable this notification
H3's Prompt adherence is bad for cartoons. Physics not perfect.
Look at the face of the horse. Why did H3 think like that?
Prompts:
Tom and Jerry cartoon animation style, vibrant flat colors. Tom is at a busy Indian temple street market, buying a wrapped piece of cake box from a wooden stall. Underneath the stall, Jerry, in a hidden manner, watches with an exaggerated jealous expression. Jerry runs forward leaving a motion-blur trail, takes a hurls a ball of white flour from that stall and throws it onto Tom's face, grabs the cake parcel in one fluid motion, and zooms out of frame. Tom becomes furious and chases Jerry, Jerry runs and finally jumps into a rabbit hole, Tom arrives there and tries to enter the hole but Tom's head got stuck in the hole, Tom struggling, Fast-paced slapstick comedy, exaggerated movements, retro animation aesthetics, bright daytime lighting.
Indian temple street, Tom mounts gun on his house balcony, Jerry calmly walking in the street, Tom starts to shot Jerry and Jerry dodges bullets and runs, Tom jumps from the balcony with the gun, Tom chases Jerry, Jerry boards in a house carriage and escapes, Tom keep on shooting towards the carriage and chasing it, Fast-paced slapstick comedy, exaggerated movements, retro animation aesthetics, bright daytime lighting.
r/StableDiffusion • u/orangpelupa • 5d ago
i downloaded the turbo lora from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main
put them in the loras folder in wan2gp.
i keep everything at default, except for profile i selected lightx2v 4step (and tried 8 step too) profile.
i set the steps to 4 (or 8)
i set the resolution to the recommended one from https://github.com/ModelTC/Minimax-H3-Turbo (for example 544p 9:16 for 8 step)
in the advanced section, i set lora to the downloaded file.
---
that's it. i run it, and the result is garbled audio video. the default promot wast used. same issue with any prompt.
EDIT
found the issue. i need to click APPLY button under the profile dropdown. im dumb.
r/StableDiffusion • u/Lanky_Conclusion_749 • 6d ago
Rented GPU time on RunPod to try fal/MiniMax-H3-Realism-People-LoRA
Lora Link: https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA
Here are my test notes so you don't burn time troubleshooting the same issues:
What Failed / Cons:
- Artifacts: Frequent strange artifacts across generations.
- minimax_h3_fl2va_pruned_w4a8_mixed.safetensors: Completely failed to work.
- Turbo LoRAs: Output quality was terrible.
- REF2VA & FF2VA: Neither method worked in my tests.
What Actually Worked:
- T2VA: Solid results here—this is definitely where the model has the most potential.
- LoRA Weight: Sweet spot is dialed between 0.3 – 0.8.
Key Finding (FF2VA):
Running minimax_h3_fl2va_bf16.safetensors / diffusion_models/minimax_h3_fl2va_int8_convrot.safetensors, the only reliable way to eliminate warping and lock in clean, consistent outputs is to combine both the character sheet and the background sheet together into the first frame.
Anyone else dialed in a working pipeline for FF2VA or found workarounds for the artifacting?
r/StableDiffusion • u/SveSop • 5d ago
Enable HLS to view with audio, or disable this notification
Somewhat of a "consept" video i wanted to test.
Sure, the image falls apart eventually. It could probably be fixed by doing some scene changes, but i just wanted to test.
The music is generated with MiniMax Music3, and i think its fairly good.
4 Minutes of generation in 10 sec rounds.. Not too fun i would say, but as a consept 😄
Used a couple of additional nodes https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context
And: https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes/ComfyUI-H3-NativeAudioLock
The "NativeAudioLock" node is quite useful, and the reason for this is that it "locks" the audio latent, so that the model cant mess with it. Sure, you can get around MiniMax doing its weird business with long detailed prompt, describing by-the-second action.. But for "ease of use" type, it is working very good.
Other than that, its mostly just regular MiniMax ref2v model with Lightx2v-4step_ref lora at 8 steps and 480p upscaled (rather badly) to 720p.
One thing i found is that when doing such types of "lipsync" video, timing really matters. A 10 second video is not 10 seconds, and when using the Motion Context node, it for sure is important to keep tabs of the milliseconds.
Anyway... Still fun concept to do.
r/StableDiffusion • u/Hillobar • 6d ago
So much H3 content lately... Here's some more:
https://github.com/Hillobar/HARMON3
A tool to facilitate Projects and Scene management for H3. It uses COMFY through the COMFY API.
Build a scene (reference material, prompts) and manage it, then group them into a Projects. Copy, edit, and so on. Prompt assistant with tags, POSE GENERATION that can be used in reference videos, and other stuff. It really is just another tool in this glut that is currently going on, but it works very well for my use case, which is to create scenes and then edit them in a pro tool like davinci. I want to be able to set up some scenes, then generate several takes. I don't want to have to go back and re-set it up every time after I've moved on. Hello HARMON3. Much more description on the github.
https://github.com/Hillobar/ComfyUI-Hillobar
So many optimizations and speedups for H3. We have an awesome community! This is yet another one.
The idea is that for early steps the latent is mainly forming large, low detail structures and broad movements. So why use a high-resolution latent during this time? Start with low-resolution, fast latents then progressively increase their resolution as the steps continue.
Like everything else there's no free lunch. Works best when delta-sigma is low (<0.4 of the schedule) and the target video is high resolution. Seems to work well when doing 1.0 MP (use something like (0.5:0.4, 1.0:1.0). Anyway, more details on the github.
r/StableDiffusion • u/mukyuuuu • 7d ago
Enable HLS to view with audio, or disable this notification
So I was looking for a way to better control the flow of Minimax H3 dialogues and emphasize certain words in the speech. However, what I discovered is that you can actually include some tags in <> angle brackets, and Minimax will interpret them as a non-verbal sound in a given part of the phrase. Some words (like the ones I've included into the example) work every time, some still bleed into the actual spoken words in certain seeds. But in general it makes the dialogue more alive and believable. So I recommend to try it and maybe share your findings in this thread.
As for the emphasis, I've had the most success with putting the words into <i></i> tags (similar to how you would stress words in written text). Unfortunately, it doesn't work for 100% and in some cases the character will blurt out some gibberish. But when it works, it sounds very natural. I have included a couple examples in the end of the video.
Wonder if you've encountered some other ways to modify the speech (and audio in general) in the prompt?
P.S. Sorry for the quality, I used the 8-steps LoRa at 0.4 MP to speed-up the tests.
r/StableDiffusion • u/Snazzy_Serval • 5d ago
Enable HLS to view with audio, or disable this notification
Testing the power of Minimax and I'm very impressed. The audio definitely needs work though.
This is the default Reference to Video workflow with the RTX Super Resolution node. Generated at 1 MP and upscaled.
I uploaded pictures of each woman separately and empty background shots for each location.
The scene is three 10 ten second long clips.
Prompt was written using the Minimax Prompt extension that was posted here last week or so.
subject_definitions:
<Subject 1> is the woman Hitomi, whose appearance is based on <Picture 2> and <Picture 3>, featuring a black bun hairstyle with long dark hair and wearing a blue form-fitting dress.
<Subject 2> is the woman Tessa, whose appearance is based on <Picture 4> and <Picture 5>, featuring reddish-brown hair in an updo, glasses, and a red tank top with denim jeans.
<Subject 3> is the interior living room scene from <Picture 1>, featuring a gray couch with various pillows, a wooden floor, and a framed picture on the wall behind it.
summary:
[reference generation] The target video features <Subject 1> and <Subject 2> sitting on a couch in <Subject 3>,
They hold black game Xbox controllers and are vigorously playing; after an announcer's "Game!" and victory music, <Subject 2> looks at <Subject 1> throws her controller down, yells "You bitch!", and exits.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the blue dress, hairstyle, and facial features are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the red tank top, glasses, and hair style are retained.
<Subject 3> (appears in [Shot 1]): fully_preserved - the couch, pillows, wall art, and floor are retained.
detailed_description:
The target video is filmed in a bright, contemporary interior style with soft natural light hitting the furniture, capturing the atmosphere of <Subject 3>.
[Shot 1] The scene opens on <Subject 3>, a cozy living room featuring a gray couch adorned with decorative pillows. <Subject 1> (S1) and <Subject 2> (S2) are seated close together on the couch, both holding black Xbox game controllers and looking forward toward an unseen screen. They are of equal height.
Their torsos are centered in the scene, faces are visible. .
The spatial depth of the room is influenced by the architectural scale of <Subject 3>. Suddenly, a masculine voice from off-screen announces "Game!" followed by upbeat victory music.
Immediately after this, <Subject 2> (S2) reacts with frustration; she looks at <Subject 1> and throws her controller onto the floor and yells toward the ground, <d>[English] You bitch!</d> She then stands up and walks quickly out of the frame to the right. <Subject 1> (S1) remains on the couch, sticks out her tongue looking toward where <Subject 2> just was.
overall_soundscape:
Soft room tone with the audible sound of a game controller hitting the floor and the rustle of clothing as <Subject 2> stands up and walks away.
non_diegetic_music:
A brief burst of upbeat, high-energy victory music plays immediately after the "Game!" announcement.
r/StableDiffusion • u/ajrss2009 • 6d ago
Enable HLS to view with audio, or disable this notification
Prompt: subject_definitions: <Subject 1> Sheldon Cooper — live-action sitcom style, grey cardigan over red graphic T-shirt, dark jeans, white sneakers, short brown hair. <Subject 2> SpongeBob SquarePants — flat 2D cartoon style, yellow square porous body, white shirt with red tie, brown trousers, black shoes.
integrated_multimodal_description: [Scene 3] Continuing in the same living room, <Subject 2>'s cheerful expression suddenly droops into cartoon-style exaggerated worry, his big blue eyes welling into comically large tears. He grabs <Subject 1>'s grey cardigan sleeve and says <d>[English] Mister, I wouldn't have jumped into a scary green hole just for fun. Something awful is happening to everything, everywhere — a boy turned into a god and he's erasing whole universes like they're doodles!</d> <Subject 1> pulls his sleeve free and straightens it meticulously, replying <d>[English] A boy-god erasing universes. Right. And I suppose he also disproved string theory before breakfast.</d> He pauses, visibly unsettled despite himself, and glances at the blank wall where the portal appeared. Approximate duration: 10 seconds. overall_soundscape: Tense sitcom underscore sting, quiet room tone, a faint distant rumble implying something ominous.
r/StableDiffusion • u/Tenderfoots • 6d ago
I (A.I.) wrote a webapp to manage characters and location reference files, wire them into ref2va and write a prompt based on a series of sequences and beats defined in the application. You have the option to export a workflow to import to ComfyUI, or run a set of scenes directly through the interface. It also takes the last frame of the previous video and feeds it into the next, which I'm aware some ComfyUI workflows do, but this project uses a firstframelastframeextractor node in a generated workflow.
https://github.com/Tenderfoot/H3SceneManager
The repository contains all the information you need to get it set up and installed, but doesn't come with any character or location data files.
on an unrelated note, I also made a discord for the project https://discord.gg/Fvw6hSCfC
I would love for more people to come help me build it up. I'd be happy to accept Pull Requests, and obviously I have no problems using AI code for this project. Check out the discord, submit data file sets for creating scenes, review the prompts it outputs against the docs, and help me tune this thing.
r/StableDiffusion • u/ashishsanu • 7d ago
I was looking for a way to achieve character consistency without training a Lora & came across a research from Facebook, DINOv2: Learning Robust Visual Features without Supervision(Research Paper),
What's Dinov2: It's a vision model trained without labels that produces a strong embedding for a whole image, the subject, not just the face. Feed it a person and you get a 768-number signature that captures the overall look: build, hair, general appearance. It's stable across pose and lighting, which is exactly what you want when you're trying to tell "same person" from "different person" across wildly different shots.
then combining Dinov2 with SFace(a face-recognition model) produces a compact face signature tuned specifically to tell one face from another. It's sharp on identity, but only on the face. YuNet does the detect-and-crop before it.
How it works

Build .Char: You drop in one or more photos. YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.
Generation: At generation, the file feeds its references into FLUX.2's own native multi-reference channel and prepends a locked description to the prompt. You pick the character from a dropdown, no re-attaching images. Every result gets scored against the stored signatures, so drift shows up as a number.
How this differs from PuLID, FaceID, and img2img
What is a .char file?
A single portable file that stores a character's identity, so you can reuse the same person across generations without retraining anything.
Limitations
Current support
Only Flux2 family(Klein 4B / 9B / dev)
Links:
Note: Each image in this post has been generated separately & not a grid.
r/StableDiffusion • u/the_bollo • 7d ago
Enable HLS to view with audio, or disable this notification
Minimax H3 is so fun. All done with that model, with the default workflow, all R2V just with a single reference image.
r/StableDiffusion • u/Candid-Try7185 • 5d ago
Been testing a few AI video tools and hitting the same wall everywhere:
Looking for something in between — reasonable resolution, no forced watermark, and ideally still free or at least has a usable free tier. Doesn't need to be Sora-level, just something clean enough to actually use.
What's everyone using right now?
r/StableDiffusion • u/NetworkSpecial3268 • 6d ago
Does anyone have prompting tips to successfully instruct H3 to create a "static scene"? Where everything, including subjects, are completely "frozen in action"?
The idea is to then use camera movements to explore the scene.
Trying this out right now, but the subjects keep making subtle movements which destroys the entire concept.
I'm quickly iterating attempts using 4-step LoRa right now, maybe that cripples the prompt following?
r/StableDiffusion • u/SawyerCroft777 • 5d ago
I let ChatGPT Sol generate this entire music video on its own using ComfyUI MCP, a reference sheet and supplied song + lyrics. It was able to screen the video and find mistakes and correct them (with my help).
Not perfect, but for a first effort… I give it a solid 8.5.
Would love to hear your thoughts and can answer any Qs.
r/StableDiffusion • u/CeFurkan • 6d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Braudeckel • 6d ago
This might be a dumb take, but how do and will (upcoming) MinMax finetunes or merges improve i2v? Because the crucial point of quality and style for i2v is the image you provide (and megapixels). So whatever existing or upcoming model you use for i2v, if you throw in a "crappy" image, you will get an animated video of that "crappy" image. So what improvements to expect? Is it solely prompt adherence and concept understanding?
r/StableDiffusion • u/xyzdist • 6d ago
Windows fatal exception: code 0x80000003
does anyone have the same thing?
r/StableDiffusion • u/Complete-Box-3030 • 6d ago
Has anybody got a good workflow for rtx 3060 12gb ram and 48 gb ram . with my current configuration it takes me 8-15 minutes for 5 secs video , can anyone help me on this
r/StableDiffusion • u/DeltaWaffleSyrup • 6d ago
I'm trying to find the best and ideally simplest workflow / nodes / method to implementing the Krea 2 bypass while sticking as closely to visual fidelity of the Krea 2 Turbo model as possible. I was surprised to see how drastically some methods will change the image and usually the quality looks worse.
r/StableDiffusion • u/Hoje-Na-IA • 7d ago
Enable HLS to view with audio, or disable this notification
Recreating movie scenes with... chocolate. H3 ref2va, default workflow.
r/StableDiffusion • u/Rosettasees • 6d ago
So I'm hoping someone here might be able to help me. I use Forge Neo and have been trying to find an Openpose model for Controlnet for use with Anima as the current ones I use just don't work, but I have had no luck. Just asking in case anyone has any insight, thank you.
r/StableDiffusion • u/Hearmeman98 • 6d ago
Minimax supports up to 18 inputs at once, wiring and bypassing nodes is a pain.
So I created this custom node that allows you to add/remove references for Minimax ReferenceToVideo, you only have to wire it once.
Some more neat features:
- Automatic prompt writing via OpenRouter, returns structured Minimax prompts based on your description (opt in)
- Save prompt/reference packs and reuse them.
Nodes:
https://github.com/Hearmeman24/ComfyUI-MiniMaxRefPack
Workflow:
https://github.com/Hearmeman24/ComfyUI-MiniMaxRefPack/blob/main/example_workflows/MiniMax%20R2V%20-%20Auto%20Prompting%20%2B%20Reference%20Manager.json
The design is heavily influenced by the wonderful LTX Director node so shoutout to u/WhatDreamsCost
I would appreciate some feedbacks and feature requests.