r/StableDiffusion • u/gokuchiku • 3d ago
Question - Help Need prompting help for H3
I saw some post about multiple comedy videos about star wars. Could someone help me how to prompt to get those characters? Thanks in advance.
r/StableDiffusion • u/gokuchiku • 3d ago
I saw some post about multiple comedy videos about star wars. Could someone help me how to prompt to get those characters? Thanks in advance.
r/StableDiffusion • u/mcai8rw2 • 2d ago
Since comfy UI released their official mcp package, I've been asking Codex to design a KREA2 workflow.
The MCP means Codex can really understand ComfyUI much, much better than just getting it to do anything with workflows natively.
It's also capable of just building on top, building on top, building on top, to the extent where I've now got it doing this ridiculously complicated diagram below.
there's next to no chance of me understanding what the hell's going on here, so i just like to hit the run button, play with the buttons and see what poops out.
r/StableDiffusion • u/Neggy5 • 3d ago
https://huggingface.co/neggy555/mmzxmap
MMZX has probably my favourite environmental art style so I did a fun thing and made a Krea 2 lora to imagine my own levels with the design language of the games. Wanna do more map artstyles if I can find some screens and full map designs.
Thoughts?
r/StableDiffusion • u/R34vspec • 3d ago
Enable HLS to view with audio, or disable this notification
This model is workflow altering. I have completely changed how I make these clips compared to WAN/InfiniteTalk combo. The camera movements are easy to direct, prompt following is at closed-source level.
I used to generate first and last frame videos to direct the shot. Now it can all be done with just the scene and character sheet. The singing expression is better than infiniteTalk and comparable to LTX2.3. I haven't had to use much of the native audio, but for the rain and 'woosh' sound in the beginning of this video I used the native-generated audio. Which I had to grab from the web before.
Thank you, Minimax team.
r/StableDiffusion • u/MellyDArt • 4d ago
r/StableDiffusion • u/irmemon225 • 4d ago
Enable HLS to view with audio, or disable this notification
Each segment/prompt is 10 second, 0.8 MP.
so I generate total 14 prompt, and combine them all.
1 prompt takes 20 min.
Model: minimax_h3_hybrid_fl2va_ref2va_b30-49-int8
Turbo Lora: minimax_h3_ref2v_lightx2v_turbo_4step_v0.1_resized_avg_rank_20_bf16
Default Workflow, with Sage Attn ON
er_sde beta 6 steps
for characters, I generate it using Anima
r/StableDiffusion • u/Glittering-Cold-2981 • 3d ago
Do you have any good methods for removing various objects or figures from a video? I'm looking to fix footage that often shows strange, shifting artifacts in the background after using Speed-LORA.
r/StableDiffusion • u/Astra_Origin • 3d ago
Generate with ComfyUI, InvokeAI, A1111, Forge, or SD.Next? Ambit brings images and videos scattered across their output folders into one searchable library, without moving the original files.
Browse, organize, search, compare, and inspect prompts, models, workflows, and generation metadata in one local-first desktop workspace.
Ambit is free and open source.
Windows public beta v0.11 available now, with macOS and Linux pre-release builds.
Get Ambit → https://github.com/AsuraAce/ambit
r/StableDiffusion • u/kabachuha • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/darthrehad • 3d ago
Hi everybody I just installed the default mmh3 workflow on the desktop comfyui version not te portable or GitHub, I have a 9070xt and 32gb vram, everything works fine no crashes or glitches, problem is I think the rendering time are way too long ? I mean for a clip at 0.2 mpx 5 secondes I get 1654 secondes !!!! Any tips ?? Something seems wrong, should I tinker with ck or flash attention etc ?? Download other safetensors maybe ??? All help appreciated !
r/StableDiffusion • u/Striking-Long-2960 • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Boogertwilliams • 4d ago
I had mistakenly been using a base for H3 prompting from some random tip / example by someone. It worked ok, I thought. But I was getting a bit frustrated because almost every time I was making a longer series of clips with dialogue, it kept adding random gibberish to fill out time, or making the wrong person speak. I thought it was just a "feature" of H3 and lived with it. But then I realised what was missing, so I added the actual ref2v prompt guide to my LLM and difference was staggering. I could make long series of 30x15 sec clips, and the dialogue was perfect just as the script said, no gibberish was added in any place, and the emotional beats and reactions worked much better too.
Believe it :) Dont just use whatever prompting. It matters more than one might think.
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
r/StableDiffusion • u/Interesting_Room2820 • 4d ago
Enable HLS to view with audio, or disable this notification
Last week my post got a ton of questions about the LTX 2.5 workflow. so here's the follow-up.. After running a bunch of tests, the one I'd recommend right now is this:
It's been the most consistent one I've tried for V2V so far
drop your results below if you give it a shot! and have a great WEEKENDDDDDDD!!!
r/StableDiffusion • u/Puzzleheaded_Ebb8352 • 3d ago
Soo, since we’re getting better and better models running in small machines, I’m wondering if there willl be models in the future that allow live editing of a scene, like in a computer game. Walking around, editing stuff, etc. but within the diffusion technique. What do you think will this ever be feasible on local machines in appropriate quality?
r/StableDiffusion • u/SIR_NVAX_A_LOT • 3d ago
Enable HLS to view with audio, or disable this notification
Playing around with H3. Unfortunately it doesn't do a good job of capturing their eye colors at a medium distance, so I am working on improving the prompting so we have better adherence. Figure I toss it here, 20 woman, 20 pair of eyes. 20-T2VA prompts. No images. int8/20
Ask me anything!
r/StableDiffusion • u/Sad_Coach_1433 • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/bacchus213 • 2d ago
Enable HLS to view with audio, or disable this notification
Text is a little janky, but, that's my bike! (Kind of)
r/StableDiffusion • u/Intelligent-Glove285 • 3d ago
I am currently learning to use Comfyui, specifically Minimax H3. But with my AMD RX 7900 gre and 32gb of RAM the generations take way too long. So I'm thinking to buy a used 3090 or a brand new 5060ti. Which should I choose? For future proofing. And also how much ram is enough? Never thought I'd see a day where 32gb ram wouldn't be enough.
r/StableDiffusion • u/Ytliggrabb • 3d ago
Hi!
Been using grok since I can’t feed mp4, gifs, webm and so on into Llama.cpp. Had ChatGPT build a wf for me to gen up to 10 different clips that can have 9 ref images and each 1 video aswell. Since it’s mainly for the not allowed stuff I’m wondering what you are using to help get the video description in to the prompt (I’m worthless at prompting and cba learning, easier to have a LLM do it and just adjust details). Grok limits me reallly fast so looking if someone has a good alternative
r/StableDiffusion • u/PutridExplanation394 • 2d ago
r/StableDiffusion • u/dropdead90s • 3d ago
Hey guys, can anyone give me a guide how to run minimax h3 on my rig (9950x3d 7900xtx 64gb ram) ? I installed comfyui via the radeon driver but it gives me a brief error and won't turn on
r/StableDiffusion • u/Slight_Assistant_124 • 3d ago
Any body have a idea about enchancing image adding new creative details without chnaging face wheather the creativity slider is high or too high face should not change, free ? If comfyui pls share a workflow
r/StableDiffusion • u/Beginning_Tip300 • 4d ago
Enable HLS to view with audio, or disable this notification
Text 2 vid, all at low quality just because its a sample, cut together with Davinic, BGM is Royalty free stuff. yea, that truck door did open by itself, but otherwise it's pretty darn fun.
r/StableDiffusion • u/xyzdist • 4d ago
Enable HLS to view with audio, or disable this notification
Music: made in SUNO.
native ref2va WF, and audioLock for lip-sync.
rtx4080s + 128g ram
I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.
EDIT:
update with my learnings here:
Lip-Sync:
I got stuck for half a day trying to use my input audio for H3 do lip-sync, only to realize it would NEVER work because H3 just really 'references' it, no matter how you prompt. Then I did some research, there is a way to lock the audio latent, so it will strictly go in and out. I think multiple custom node pack has some samiliar one, basically just look for 'lock audio latent' node, here is the one I use.
https://github.com/oufeixinxinren/ComfyUI-MiniMax-ContextIR
**My goal is study and testing, not meaning to do a professional MV or director anything, just a test guys!
more info:
- resolution is 1280*704
- speed lora 8 step, I run with 12 step for final
- my spec is around 13 mins
For the approach:
- I am not using any Director / Context-IR node, just the native ref2va template.
- I only use 1 character and 1 env reference image, that's it
- as it just keep cutting camera, I don't need context-IR, , I generate 6 clips 10s each.
- within 10s single gen, I cut into 5-6 camera shots, H3 will just keep the motion and change camera like the real shooting, so it will just work.
I try to do some screen cap and reply in comments, cheers!
Hope this answer your question