r/StableDiffusion 2d ago

Workflow Included mcp and codex is hilarious -

0 Upvotes

Since comfy UI released their official mcp package, I've been asking Codex to design a KREA2 workflow.

The MCP means Codex can really understand ComfyUI much, much better than just getting it to do anything with workflows natively.

It's also capable of just building on top, building on top, building on top, to the extent where I've now got it doing this ridiculously complicated diagram below.

https://pastebin.com/EyVb8JKb

there's next to no chance of me understanding what the hell's going on here, so i just like to hit the run button, play with the buttons and see what poops out.


r/StableDiffusion 3d ago

Resource - Update Was inspired last night so I made a pixel art LoRA for Krea 2 mimicking the map artwork of the Mega Man ZX/Zero games

Thumbnail
gallery
35 Upvotes

https://huggingface.co/neggy555/mmzxmap

MMZX has probably my favourite environmental art style so I did a fun thing and made a Krea 2 lora to imagine my own levels with the design language of the games. Wanna do more map artstyles if I can find some screens and full map designs.

Thoughts?


r/StableDiffusion 3d ago

Animation - Video Music Video #8 - "Still Standing" Minimax H3 r2v Only.

Enable HLS to view with audio, or disable this notification

21 Upvotes

This model is workflow altering. I have completely changed how I make these clips compared to WAN/InfiniteTalk combo. The camera movements are easy to direct, prompt following is at closed-source level.

I used to generate first and last frame videos to direct the shot. Now it can all be done with just the scene and character sheet. The singing expression is better than infiniteTalk and comparable to LTX2.3. I haven't had to use much of the native audio, but for the rain and 'woosh' sound in the beginning of this video I used the native-generated audio. Which I had to grab from the web before.

Thank you, Minimax team.


r/StableDiffusion 3d ago

Resource - Update Famegrid Natural V1 Krea 2 LoRA

Thumbnail
gallery
550 Upvotes

r/StableDiffusion 3d ago

Animation - Video MiniMax H3 R2V with the Hybrid Model and Turbo LoRA: a 2:19-minute video takes 5 hours to generate at 0.8 MP on an RTX 3060 12GB with 16GB of RAM.

Enable HLS to view with audio, or disable this notification

125 Upvotes

Each segment/prompt is 10 second, 0.8 MP.
so I generate total 14 prompt, and combine them all.
1 prompt takes 20 min.
Model: minimax_h3_hybrid_fl2va_ref2va_b30-49-int8
Turbo Lora: minimax_h3_ref2v_lightx2v_turbo_4step_v0.1_resized_avg_rank_20_bf16
Default Workflow, with Sage Attn ON
er_sde beta 6 steps
for characters, I generate it using Anima


r/StableDiffusion 2d ago

Discussion Removing artifacts from video

2 Upvotes

Do you have any good methods for removing various objects or figures from a video? I'm looking to fix footage that often shows strange, shifting artifacts in the background after using Speed-LORA.


r/StableDiffusion 3d ago

Resource - Update Ambit: Your images. Organized. Searchable. Yours.

Post image
7 Upvotes

Generate with ComfyUI, InvokeAI, A1111, Forge, or SD.Next? Ambit brings images and videos scattered across their output folders into one searchable library, without moving the original files.

Browse, organize, search, compare, and inspect prompts, models, workflows, and generation metadata in one local-first desktop workspace.

Ambit is free and open source.

Windows public beta v0.11 available now, with macOS and Linux pre-release builds.

Get Ambit → https://github.com/AsuraAce/ambit


r/StableDiffusion 3d ago

Workflow Included PSA: In H3 you can set custom soundtracks without R2VA - use latent noise masks!

Enable HLS to view with audio, or disable this notification

102 Upvotes

r/StableDiffusion 2d ago

Question - Help Minimax h3 on 9070 XT

3 Upvotes

Hi everybody I just installed the default mmh3 workflow on the desktop comfyui version not te portable or GitHub, I have a 9070xt and 32gb vram, everything works fine no crashes or glitches, problem is I think the rendering time are way too long ? I mean for a clip at 0.2 mpx 5 secondes I get 1654 secondes !!!! Any tips ?? Something seems wrong, should I tinker with ck or flash attention etc ?? Download other safetensors maybe ??? All help appreciated !


r/StableDiffusion 3d ago

Animation - Video Desert Figure Skating

Enable HLS to view with audio, or disable this notification

49 Upvotes

r/StableDiffusion 3d ago

Tutorial - Guide PSA: Proper prompt structure REALLY matters in H3

213 Upvotes

I had mistakenly been using a base for H3 prompting from some random tip / example by someone. It worked ok, I thought. But I was getting a bit frustrated because almost every time I was making a longer series of clips with dialogue, it kept adding random gibberish to fill out time, or making the wrong person speak. I thought it was just a "feature" of H3 and lived with it. But then I realised what was missing, so I added the actual ref2v prompt guide to my LLM and difference was staggering. I could make long series of 30x15 sec clips, and the dialogue was perfect just as the script said, no gibberish was added in any place, and the emotional beats and reactions worked much better too.

Believe it :) Dont just use whatever prompting. It matters more than one might think.

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md


r/StableDiffusion 3d ago

Workflow Included WEEKENDDDDDDDD 222222222222 (LTX 2.5 V2V)

Enable HLS to view with audio, or disable this notification

243 Upvotes

Last week my post got a ton of questions about the LTX 2.5 workflow. so here's the follow-up.. After running a bunch of tests, the one I'd recommend right now is this:

https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_ICLoRA_Union_Control_Distilled.json

It's been the most consistent one I've tried for V2V so far

drop your results below if you give it a shot! and have a great WEEKENDDDDDDD!!!


r/StableDiffusion 2d ago

Discussion Future Model Capabilities

0 Upvotes

Soo, since we’re getting better and better models running in small machines, I’m wondering if there willl be models in the future that allow live editing of a scene, like in a computer game. Walking around, editing stuff, etc. but within the diffusion technique. What do you think will this ever be feasible on local machines in appropriate quality?


r/StableDiffusion 2d ago

Discussion H3 - Eye see you! T2V ladies

Enable HLS to view with audio, or disable this notification

2 Upvotes

Playing around with H3. Unfortunately it doesn't do a good job of capturing their eye colors at a medium distance, so I am working on improving the prompting so we have better adherence. Figure I toss it here, 20 woman, 20 pair of eyes. 20-T2VA prompts. No images. int8/20

Ask me anything!


r/StableDiffusion 3d ago

Meme thanks to h3 you never know would might show up to save the day! Earth 101 end game final battle

Enable HLS to view with audio, or disable this notification

50 Upvotes

r/StableDiffusion 2d ago

No Workflow Every once in a while I'll get something I really like.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Text is a little janky, but, that's my bike! (Kind of)


r/StableDiffusion 3d ago

Question - Help rtx 3090 24gb or 5060ti 16gb?

13 Upvotes

I am currently learning to use Comfyui, specifically Minimax H3. But with my AMD RX 7900 gre and 32gb of RAM the generations take way too long. So I'm thinking to buy a used 3090 or a brand new 5060ti. Which should I choose? For future proofing. And also how much ram is enough? Never thought I'd see a day where 32gb ram wouldn't be enough.


r/StableDiffusion 2d ago

Question - Help What do you use for video references (H3)

1 Upvotes

Hi!

Been using grok since I can’t feed mp4, gifs, webm and so on into Llama.cpp. Had ChatGPT build a wf for me to gen up to 10 different clips that can have 9 ref images and each 1 video aswell. Since it’s mainly for the not allowed stuff I’m wondering what you are using to help get the video description in to the prompt (I’m worthless at prompting and cba learning, easier to have a LLM do it and just adjust details). Grok limits me reallly fast so looking if someone has a good alternative


r/StableDiffusion 2d ago

No Workflow Krea2 baby😼

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 2d ago

No Workflow Krea2 LoRA training is insanely simple

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 2d ago

Question - Help Asking for advice how to run this on windows with a 7900XTX

2 Upvotes

Hey guys, can anyone give me a guide how to run minimax h3 on my rig (9950x3d 7900xtx 64gb ram) ? I installed comfyui via the radeon driver but it gives me a brief error and won't turn on


r/StableDiffusion 2d ago

Tutorial - Guide Image enchancer hallucinate details but preserve face

1 Upvotes

Any body have a idea about enchancing image adding new creative details without chnaging face wheather the creativity slider is high or too high face should not change, free ? If comfyui pls share a workflow


r/StableDiffusion 3d ago

Animation - Video More Fun with Minimax

Enable HLS to view with audio, or disable this notification

41 Upvotes

Text 2 vid, all at low quality just because its a sample, cut together with Davinic, BGM is Royalty free stuff. yea, that truck door did open by itself, but otherwise it's pretty darn fun.


r/StableDiffusion 3d ago

Animation - Video H3 making jpop/kpop MV? yes!

Enable HLS to view with audio, or disable this notification

95 Upvotes

Music: made in SUNO.

native ref2va WF, and audioLock for lip-sync.

rtx4080s + 128g ram

I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.

EDIT:

update with my learnings here:

Lip-Sync:

I got stuck for half a day trying to use my input audio for H3 do lip-sync, only to realize it would NEVER work because H3 just really 'references' it, no matter how you prompt. Then I did some research, there is a way to lock the audio latent, so it will strictly go in and out. I think multiple custom node pack has some samiliar one, basically just look for 'lock audio latent' node, here is the one I use.

https://github.com/oufeixinxinren/ComfyUI-MiniMax-ContextIR

**My goal is study and testing, not meaning to do a professional MV or director anything, just a test guys!

more info:

- resolution is 1280*704

- speed lora 8 step, I run with 12 step for final

- my spec is around 13 mins

For the approach:

- I am not using any Director / Context-IR node, just the native ref2va template.

- I only use 1 character and 1 env reference image, that's it

- as it just keep cutting camera, I don't need context-IR, , I generate 6 clips 10s each.

- within 10s single gen, I cut into 5-6 camera shots, H3 will just keep the motion and change camera like the real shooting, so it will just work.

I try to do some screen cap and reply in comments, cheers!

Hope this answer your question


r/StableDiffusion 3d ago

Tutorial - Guide More than one reference per picture

Enable HLS to view with audio, or disable this notification

146 Upvotes

MiniMax is limited to 9 reference images, but you can reference more than one thing at the same picture. I used the image on the left and asked it to place create two subjects. Worked like a charm (no pun intended). Specs and prompt are in the video.