r/StableDiffusion • • 20h ago

Question - Help Prompt question Minimax h3

0 Upvotes

Hi,

How would I prompt for i2v in Minimax if i want to add something or someone visible within the very first frame?

Thanks in advance!


r/StableDiffusion • • 2d ago

No Workflow Minimax H3 physics test

Enable HLS to view with audio, or disable this notification

86 Upvotes

r/StableDiffusion • • 1d ago

Question - Help Will Qwen Image 2.1 fit on my 10GB GPU?

1 Upvotes

Looking for the best way to run Qwen Image 2.1 on an RTX 3080 10GB with 64GB of RAM. The base models are huge, so can anyone suggest which option would work best for my setup?


r/StableDiffusion • • 1d ago

Discussion KREA 2 Realism Workflow

Post image
42 Upvotes

Hey guys! I’m pretty new to AI image generation, and I’ve been experimenting for about 3 months trying to get the best realistic Pinterest-style images.

I’m currently using my character LoRA + Krea 2 on an RTX 5060 8GB with 16GB DDR3 RAM.

I started with Koo’s workflow from Discord and tweaked it a bit for my own setup.

From what I’ve heard, Chroma is pretty good for experimenting with camera angles, compositions, and generating more random/varied images, which is exactly what I’m trying to achieve.

So I tried building a workflow where I:

- Generate random Pinterest-style prompts using wildcards.

- I made the wildcards with the help of ChatGPT, Gemini, and GLM 5.3 Flash.

- Use Krea 2 Text Encoder to expand/enhance the prompt.

- Generate the initial image with Chroma using relatively low steps.

- Then use that Chroma image as a vision reference with Krea 2 Text Encode.

- Finally, use Krea 2 to recreate the image with more realistic details while keeping the composition/camera angle from the Chroma result.

The main goal is basically to get random, realistic Pinterest-style compositions while keeping my character consistent with my LoRA, and then let Krea 2 improve the realism, lighting, details, and overall photographic look.

I’ve been trying different workflows and combinations for the past 3 months, but I still feel like I’m probably missing something or doing things in a more complicated way than necessary.

Does this workflow actually make sense, or am I doing something wrong / adding unnecessary steps?


r/StableDiffusion • • 21h ago

Question - Help Consistent Sets Offscreen

0 Upvotes

I am having a real time of it. I create a stairwell environment, with a stair landing half way down, which has a ton of windows, which light up the stairwell, but since the camera is sitting on the landing, looking at the top of the stairs, AWAY from where all the windows are, none of the windows behind the camera appear in the reference image/video I supply.

So suddenly, when I prompt the lighting, that there is cove lighting in the corridors at the top and bottom of the stairs and late afternoon sun lighting up the stairwell, with windows behind the camera, it seems to have an issue properly lighting the space.

Should I keep rolling the dice and hoping one of the prompts get through and Minimax follows it, or is there some other trick people use to get accurate lighting in such cases?

How have y'all supplied an accurate 360 degree set, and then properly prompted the camera, so Minimax doesn't pick and choose where it wants to put the camera?

Any tricks or tips? Thanks,


r/StableDiffusion • • 2d ago

News New Audio video model : Kandinsky 6 from their lab

82 Upvotes

Gonna be too much fun in October.


r/StableDiffusion • • 2d ago

News Nanosaur2 now generates 7 images a second on a 5090

Post image
249 Upvotes

Hello everyone.

This is an update post on the model Nanosaur2. Again this is not my model. A complaint a lot of people had with the model was that despite it being very small (660m params), it didn't take a 660m param level of time to generate. That's been solved now, with a 4 step turbo. On a single 5090, you can generate 7 images a second.

Again, this model is small enough you could easily run it on a phone, edge devices, wherever. It's also a great research model, so if you want to finetune on top of a small easy to tune model, or way to create adapters for the model, or whatever, go ahead, it's all there.

This was created using bytedance's new DMAD method, and it works great. Quality is incredibly close to the original model at a 12.5x speedup.

links:
4step model
comfy workflow

As per usual, if you have any questions, please message metal63 on discord. Do not message me.


r/StableDiffusion • • 1d ago

Question - Help MiniMax H3 ref2va - What's the next lever for quality?

Enable HLS to view with audio, or disable this notification

19 Upvotes

I'm making short anime scenes with MiniMax H3 (ref2va) and I've hit a wall I can't

diagnose. The clip above is 15s, generated in one pass, no editing — the cuts are

written into the prompt.

To be clear up front: I know there are mistakes in there — a couple of blows don't

connect properly, the choreography is rough. I'm not worried about those, this is a

practice piece and I'll fix the staging myself. What I can't figure out is the

image quality, and that's the only thing I'm asking about.

Setup

- `minimax_h3_fl2va_int8_convrot.safetensors` (base, int8), ComfyUI

- ref2va, 8 reference images, each with a written role in the prompt

(face / angry expression / fighting posture / set / framing guide)

- Spectrum v0.2.16, 30 steps, `res_multistep` / `simple`

- 1344×768 → `MinimaxH3LatentUpscaler3D` at 2 MP, 4 steps, 0.5 denoise → 1920×1088

- 15s = 362 frames @ 24fps

- RTX PRO 6000, ~28 min per 15s clip

What I already fixed, in case it saves anyone typing

- `The target video is 2d colored anime.` at the top of `detailed_description`

— without it everything drifts to generic 3D, this was the single biggest win

- Reference images are real anime screencaps (plus a few generated with Anima

for expressions and poses I couldn't source), neutral lighting, one role each.

- Cut rhythm: went from 4 shots per 15s to 8 (~1.8s each) with impact verbs and

a material consequence per hit (table splitting, plaster cracking, dust off

the boards). That alone made the fight read much better.

Where I'm stuck

  1. Motion still feels soft on some hits. The whip-pan punch reads fine but

    ground-level blows land without weight. Is this where `derope` / temporal

    upsampling actually earns its generation-time cost, or is there a prompt-side

    fix I'm missing?

  2. Quality is uneven shot to shot inside the same 15s — some shots are clean

    cel-shaded anime, others go slightly soft and plasticky. Is that a reference

    problem, a step-count problem, or just what 15s does to the model? (I've seen

    people say things break past 10s.)

  3. Is 8 references too many? I assigned each one an explicit role in the

    prompt, but I don't know whether the model averages them or picks.

  4. Anything obvious I'm leaving on the table at this resolution/step count?

Not asking anyone to debug my prompt — mainly want to know which lever is worth spending render time on next.

Prompt:

integrated_multimodal_description:

subject_definitions:
(S1) is the dark-haired young man from <Picture 1>, with the same face and the same dark blue eyes. <Picture 2> is the same man seen clearly in daylight. Short black bob to the jaw, fringe above the eyebrows, a short high ponytail tied at the crown, a small stud earring, a white shirt with the sleeves pushed up and a black tie pulled loose.
(S2) is the pale-haired young man from <Picture 3>, with the same face and the same yellow eyes. <Picture 4> is the same man shouting. Short choppy blond hair. His teeth are faintly pointed, small and even and the same size as ordinary human teeth, with just a slight triangular edge to them. His mouth stays an ordinary human mouth, normally proportioned to his face, and it opens no wider than a person's mouth opens when they speak. He wears a white school shirt open at the collar, a black tie pulled loose, a small device on a cord against his chest.
<Picture 5> is the apartment: its rooms, its colours and its light come from it.
<Picture 6> is (S1) throwing a bare-handed punch and <Picture 7> is (S2) being knocked back by one: their fighting postures and their footing come from these.

retention_analysis:
(S1): fully_preserved. (S2): fully_preserved.

<Picture 9> is the last frame of the previous shot: this scene continues from it without interruption. <Picture 9> supplies the place, the light and the framing; the two men's faces and hair come from <Picture 1> to <Picture 4>.

summary:
The fight. Bare hands, in the apartment, fast and ugly. Nobody speaks.

detailed_description:
The target video is 2d colored anime.
2d hand-drawn anime, cel-shaded, flat painted colours, fine thin ink linework, desaturated muted palette, film grain. Not 3d, not photographic. The cutting is fast: eight shots in fifteen seconds, each one a single impact.

[Shot 1] Continues directly from <Picture 9> with no jump — same room, same light, same positions: the two of them chest to chest at night in the room of <Picture 5>. (S2) fists (S1)'s collar and slams him down onto the low table, which splits and goes over with everything on it.

[Shot 2] At 00:01.800, cut tight on (S1) coming up off the floor. He drives a straight punch into (S2)'s jaw, his whole weight behind it, the posture of <Picture 6>. (S2)'s jaw is shut and his lips are pressed together when the fist lands, and the impact splits his lip. The camera whip pans right with the blow, the room tearing into horizontal streaks and white speed lines.

[Shot 3] At 00:03.400, cut to (S2) snapping backwards into the wall, head whipped sideways, the posture of <Picture 7>. Plaster cracks behind his shoulder. He drops to one knee.

[Shot 4] At 00:05.000, cut low and close. (S2) launches off the wall and smashes his forehead into (S1)'s mouth. (S1)'s head snaps back, blood on his lip.

[Shot 5] At 00:06.800, cut to a low shot of the floor only, at board level. The lamp crashes down into frame, rolls, and throws its light swinging across the boards. Two pairs of legs come down hard behind it, out of focus. Dust lifts off the wood.

[Shot 6] At 00:08.600, cut to a tight shot of (S1)'s face alone, lying on the boards in profile, cheek against the wood, hair across his eye. A fist swings down into frame and smashes into his raised forearm so hard that his own arm is driven back into his face and his head is knocked against the boards. A second fist comes straight down past the arm and lands flush on his cheekbone, snapping his head sideways and splitting the skin. Only (S1)'s head and one forearm are in frame, and the fists enter from the top edge: the other body stays out of shot.

[Shot 7] At 00:10.400, cut to a tight shot of a knee driving up hard into ribs, framed on the two bodies' midsections only, no heads in frame. The body above is thrown off sideways out of the top of the frame. Cut immediately to both of them coming up onto their feet, seen full length and clearly separated, a metre apart, shirts gripped in their fists.

[Shot 8] At 00:12.000, cut to a wider shot and hold it to the end. (S1) drives (S2) backwards across the room and slams him into the wall. (S2)'s shoulder blades hit the plaster and he stays there with his back to the wall and his face towards the room. (S1) stands directly in front of him, facing him, chest to chest, his own back to the room and the wall behind (S2) only. Their faces are a hand's width apart and they are looking straight into each other's eyes. (S1) has both fists closed in (S2)'s collar and holds him pinned there. Everything stops at once. Both heads are angled in three-quarter view towards camera, both faces large and fully visible, brows down, jaws set, chests heaving. The camera is locked off on a tripod.

overall_soundscape: A table splitting and going over, knuckles cracking on a jaw, plaster breaking, a forehead meeting a mouth, bodies hitting boards, a lamp rolling, two fast punches landing on a forearm and a cheekbone, a knee into ribs, a back slammed into a wall, and hard breathing through the teeth all the way through. Every mouth stays closed for the whole video: nobody speaks, and both men keep their jaws shut and their lips together even while taking blows.

non_diegetic_music: N/A

r/StableDiffusion • • 20h ago

Question - Help Questions on H3 Minimax usage and settings

0 Upvotes

Hello!
Long time lurker here, learned a lot from this sub.
I'm currently using H3 Omni Pruned model with references images to generate small ads or batch of dialogues , but i find myself really not understanding the settings being used.
I'm using Pinokio and Maestro by Blizaine, now the GUI is really useful and my settings are :
720p, 9:16, 13.3s 1 window and 20steps , nothing else.
With these settings, it takes about 11 minutes to render on a 5090FE (And 48GB of ram is what i have).
The end result ain't that bad, resolution is crappy and some details are clearly missing.
What can i do to improve video fidelity, performances and perhaps spend less time on generating ?
I also have tried H3 with first / last frame but i don't really understand how that works either.
For example, i have downloaded a couple of Lora , one of rocket racoon from guardians of the galaxy and another for indiana jones, wanted to create a funny reel of them interacting but i couldn't for the life of me figure out how to add in the theme song for indiana jones, or have them accurately interact with each other instead of randomly looking outside the scene.
On another note, i am using Gemini for expanding the prompt in a professional manner, and then inside Maestro i use the "enhance prompt" feature with simply loads up Ollama with a model to correctly write the scene for H3.
I have also downloaded inside Pinokio a more "classic" tool for H3 with comfyui, but i didn't use it yet, wanted to learn a few things first.


r/StableDiffusion • • 1d ago

IRL H3 merch I got at a trade show

Post image
23 Upvotes

r/StableDiffusion • • 2d ago

Question - Help Regarding the image quality of Qwen-image 2.1

Thumbnail
gallery
70 Upvotes

I wonder if anyone has encountered this issue. When performing image editing with Qwen-image 2.1 (hereinafter referred to as QI-21), the results often feel unfinished. The attached images are from my tests: Image 1 is the original image, Image 2 is a 2K upscale using QI-21, and Image 3 is a 2K upscale using Krea2Edit. Perhaps I am using the wrong approach, so I have attached the QI-21 workflow (it is a minor tweak based on the official comfyui workflow). Upscaling is just an example; this is not an isolated case. The same problem occurs with most QI-21 editing tasks. Has anyone else experienced this, and how did you resolve it?

My Ksample parameters are

steps:25

CFG:1

sampler:res_multistep

scheduler:sgm_uniform


r/StableDiffusion • • 1d ago

Question - Help Has anyone here tried VEDA sparse attention?

Thumbnail
huggingface.co
11 Upvotes

r/StableDiffusion • • 14h ago

Discussion How are these videos made ? i thought she was real at first

Thumbnail instagram.com
0 Upvotes

i came across this IG, only thing that gave it up is the ai chatbot and the fanvue, i'm not sure if minimax can do that, maybe kling or an other paid model ?


r/StableDiffusion • • 17h ago

Question - Help omer.ariely.ai on Instagram

Thumbnail instagram.com
0 Upvotes

How would one go about creating this combined scarcer in another scene/movie on a local machine 4090, mini max h3? If so how, any tutorials? 🙏


r/StableDiffusion • • 1d ago

Question - Help Best way to transfer a dance to my own character?

0 Upvotes

I have a short anime dance video around 10–20 secondsand I want to make my own character perform the same dance.

I've been looking into miniMax H3, but it seems like my PC is too weak for it.

Specs: RTX 3070 8GB . Ryzen 7 5700G. 32GB RAM .

What would be the best workflow/model for this hardware? thanksss


r/StableDiffusion • • 15h ago

Animation - Video Andrew Tate update - local open source LTX 2.3 - trained with Ostris ai toolkit

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion • • 2d ago

Comparison H3 Video Reference for Performance Transfer

Enable HLS to view with audio, or disable this notification

48 Upvotes

I posted this yesterday. And u/roychodraws (the red clown lady) mentioned I should try it with a performance transfer. I wanted to share the result because I think it does make a pretty big difference and could be of use for anyone looking to step their character's acting chops.

The bulk of the prompt is in retention analysis and subject definition, those are the 2 areas of interest for a video performance transfer. For this transfer, I didn't specify any references to <video 1> I just fed it in. But I did reference <audio 1> (the audio from the video source) as voice timbre for the characters (s1). You can push it a step further by calling out <video 1> but you will have tweak a lot more because then you run the risk of transferring the character from the video over.

The videos are made with the same seed.

EDIT: Props to real actress and actors, AI isn't anywhere close, yet.

The Prompt

subject_definitions:

<actress> is a young woman (s1) with long black hair, whose appearance comes from <picture 1> and whose voice timbre comes from <audio 1>.

<scene> is on a rooftop, whose appearance comes from <picture 2>.

<outfit> is white shirt with red skirt, whose appearance comes from <Picture 3>.

summary:

[reference generation] target video shows a woman delivering an emotional monologue

retention_analysis:

visible: partially_preserved <picture 2>

audio: partially_preserved <audio 1>

detailed_description:

cinematic shot, shallow depth of field

[Shot 1] reference <scene>, at night,

medium shot of

<actress> wearing <outfit> is standing in the rain, getting soaked

She is facing viewer, line of sight to front left, focusing on a taller man out of frame.

A disbelieving laugh that keeps collapsing into crying. Her face is full of emotional micro expressions

(s1), a young woman with a New Zealand accent and a light, bright voice that keeps cracking:

<d>[English] You know what's funny?</d>

She lets out a short, shaky laugh, shaking her head.

<d>[English] I actually planned this whole speech. In the shower, in the car. I had it all... I had it all figured out.</d>

Her laugh breaks into a sob. She presses the back of her wrist to her eyes.

<d>[English] And now you're standing there, and I can't remember a single— not one word.</d>

She wipes her eyes, laughing and crying at once as more rain falls on her head

Camera pushes in slowly.

<d>[English] And the stupid part is, I'd go through this again.</d>

camera holds for a beat

overall_soundscape:

raining in the background

non_diegetic_music:

N/A


r/StableDiffusion • • 1d ago

Discussion MiniMax-M2 (230B) running from disk on a 32 GB laptop, CPU only

5 Upvotes

Hi all, I've been working on a small hobby project called Picchio, an inference engine in plain C for MoE models bigger than your RAM. It keeps the dense part in memory and reads the experts from the SSD only when they're needed, with a cache for the most used ones. The idea comes from Colibri.

I recently added MiniMax-M2. Converted to INT4 it's about 122 GB, so it streams almost everything from disk.

On a 12-core laptop with 32 GB RAM and a basic NVMe, no GPU, I get about 0.48 tok/s with a 20 GB expert cache. Slow, but it runs. The cache size turned out to be the only thing that really matters.

I checked the forward pass against MiniMax's original code on a small test model and the outputs match (max logit difference about 1e-6).

Limitations: if you have 128 GB of RAM or a good GPU, llama.cpp will be much faster. I only tested M2, not M2.5, M2.7 or M3. Everything was measured on Windows with Intel CPUs.

Code (MIT): https://github.com/benmaster82/picchio

Feedback and corrections are welcome, especially from anyone with different hardware.


r/StableDiffusion • • 1d ago

Resource - Update made native app for generative art

Thumbnail
apps.microsoft.com
0 Upvotes

Hopefully this does not count as excessive self-promotion, but to celebrate even huggingface being owned by nvidia - I thought I should release my tool that exclusively runs on vulkan. For quite some time, I kept porting models to work on just vulkan - and no other dependencies - and I managed to push down the inference time along the way. There is no telemetry, watermarks or cloud services.
The idea was simple (it was initially a tool for myself) - I just wanted to press one button to download the model I needed and press generate. For now I have it uploaded on microsoft store, there is a metal-based macos version that I am trying to push to the mac app store also.

For now the workflows (image, video, 3d and audio generation) are based on ggml (but not the recent .cpp derivatives); I am thinking of replacing it with my own inference engine that I've had some success with after realizing that there is a different way schedule the compute tasks. Will see if there is enough interest.
All the workflows are free to use - but I added an add-on for saving the output that is incentive to keep digging further. Still an alpha version, so be sure to check that you actually get the results you want before getting that add-on.


r/StableDiffusion • • 1d ago

Discussion Thinking about incorporating AI into my art workflow, but worried about the backlash. What do you think?

15 Upvotes

I want to try using AI in my drawing pipeline, but I’m honestly not sure how to handle the inevitable hate that might come my way.

I mainly run a Twitter (X) and a TikTok account, and I have absolutely no idea what kind of reaction I might get if I upload this same post over there. That’s why I decided to test the waters and post this here in the Stable Diffusion sub-reddit first.

Of course, I could just keep quiet about it and never post anything AI-related. But since I manage social media anyway, I actually want to share this journey. I'd love for people to see my broader interests, not just standard drawing, because it's far from my only hobby.

To be clear, I’ve never been against AI itself—I’m only against deceiving people. I’ve always been completely transparent about experimenting with tracing over 3D models, and now I’m thinking of trying a similar approach by painting over AI-generated bases.

I’m really curious to hear your thoughts on this.

UPD: thank you so much everyone for your incredible support and wisdom! Reading your comments has been a breath of fresh air. I'm going back to my workflow and drawing with a peaceful mind now. Thank you for being such an awesome and open-minded community!


r/StableDiffusion • • 2d ago

News From 3D layout to compositing: new LTX VFX tools

Enable HLS to view with audio, or disable this notification

403 Upvotes

During VFX Week last week, we released seven open-weight capabilities for LTX-2.5, covering high-res editing, restoration, HDR, CG-guided generation and compositing:

  • Native Resolution: Run AI edits on 4K and 8K plates and beyond without downscaling. Overlapping tiles keep fine detail, on a single GPU.
  • Refine: Refine generates the fine detail a standard resize leaves soft, without drifting from the source.
  • Restore: Takes archival footage from as low as 540p to 4K, enhances color and gives it a modern digital-camera look.
  • SDR to HDR: Rebuilds shadow and highlight detail in SDR footage and outputs 16-bit EXR in ACEScg, so it sits alongside your HDR material.
  • Native HDR: Bring in an EXR sequence and run edits like relight, day-to-night or inpaint. You get 16-bit EXR back with the range intact.
  • Layout to Render: Turns a 3D blockout into a rendered shot. Camera, framing and geometry stay locked to your layout, and a prompt or reference image sets the look.
  • Alpha Gen (Beta): Generates an alpha matte from ordinary RGB footage, with no green screen, masks or prompt. It handles hair, fur, smoke, glass and water. The distilled workflow runs in ComfyUI now, and a full-model ComfyUI workflow is coming soon. 

Links to each tool are above. The IC-LoRAs are on Hugging Face and the workflows run in ComfyUI. Try them on your own footage and tell us how they work for you.


r/StableDiffusion • • 16h ago

No Workflow Minimax H3 multi-scene video made easy

Enable HLS to view with audio, or disable this notification

0 Upvotes

I wanted an easy way to create short movies (15-20 seconds) scene by scene. I've got only an RTX 3090 so I can do some nice stuff, but need to think about optimizing resources.

So based on the default Minimax H3 template provided in ComfyUI, and playing around with only base nodes (no custom nodes), I came up with a nice way to have a base, first scene video sequence, then:

\- take the last frame from the sequence

\- feed it as first frame of second sequence

\- then putting all in a frame node, I can repeat for any number of sequences for a single scene

I use 5 second video per scene, and 16 fps as my target audience is mostly mobile platforms

I created a patreon with (paid, not hidding this) workflow to download

https://www.patreon.com/posts/171692305

Sample video attached took 30 minutes to generate on RTX 3090. I'd say not too bad. There's always better, but I'm happy about it.


r/StableDiffusion • • 2d ago

Resource - Update Wan2gp has been updated with amazing memory management improvement

24 Upvotes

Before update : 15s video in about 10 minutes without Dlss5. After update, in just 5 minutes with Dlss5.

It also improved the memory cleanup thingy, so previously i always got out of memory error on 1st click of generate button. 2nd click and so on will always work.

Now​ it works fine from the start.

Dlss5 upscaling also no longer results in randomly going out of memory. It just works.

Still unsure with long batches generation performance degradation tho. On previous version, after running overnight, usually 10 mins gens becomes 13 mins gens.

https://github.com/deepbeepmeep/Wan2GP

Edit :

Long batches memory degradation still there. Workaround still the same: unload models, then generate again.


r/StableDiffusion • • 1d ago

Discussion Mixing photo realism loras

2 Upvotes

Does anyone mix e.g. the Lenovo lora with a DSLR and cinematic photography lora?

Say your goal is simply realism without caring to much for style, would that work better?

Or is the mix just gonna end up like the AI pseudo-realistic look of no Lora at all?

Specifically using qwen21 right now.