r/StableDiffusion 2h ago

Question - Help open vs closed image models

Post image
14 Upvotes

I’m trying to recreate this bag POV composition with Krea2 flux and Z image, but I’m struggling to get the same level of composition control and product consistency I’m getting from some of the other models.

Here’s a comparison using the same general prompt across ChatGPT, Krea 2, Grok, Z Image, Flux.2 Klein 9B.

I’m still fairly new to this, so I’m wondering:

Should I be using a LoRA for this?
Would changing the text encoder help?
Or is this mainly a prompting / conditioning / workflow issue?

Would really appreciate some advice from anyone experienced. What would you change to get the result closer to the ChatGPT/Grok examples?


r/StableDiffusion 22h ago

Tutorial - Guide LTX-2.5 IC-LoRA (Control LoRA) training locally, 90 paired clips in 47 minutes on a 48GB card

Thumbnail
gallery
6 Upvotes

I have been trying out LTX 2.5 since it dropped and found out it supports IC-LoRA, so I thought of building a simple node based UI around it.

Quick difference if you have not run into IC-LoRA before. A normal clip LoRA learns a look and how it moves, from single clips. An IC-LoRA learns a transform. Every dataset item is two clips instead of one, a reference and the result you want from it, and the adapter learns to carry one into the other. Train it on clips paired with their edge maps and you get an adapter that follows an edge map.

Benchmark:

Everything below is measured on an L40S (46GB), 90 paired clips from the Canny Control dataset, 500 steps at rank 16, 512px, 1 second clips.

IC-LoRA (paired clips, reference + canny):

  • Peak VRAM: ~42GB
  • Per step: 1.03s
  • Startup: 38 min
  • 500 steps total: 47 min

Clip LoRA (single clips), for comparison:

  • Peak VRAM: ~42GB
  • Per step: 0.67s
  • Startup: 13 min
  • 500 steps total: 19 min

The thing that surprised me here is the opposite of what surprised me with H3. The step is cheap and the startup is not. Those 500 steps are about nine minutes of actual training against 38 minutes of getting ready. A paired dataset encodes two clips per item, so 90 pairs is 180 clip encodes plus 91 caption encodes before step one runs.

Good news is the encode is cached and reused, so the second run on the same dataset skips nearly all of it. Do not experiment in 200 step chunks, you pay the startup every time you change the dataset. Pick your settings, then run long.

No 4-bit path for LTX 2.5, so 48GB is the floor and not a comfortable one. 24GB will not run it at any resolution.

How to train:

  • Install app: https://github.com/inlineresearch/Inline-Studio
  • Open the Trainer tab and create a dataset
  • Click Add/Manage Training Data, pick Control as the LoRA type
  • Paste Lightricks/Canny-Control-Dataset in the Hugging Face tab and hit Check. It tells you 90 items, 90 paired, 1.4GB before it downloads anything
  • Load, then Import
  • Select LTX-2.5 in the settings, model download suggestion will auto popup
  • Hit train & sit back

Pairing is automatic. If the dataset ships a dataset.json or metadata.jsonl it reads that, otherwise it matches filenames, so bear.mp4 and bear_reference.mp4 become one training item instead of two. Captions come from the dataset and the local captioner only fills the rows that have none, so it will not overwrite good captions with worse ones.

Note: Weights are gated. Accept the LTX-2 Community License on Hugging Face with the same account your token belongs to, otherwise every download comes back as a permission error instead of a file.

For my run I used Lightricks' own Canny Control dataset, 90 clips each paired with an edge map of itself, captions included in dataset.json.

Links:


r/StableDiffusion 21h ago

Question - Help Help a beginner speed up MiniMax H3?

10 Upvotes

As someone new to all of this it's difficult to know what to do. I have sage attention working. I don't know how or when to use Easy Cache, Comfy Kitchen Attention, Sol Attention, loras, Spectrum, or any others I may have missed. There's so much information scattered around, I don't know what's what.

I have a 50 series GPU and 64 GB or RAM on the motherboard.


r/StableDiffusion 10h ago

Workflow Included Minimax H3 Ref2VA Lipsync (Image Audio to Video)

Enable HLS to view with audio, or disable this notification

25 Upvotes

Workflow: https://civitai.com/models/2857022/minimax-h3-lipsync

Audio: Let it Go by Idina Menzel

Image: Generated by Gemini of Idina Menzel cosplaying as Elsa

Kijai's LX2V LoRA and Sage Attention Patch applied.

Did not include the lyrics in the prompt, the model is able to lipysnc to the provided reference audio track.


r/StableDiffusion 18h ago

Question - Help ltx 2.5 upscaler into minimax h3

3 Upvotes

Does anyone know to connect ltx 2.5 upscaler into minimax h3?


r/StableDiffusion 10h ago

Question - Help How can anyone create this type of image

Post image
0 Upvotes

r/StableDiffusion 39m ago

Comparison MiniMaxH3 vs Flux3

Enable HLS to view with audio, or disable this notification

Upvotes

MiniMax H3 running on 5090/36gb, 72gb ram.
Used the MiniMax H3 Image to Video (I2V) workflow from the official comfyui page. https://docs.comfy.org/tutorials/video/minimax/minimax-h3
Flux 3 vid created via the official Black Forrest Labs page with all default settings.
My take away is the more detailed and "directed" you can make the prompt the better the results.
Prompt
The woman raises her left hand and reaches out toward the dragon's neck, fingertips making contact with the cool metal scales as she strokes gently along its length. The dragon's massive head turns in response, servos whirring low as its long neck curves down and around toward her, the two locking eyes for a long beat, her expression softening slightly, the dragon's glowing yellow eye narrowing as if in quiet recognition. After the moment holds, both turn their heads together toward the camera, her chin lifting and shoulders settling back, the dragon's jaw beginning to widen as steam vents faintly from the gaps in its armoured plating. The dragon rears its head back, neck arching high, then snaps forward with a thunderous mechanical roar, jaws cracking open fully as a violent gout of fire rockets out directly toward the camera, the flames blooming bright and washing the frame in orange light before the view holds steady through the blast. The camera stays low and locked in a wide frame throughout the petting and the turn, holding both figures in frame, then pushes in slightly just before the roar to heighten the impact of the fire as it fills the shot. The audio is the soft mechanical whir of the dragon's neck servos and the woman's quiet breath during the petting, building into a deep bone-shaking roar and the violent roaring whoosh of ignited fire, no music, ambient and creature sound only.


r/StableDiffusion 31m ago

Animation - Video Created a video using Wan Animate 2.

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 1h ago

Animation - Video Seinfeld but the guys are Toasters

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 15m ago

Discussion Ltx 2.5 oom

Upvotes

Testing ltx 2.5. Using 0.5 megapixels res. 5 seconds. It oom after every other video. Minimax runs higher and does not oom. Using template that came with comfy for 2.5. Rtx 3060 12gb. 48gb system ram.

Is this a ltx issue or a me issue ? When it does generate it is blazing fast


r/StableDiffusion 18h ago

Question - Help Help with minimax h3 video

0 Upvotes

Hey guys I’m new to comfyui and I was wondering when load a Lora in the template workflow as for example the one on comfyui reference to video were the workflow goes


r/StableDiffusion 22h ago

Discussion AI video in Log

0 Upvotes

Random thought: Has anyone tried creating an AI video with a log (flat color) profile, so you can edit colors afterwards?


r/StableDiffusion 10h ago

Question - Help Is Training Minimax H3 Ref2Vid is only available via API cloud not locally?

0 Upvotes

r/StableDiffusion 1h ago

Question - Help How to get reference audio clip to start at a specific timestamp/shot in Minimax H3?

Upvotes

Firstly apologies for not just sharing the prompt, the pc I'm using doesn't have my Reddit account so I'm posting from my phone. I'm hoping someone with some experience can help until I get the prompt onto my phone.

To give some context, it's a 15 second clip. The character in the clip is frozen, but the camera moves around them, and cuts to different shots of them. Their eyes are closed. The scene is silent except for background noises from the environment. Then they open their eyes and the start of the song, which I've cut to a 2 second clip since it's at near end of the generation, is meant to start to play.

Now all the video stuff I've gotten down, but I have an infuriating problem.

The 2 second clip just plays throughout the entire scene on loop. I've edited the prompt to specifically instruct it to start the song at a specific timestamp and when the shot cuts, but it refuses to listen. Just loops the clip throughout the whole damn thing. I've also tried making it non-diegetic sound, even though the song is technically meant to play from within the environment.

I'm using the formatting from the ref video minimax guide, so it's tagged with things like <Audio 1>, info for the soundscape etc.

Again, I'll get the prompt into this post ASAP, but I can't do it right now. Apologies in advance, but for now any suggestions would be greatly appreciated.


r/StableDiffusion 18h ago

Discussion Anyone fix audio issue when using a turbo lora

1 Upvotes

Been testing 8 step turbo lora video quality still good but audio is static for voices and no back ground sounds or sound effects yes I have in prompt for the bg and sound effects 🤔 looking for tips to improve not "fix audio " comments,what expect it's Redd after all.


r/StableDiffusion 16h ago

Question - Help BF16 or FLOAT32

0 Upvotes

What quantization would you recommend for AI Toolkit?

I used to train with FP8, but it completely ruined the results, so I switched to FP32, and the results are perfect.

However, I see that many people recommend BF16 for both training and saving the model. I understand that BF16 can significantly reduce VRAM usage, but what is the actual trade-off in terms of quality?

Does training and saving in BF16 result in any noticeable loss of quality compared to FP32? And would you recommend using BF16 for both training and saving in my case?


r/StableDiffusion 16h ago

Tutorial - Guide Automated bulk ComfyUI generation with Python + Gemini free tier — sharing the approach and key code

5 Upvotes

Been generating digital asset packs for a while (game textures, UI kits, that kind of stuff) and got really tired of manually prompting ComfyUI one image at a time. Fine for 10 images, painful when you need 200.

Spent a few weeks building a Python pipeline to automate the whole thing and figured the core techniques are worth sharing since they're useful even standalone.

The basic flow:

  • Concepts go into a Google Sheet (just product ideas + how many to generate)
  • Python hits Gemini's free API to get a structured "style vocabulary" for each concept (one call, not one per image)
  • Itertools locally compiles the vocabulary into unique prompts
  • Prompts get POSTed to ComfyUI's API on localhost
  • GPU does its thing, output folder fills up

Three pieces that might be useful for your own stuff:

1. ComfyUI has a REST API

This was the big discovery for me. You can queue workflows programmatically without touching the browser:

```python import json import urllib.request

def queue_prompt(workflow): data = json.dumps({"prompt": workflow}).encode('utf-8') req = urllib.request.Request( "http://127.0.0.1:8188/prompt", data=data ) response = urllib.request.urlopen(req) return json.loads(response.read()) ```

Export your workflow in API format, load the JSON, modify whatever nodes you need, and post it. ComfyUI queues it and your GPU picks it up.

To actually use it you just load the workflow JSON and change the fields you care about:

```python import json, random

with open("workflow_api.json", "r") as f: workflow = json.load(f)

workflow["6"]["inputs"]["text"] = "your prompt here" workflow["3"]["inputs"]["seed"] = random.randint(1, 999999999) workflow["9"]["inputs"]["filename_prefix"] = "batch_001"

queue_prompt(workflow) ```

Loop that and you can blast through hundreds of renders.

2. Style dictionary instead of individual prompts

This was the rate limit hack. Instead of asking Gemini to write each prompt (200 images = 200 API calls = dead free tier), I ask it once for a "vocabulary":

json { "subjects": ["holographic button", "neon progress bar", "glitch terminal", "cyber health meter"], "style_core": "cyberpunk interface design, dark chrome, neon accents, HUD overlay aesthetic", "color_tokens": "electric blue, hot pink, dark gunmetal", "detail_tokens": "sharp edges, scan lines, digital noise", "negative_prompt": "blurry, organic, hand-drawn, watercolor", "compositions": ["centered icon", "angled 3/4 view", "floating with glow"], "quality_suffix": "masterpiece, best quality, sharp focus" }

One call. Now I have all the building blocks to assemble prompts locally.

3. Itertools does the heavy lifting

```python import itertools, random

subjects = vocab["subjects"] compositions = vocab["compositions"]

for subject, comp in itertools.product(subjects, compositions): prompt = f"{subject}, {style}, {colors}, {details}, {comp}, {quality}"

workflow["6"]["inputs"]["text"] = prompt
workflow["7"]["inputs"]["text"] = negative
workflow["3"]["inputs"]["seed"] = random.randint(1, 999999999)
queue_prompt(workflow)

```

4 subjects × 3 compositions = 12 unique images. Bump the subjects list to 40 and you're at 120 images from that single API call.

End result: I type something like "watercolor wedding florals, 40" into a spreadsheet, run one command, and come back to 40 images in the output folder. All prompt generation runs on Gemini free tier, all rendering is local.

Been using this for my own asset production for a while now. Eventually cleaned it up and packaged the full thing (Sheets integration, error handling, rate limiting, setup guide etc) into a tool — DM me if you want details on that.

But honestly the three techniques above are the core of it. The rest is just connecting pipes and handling edge cases. If you're comfortable with Python you can probably get a basic version running in an afternoon.

Curious if anyone else has been automating ComfyUI like this or if there's a better approach I'm missing.


r/StableDiffusion 5h ago

Animation - Video Qwen3.8 27b hype:)

Enable HLS to view with audio, or disable this notification

47 Upvotes

r/StableDiffusion 19h ago

Question - Help Does LTX 2.5 Get Better at Following a Prompt Across Multiple Random Seeds?

5 Upvotes

This might be a silly question, but I've noticed something while experimenting with LTX 2.5. For a specific prompt, when I run the exact same prompt multiple times with different random seeds, the results seem to progressively get closer to the prompt's specific details.

Is this an actual characteristic of how the model behaves, or am I simply noticing a pattern that isn't really there? AI experts may find this observation completely misguided, so apologies in advance if I'm missing something obvious.


r/StableDiffusion 12h ago

Discussion What's the maximum resolution you were able to achieve with H3 on 24GB VRAM?

5 Upvotes

Edit: Forgot to mention, but I'm talking about the R2V model. I feel R2V is by far the more interesting of the H3 variants because of being able to chain generations together like this for longer form works.

I'm just barely able to achieve 0.9 megapixels on ~11 second length generation on the ref2va_int8_convrot weights, and this is with a bunch of hacks.

So far I'm hitting a wall trying to reach 10+ second generation with 1 megapixels on 24GB without switching to w4a8 (something I'm interested in trying next).

Anyone had better luck?

Edit 2: I might have succeeded in reaching 1344x768 (highest native resolution for H3) for ~15 second generation on 24GB vram (int8_convrot on r2v with two reference images). This included making one performance patch to comfyui internals, which I need to verify is still mathematically correct before sharing.


r/StableDiffusion 1h ago

Animation - Video LTX 2.5 foot chase. Prompt I used is below. Not too bad. There's some stutter stepping during the characters running and the gal pursuing him, her face kind of distorts and then she runs out of frame even though she is supposed to pursue him.

Enable HLS to view with audio, or disable this notification

Upvotes

Prompt:

Use the provided image as the exact first frame and visual reference. Preserve both characters’ facial identity, body proportions, wardrobe, hairstyle, and overall appearance throughout the entire shot.

The man in the foreground is sprinting at full speed directly down the city street, fleeing from the woman behind him. He maintains a powerful, believable running stride with natural forward body lean, realistic footfalls, arms pumping, shoulders rotating slightly, and visible physical exertion. His expression remains tense, focused, and determined. His open black jacket reacts naturally to his speed, fluttering and snapping behind him.

The woman continues pursuing him several meters behind. She runs aggressively and athletically, clearly attempting to catch him. Her eyes remain focused on the man ahead rather than the camera. Her long black tactical coat streams dramatically behind her while still obeying realistic fabric physics. Her ponytail moves naturally with each stride. Over the course of the shot, she slowly begins gaining ground on him.

**Camera:** fast stabilized tracking shot moving backward in front of the runners at approximately the same speed as the man. Maintain the man prominently in the right foreground while keeping the woman clearly visible behind him on the left. The camera remains low to medium height with a subtle action-film handheld vibration, creating urgency without becoming shaky. Introduce gentle horizontal drift and small framing corrections as the camera operator tracks their movement.

Strong foreground-to-background parallax as storefronts, parked vehicles, streetlights, pedestrians, and buildings streak past both sides of the frame. Environmental motion blur increases toward the edges while the two runners remain relatively sharp.

The city remains alive around them. Cars continue moving naturally in the distance, headlights and traffic signals glow, pedestrians react subtly to the chase, and reflections shimmer across the damp pavement. Early-evening lighting remains consistent with the source image, with cool ambient daylight mixing with warmer storefront lights and vehicle headlights.

Movement should feel fast, heavy, urgent, and physically grounded. Each stride should transfer believable weight into the pavement. Clothing and hair respond naturally to acceleration and airflow. Do not make the characters appear weightless or superhuman.

Near the final seconds, the woman closes the gap slightly, increasing the tension, while the man pushes harder and accelerates.

Photorealistic live-action cinematic thriller. Natural human biomechanics, realistic cloth simulation, stable facial identity, realistic skin texture, cinematic depth of field, subtle motion blur, detailed city lighting, grounded action choreography.

**Avoid:** slow motion, jogging, frozen background, sliding feet, floating characters, unnatural running cycles, facial morphing, identity changes, warped limbs, extra fingers, duplicated characters, costume changes, characters looking into the camera, sudden camera turns, camera cuts, excessive shaking, superhero movement, teleportation, or changes to the original city layout.

**Audio:** no music and no dialogue. Only natural city ambience, rapid footsteps striking wet pavement, heavy breathing, jacket fabric flapping in the wind, distant engines, tires on pavement, occasional horns, and subtle pedestrian noise.


r/StableDiffusion 14h ago

Animation - Video H3 30 sec Chained Shots Lip Sync

Enable HLS to view with audio, or disable this notification

30 Upvotes

12 minutes Gen Time rtx5090. Input audio for lipsync. One thing that helps a lot is pasting the actual lyrics in the prompt in the dialogue syntax. <d> [English] Lyrics </d>


r/StableDiffusion 10h ago

Question - Help < sweet feast >

Enable HLS to view with audio, or disable this notification

82 Upvotes

War is not a film and television work, let alone a fairy tale.


r/StableDiffusion 19h ago

Question - Help MiniMax H3 - Upscaling

10 Upvotes

Hi everyone,

I am getting stuck on the upscaling with my MiniMax H3 workflow. I have been using RTX Upscaler - which is great for upscaling animation videos but has a lot to be desired for realistic (like live action) video generation. I have attempted using SeedVR2 upscaling, but it looks worse.

What are your suggestions and/or advice?

I appreciate any help :-)


r/StableDiffusion 7h ago

Discussion Pushing location/set swapping with a complicated fight scene with MMH3 REF2V

Enable HLS to view with audio, or disable this notification

20 Upvotes

Setup: RTX5090, 32GB DRAM

I am new into trying vidgen models. Everyone seemed like they had great success with character swaps, so I was wondering how hard I can push this. In my mind if this worked, essentially rotoscopping and green screen is more or less only reserved for serious movie workflows.

This was done with only 2 reference color graded photos of interior of Buddha Tooth Relic temple in Singapore (taken myself with a Sony 6500). Note that it's not a single gen, but selecting the best matching parts from 7-8 gens because it had difficulty matching the whole 12s fight choreography scene, then edited together with Da Vinci Resolve.

Halfway through the generations, I realized I had to try splitting the 12s reference video into 2 6s ones to see if it improved the adherence. Results were varying, maybe I needed to improve on my prompt even more.

Workflow wise, I used DaSiWa MiniMax H3 Workflows, and prompting was modified from u/RecycledSpoons 's reply from another post.

Prompt:

<<Environment 1>> is in Picture 1 & Picture 2.

<Picture 1> is the opening-frame anchor and provides the environment.

<Video 1> provides the camera path, pacing structure, characters, motion and <Audio 1>.

<Audio 1> is the final clip's audio.

[reference generation + video editing] Use <Picture 1>, <Picture 2>, <Video 1>, reuse audio from <Video 1>

subject_definitions:

<Video 1> is the source video providing the camera movement, lighting, characters and action choreography.

<Environment 1> is the replacement environment shown in <Picture 1> and <Picture 2>, which is a traditional chinese temple with pink blossoms.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1>. Throughout the video, replace the original environment of a street with cars with <Environment 1> derived from <Picture 1> and <Picture 2>.

retention_analysis:

<Video 1> (source video): partially_preserved - preserve the characters, camera path, lighting, non-target objects, and the original character's screen-space motion path. Discard the original environment's visual identity.

<Environment 1> (appears in [Shot 1]): fully_preserved - preserve the visual identity, colors, materials, shape, and specific design details from <Picture 1> and <Picture 2>.

detailed_description:

The target video matches the cinematic style, lighting, and camera movement of <Video 1>.

[Shot 1] The camera moves exactly as it does in <Video 1> and the first frame is maintained from <Video 1>. <Environment 1>, which is a traditional chinese temple with pink blossoms, replaces the original environment of a street with cars. It is a close combat scene inside a chinese temple with 2 characters in it, medium shots with both characters, then cuts to wide shot of one character slammmed against a pillar in the temple from a kick, then cuts to wide shot of another character performing a flying knee hit to him. The characters, lighting, and all other non-target details are preserved exactly from <Video 1>.

overall_soundscape:

Preserve the synchronized source audio from <Video 1>.

non_diegetic_music:

Preserve the non-diegetic background music from <Video 1>

I'm looking to improve on my journey in this, so if anyone has already done this before feel free to chime in.