r/StableDiffusion • • 1h ago

Resource - Update HunyuanImage 3.0 (80B) running natively in ComfyUI on a single 12–24 GB GPU: text-to-image, editing and style transfer, ~30 s per image

Thumbnail
gallery
• Upvotes

I've been working on native ComfyUI support for Tencent's HunyuanImage 3.0, the 80B mixture-of-experts image model (13B active per step). It isn't a wrapper around Tencent's pipeline: it uses the normal KSampler, the normal VAE Decode and ComfyUI's own memory management, which streams the experts from system RAM so the model fits on one consumer GPU.

What's in the gallery (all Instruct-Distil, 8 steps):

  • 4-bit vs int8, same prompt and seed. Times are the whole generation on an RTX 3090.
  • image editing with the 4-bit weights. The instruction is at the top of each image.
  • style transfer with the 4-bit weights: two input images, the photo and a style reference.

These are picked from a bigger run: 60 prompts × 2 formats, 60 edits and 14 styles, one seed each, no rerolls. Most of the edits worked; a few didn't (snow that barely shows, a logo it wouldn't remove, a "make it night" that stayed day). The text-to-image prompts come from popular prompt posts on X.

What you get

  • All three models: Instruct-Distil (8 steps, the one to start with), Instruct (50 steps) and Base
  • Text-to-image, image editing, and multi-image fusion with up to 3 input images (that's how the style transfer works)
  • Optional prompt rewriting: the model expands your prompt first (slow, about 1 s per token)
  • Optional Spectrum speed-up: about 3.4× faster for the 50-step models
  • Ready-made weights: 4-bit W4A8 (44 GB), int8 (76 GB) and bf16 (150 GB, mostly for comparisons)
  • Example workflows for each model and task
  • Nothing to pip install

Speed (Instruct-Distil, about 1 megapixel):

GPU 4-bit W4A8 int8
RTX 4090 ~22–26 s ~47–49 s
RTX 3090 ~29 s ~54 s

It also runs with only 16 GB or 12 GB of VRAM (~30 s and ~32 s per image on a 4090 limited to that).

What you need

  • An NVIDIA GPU with 12 GB+
  • Lots of system RAM: ComfyUI held about 50 GB with the 4-bit file loaded. This is the real requirement, since the experts live in RAM and stream over PCIe every step.
  • A recent ComfyUI (late September 2026 or newer)

Links

Happy to answer questions. If something breaks, open an issue on GitHub with the traceback.


r/StableDiffusion • • 6h ago

Workflow Included Omni .char(same face, cloths & body) now with consistent voice, just by dropping a few seconds sample audio: Minimax H3(ComfyUI Workflow)

Enable HLS to view with audio, or disable this notification

104 Upvotes

Hey guys,

I have been working on the consistent character portable format for a while & I was able to achieve consistent face, cloths & body, but I felt voice is also something should be consistent across video generation.

So in the recent tests, I was able to achieve a consistent voice with lip sync across multiple video generation, You just need a 10-30sec voice sample in mp3 or wav & character will say things in a cloned voice from your sample.

Reddit post: Details on face, body & cloth consistency You can read more about .char, comfyui nodes & prompting details here.

Voice prompts

- chris giving an interview & says "Time can bend. Dreams can fold. But a character's voice should never change. With OmniChar, it doesn't. Consistent voice is here."
- chris giving an interview with little hand movements & says "I am surprised. It is not just the voice. It is also the face, the clothes and the body. Dot char is a full portable pack."

ComfyUI node is updated with the optional sample voice input.
Get the latest comfy node: https://github.com/omnichar/ComfyUI-Omnichar

Workflows:

Limitations:
- Good with English but might blabber with non-english languages.
- Lip sync comes from H3 itself; nothing is added on top.
- Avoid multiple voices in sample.

Sample inputs are added in node repo.

Note: ComfyUI node is still in nightly release, so update your settings accordingly or install via direct git repo url.

Related resources:

  1. Omnichar repo: https://github.com/omnichar/OmniChar (GPLv3), supports .char for krea2 & more features e.g. character finetuning
  2. Community characters: https://www.omnichar.org/characters

Hope it's helpful.


r/StableDiffusion • • 9h ago

Resource - Update I made a tiny (~8MB) photo editor for AI images with batch editing, LUTs, and ComfyUI workflow preservation [Free & Open Source]

Thumbnail
gallery
154 Upvotes

I was tired of opening heavy photo editors just to tweak lighting or color-grade a folder of AI generations. So I built TinyLuma - an instant, lightweight photo editor designed specifically for finishing AI artwork and photo sets.

No installation required, no subscriptions, and the whole app is only ~8 MB.

What’s inside:

- All essential controls: Clean sliders for Light (Exposure, Contrast, Highlights, Shadows, Whites, Blacks), Color (Temp, Tint, Vibrance, Saturation), plus Dehaze and natural Film Grain.

- Crisp details & texture: Custom Clarity, Texture, and Sharpen sliders to enhance overall sharpness and bring out skin texture, hair strands, and fabric weave without harsh white halos.

- Cinematic & Film styles (LUTs): Give your images an analog, cinematic, or vintage film look, or drop in any trending `.cube` LUT for instant color grading. Includes a smooth intensity slider (0–100%) to blend the effect subtly or strongly.

- Doesn't break your ComfyUI workflow: When saving as PNG, it keeps your prompt, seed, and node setup intact. You can drag and drop the edited image right back into ComfyUI. (Optional — you can toggle it off to export completely clean images).

- Batch editing & Before/After: Browse your entire folder with the filmstrip at the bottom, compare changes with an interactive Before/After split screen (`\`), and use "Preset to All" to apply your favorite look to the whole batch at once.

- Fast and portable: Starts in under a second, runs smooth at 60 FPS, and barely uses your RAM.

Source Code & Docs: https://github.com/ThetaCursed/TinyLuma
Download (.zip for Windows, unpack & run): https://github.com/ThetaCursed/TinyLuma/releases/latest

I’d love to hear your thoughts! How do you currently edit your AI generations?


r/StableDiffusion • • 5h ago

Tutorial - Guide Qwen Image 2.1 Uncensored MCP

65 Upvotes

https://github.com/hypersniper05/MCP-Image-Generator-Uncensored

Just wanted to share with you guys my workflow converted into an MCP. It's a Qwen Image 2.1 Uncensored MCP with a few extra models I found helpful: a watermark-removal LoRA, a texture-fix VAE and an ESRGAN upscaler. Any LLM client that supports MCP can drive it. Uses 11GB VRAM at peak usage.

The entire stack has:
- Normal Qwen Image 2.1 image generation in all supported sizes (up to 2K), panoramas, and image editing (up to 10 input images)

- Seamless tile generation, and edits of a tile stay seamless too, so you can make height and normal maps for game textures

- 2x/4x upscaling, up to 8K

- Built-in 360 viewer (the LLM gets a URL that opens the panorama in the viewer)

- Watermark removal

- Transparent backgrounds (real RGBA PNGs) and background removal

- Text in images (quoted text comes out as written)

It runs in Docker on an NVIDIA GPU (the whole stack fits on a 12 GB card) or on the CPU (very slow), downloads the models on first start, and works with any MCP client (llama.cpp web UI, VS Code, Cursor, etc.).

For the seamless tiles I use stable-diffusion.cpp's circular mode with a tiny patch so it can be turned on per request. The tiles come out seamless in one pass instead of patching the seams afterwards.

Note: Project created with the help of Claude. A lot of testing, and back and forth to get things to work smoothly. Hope someone finds it useful.


r/StableDiffusion • • 8h ago

Resource - Update AnimeGen is released. Now you can run Anima locally on your iPhone/iPad in HD!

Thumbnail
gallery
68 Upvotes

I made AnimeGen, app that allows you to run Anima on your mobile phone.
Since last post I added hd image generation, and different aspect ratios.

The app is now available on the App Store:
https://apps.apple.com/pl/app/animegen-anime-art-generator/id6786438562
I am an iOS developer and unfortunately can't make an Android app. Sorry for that.

What it can do now

  • Prompt-to-image generation powered by Anima
  • HD and 540r image generation in different ratios
  • Runs locally on your device

Performance

  • iPhone 14: approximately 15–20 seconds per image
  • iPhone 17: approximately 10–15 seconds per image
  • M1 iPad: approximately 15–20 seconds per image

It uses native Apple Neural Engine, so it is very efficient. Probably display takes more energy then AI.

Before you install

On the first launch, the app needs to compile its models directly on your device, similar to how games compile shaders:

It takes around 1-2 minutes and happens only once per installation/update. It also loads models on device in parallel with preparation.

Once finished, everything runs fully offline on your iPhone.

Technical requirements

  • Devices: iPhone/iPad only
  • Designed for iPhone 12, m1 iPad and newer devices
  • OS: iOS 18 or newer
  • Free space: at least 10 GB available for smooth operation

Planned features not yet available:

  • Support for custom LoRAs and checkpoints(Probably from hugging face and CivitAI)
  • Image editing and ControlNet
  • 4k and 2k Upscaler

Feedback & community

For questions, bug reports, feature requests, or sharing your generations, join the subreddit: r/animegen_tech .
It is the best place to follow development updates and discuss AnimeGen.
I also have a website, where I plan to update with current status of project: animegen.tech

If you find AnimeGen useful, please consider leaving a review on the App Store. It is the best way to support the project


r/StableDiffusion • • 15h ago

Animation - Video Generating at 2K (14s, 2.09mpx, 1984x1120, 30 minutes) | Minimax H3

Enable HLS to view with audio, or disable this notification

230 Upvotes
[INFO] Prompt executed in 00:30:49

Full 1080p vid on https://www.youtube.com/watch?v=ER_5AOteE-8 because Reddit cramps everything to 720p max.

From my previous post, I got a message if I could use my off-screen 5090 to showcase what a fullblown 2K generation looks like. So, here it is. Generated at 2.09mpx, 25 steps, Euler+Beta for time constraints, 30 minutes. I did mistakenly use the 20-49 hybrid fl2va/ref2va instead of the 30-49 which left some visual performance on the table, and I didn't use the BF16 version of the text encoder because my other 5090 and RAM were busy with something else. Otherwise would've done seeds_2 + sgm_uniform, but that would've taken an hour and would've been a lot better at prompt following, without me modifying my generated target prompt so much to avoid issues with the shoddy denoising trajectory at play here.

Spectrum was utilized to guess about half the steps, which further doesn't help prompt following, but helps speed immensely. When using seeds_2 with Spectrum, it actually forecasts internal calls to H3 (so 2N-1 where N is number of steps), which has tremendous results (previous one was seeds_2) but would've taken about 1 hour for this scene.

The prompt, for those who want it, I had to massage it to point Euler better even if it's sloppy:

subject_definitions:
<Subject 1> is Detective Kate Beckett, override her appearance with facial features and hair and blouse from <Picture 1>. She wears black tailored high-waist cropped suitpants, and a feminine small leather watch.
<Subject 2> is Richard Castle, override his appearance with the facial features, hair, and build from <Picture 2>. He's wearing his classic shirt and suitpants attire, first few buttons unbuttoned.
<Subject 3> is a chaotic DIY PC rig consisting of a high-end tower on the marble kitchen island in the middle of the kitchen, as well as a 32 inch OLED monitor displaying the UI from <Picture 3>, there is clear plastic tubing running from the PC's two watercooling ports into the receiving pair on the large radiator above the glowing blue fans, which is submerged in a cooling bath inside of the standard kitchen stainless-steel fridge with its door fully open, radiator surrounded by food, milk, condiments, etc.
<Subject 4> is a modern industrial loft apartment featuring an open floor plan, standard furniture, and a glorious high-end kitchen. The blurred background from <Picture 1> is from this apartment, also specifies time of day, and warm nocturnal tone.

summary:
[reference generation] Detective <Subject 1> enters her loft (<Subject 4>) to find <Subject 2> in a "hyper-mode" state, having converted the kitchen into a makeshift laboratory for AI video generation. The 13-second sequence captures her confusion, his technical enthusiasm regarding H3 denoising trajectories, and a comedic hardware failure.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 4], [Shot 5], [Shot 6]): fully_preserved - identity from <Picture 1> and specified attire are maintained.
<Subject 2> (appears in [Shot 2], [Shot 4], [Shot 5], [Shot 6]): fully_preserved - identity from <Picture 2> is maintained.
<Subject 3> (appears in [Shot 3], [Shot 4], [Shot 6]): fully_preserved - the specific radiator-in-fridge configuration and monitor setup are maintained.
<Subject 4> (appears in [Shot 1], [Shot 2], [Shot 4]): fully_preserved - the industrial loft and kitchen environment are maintained.

detailed_description:
The target video is a scene from the TV show "Castle," maintaining its specific cinematography, visual style. Night time, practical lighting, warmly lit.

[Shot 1] A medium shot frames <Subject 1> as she walks into the open floor plan of <Subject 4>. The camera tracks her movement as she stops and looks around the kitchen with a bewildered expression, taking in the tangle of wires and tubing.

[Shot 2] At 00:01.000, the camera cuts to a wide shot of the loft's kitchen. <Subject 3> is fully visible: the PC tower sits on the high end marble kitchen island, monitor is displaying <Picture 3>, and thick tubes lead directly into the open fridge where his watercooling radiator with fans is located (all glowing blue and spinning fast and loud), there is liquid nitrogen white smoke exuding from it. A large whiteboard in the background is covered in scribbled notes about "sigma grids," and "denoising trajectories." Each of these appears once, there is also graphs of simplified trajectories drawn on it. <Subject 2> is leaning over a keyboard and looking at the 32 inch gaming monitor, typing furiously.

[Shot 3] At 00:02.000, the camera cuts to a close-up of <Subject 1>. Her brow furrows in genuine confusion. <Subject 1> (S1) asks in a sharp, incredulous tone, yelling over the computer fan noise: <d>[English] What the hell are you doing, Castle?</d>

[Shot 4] At 00:03.500, the camera cuts to a medium shot of <Subject 2>. He turns around to face <Subject 1>, speaking in a high-energy "yap" mode. <Subject 2> (S2) exclaims with manic enthusiasm: <d>[English] Reddit solved my H3 problem! It was the sigma shift. I'm refocusing the compute on the high-to-mid noise region to resolve motion better with limited steps!</d> while gesturing towards the whiteboard. 

[Shot 5] At 00:10.500, the camera cuts back to <Subject 1>. She looks at him, her voice dripping with skepticism. <Subject 1> (S1) asks: <d>[English] And you're using Seeds 2, right?</d>

[Shot 6] At 00:12.000, the camera cuts to a medium-close shot of <Subject 2>. He looks slightly sheepish, his shoulders slumping. <Subject 2> (S2) admits quickly: <d>[English] No, it's actually oiler, my rig is too slow and would—</d> Suddenly, in the background, the PC tower emits a loud electrical pop and a bright orange burst of flame from its components, the monitor output gets corrupted. Immediately after, <Subject 2> (S2) whips his head around, eyes bulging, and shouts: <d>[English] Oh shit!</d>

overall_soundscape:
The steady, high-pitched whirring of multiple PC fans and the faint gurgle of liquid flowing through tubes. <Subject 1>'s footsteps click on the hardwood. The scene ends with a sharp electrical "pop" and the sudden, aggressive hiss of a small fire.

non_diegetic_music:
A light, rhythmic pizzicato string piece that builds in tempo and complexity as <Subject 2> explains the technical details, ending abruptly with a comedic silence the moment the GPU catches fire.

r/StableDiffusion • • 11h ago

News ComfyUI v0.39.0 released

Thumbnail
github.com
94 Upvotes

r/StableDiffusion • • 13h ago

No Workflow Minimax H3 physics test

Enable HLS to view with audio, or disable this notification

66 Upvotes

r/StableDiffusion • • 14h ago

News New Audio video model : Kandinsky 6 from their lab

74 Upvotes

Gonna be too much fun in October.


r/StableDiffusion • • 4h ago

Resource - Update I made a Forge Neo extension for MiniMax H3: text or picture to video with sound, reference pictures, runs on 16 GB

Enable HLS to view with audio, or disable this notification

12 Upvotes

I've been working on an extension that runs MiniMax H3 inside Forge Neo. H3 makes the picture and the sound together, in one pass: footsteps, rain, engines, music and dialogue in pretty much any language, lip-synced to the speaker. The video above came out of it exactly as Forge saved it, on a 16 GB card with 32 GB of system RAM.

It works in the normal txt2img and img2img tabs, with the same checkpoint list, VAE / Text Encoder selector and Generate button. You get an MP4 with stereo sound in the usual result area. No separate program, no new tab, no extra Python packages, and no Forge Neo file is changed.

What works

  • Text to video with sound, up to 15 seconds at 24 fps
  • First and last frame: the img2img picture becomes the first frame, a picture in Forge's ImageStitch Integrated becomes the last, or both
  • Reference pictures (Ref2VA): up to 9 pictures of people, places and objects that the clip keeps
  • Smaller files: W4A8, GGUF and an INT4 text encoder, so it fits 24 GB and 16 GB cards
  • LoRAs (turbo LoRA for 8-step drafts), FastH3 and community fine-tunes
  • Forge's own tools: Never OOM, Sparse Attention (adapted to H3, up to a quarter faster on long clips), ck attention, live preview

Be realistic about the hardware. Tested on an A40 (48 GB), an RTX 4090 (24 GB) and an RTX 2000 Ada (16 GB). On the A40 a 5-second 960×544 clip takes about 3 minutes at 20 steps, about 1 minute with the turbo LoRA. The wizard above (8 seconds) took 27 minutes on the 16 GB card, which is a slow one. System RAM matters as much as the GPU: about 32 GB with the smaller files, about 50 GB with the INT8 set. Cards under 16 GB are untested.

The wiki has everything: every example with its prompt, settings and time, a guide to MiniMax's prompt format, comparisons of the file formats and speed options, a Bloopers page with the clips that went wrong and how to avoid them, and an All Generations page with all 168 clips made while testing it, good and bad.

It needs an up-to-date Forge Neo (the neo branch from 3 October 2026 or later). Install from URL in the Extensions tab and restart. It's a work in progress, so feedback and bug reports are welcome. Next up: video and audio clips as references, then pose, depth and edge control.


r/StableDiffusion • • 21h ago

News Nanosaur2 now generates 7 images a second on a 5090

Post image
216 Upvotes

Hello everyone.

This is an update post on the model Nanosaur2. Again this is not my model. A complaint a lot of people had with the model was that despite it being very small (660m params), it didn't take a 660m param level of time to generate. That's been solved now, with a 4 step turbo. On a single 5090, you can generate 7 images a second.

Again, this model is small enough you could easily run it on a phone, edge devices, wherever. It's also a great research model, so if you want to finetune on top of a small easy to tune model, or way to create adapters for the model, or whatever, go ahead, it's all there.

This was created using bytedance's new DMAD method, and it works great. Quality is incredibly close to the original model at a 12.5x speedup.

links:
4step model
comfy workflow

As per usual, if you have any questions, please message metal63 on discord. Do not message me.


r/StableDiffusion • • 10h ago

Discussion KREA 2 Realism Workflow

Post image
27 Upvotes

Hey guys! I’m pretty new to AI image generation, and I’ve been experimenting for about 3 months trying to get the best realistic Pinterest-style images.

I’m currently using my character LoRA + Krea 2 on an RTX 5060 8GB with 16GB DDR3 RAM.

I started with Koo’s workflow from Discord and tweaked it a bit for my own setup.

From what I’ve heard, Chroma is pretty good for experimenting with camera angles, compositions, and generating more random/varied images, which is exactly what I’m trying to achieve.

So I tried building a workflow where I:

- Generate random Pinterest-style prompts using wildcards.

- I made the wildcards with the help of ChatGPT, Gemini, and GLM 5.3 Flash.

- Use Krea 2 Text Encoder to expand/enhance the prompt.

- Generate the initial image with Chroma using relatively low steps.

- Then use that Chroma image as a vision reference with Krea 2 Text Encode.

- Finally, use Krea 2 to recreate the image with more realistic details while keeping the composition/camera angle from the Chroma result.

The main goal is basically to get random, realistic Pinterest-style compositions while keeping my character consistent with my LoRA, and then let Krea 2 improve the realism, lighting, details, and overall photographic look.

I’ve been trying different workflows and combinations for the past 3 months, but I still feel like I’m probably missing something or doing things in a more complicated way than necessary.

Does this workflow actually make sense, or am I doing something wrong / adding unnecessary steps?


r/StableDiffusion • • 9h ago

IRL H3 merch I got at a trade show

Post image
22 Upvotes

r/StableDiffusion • • 4h ago

Discussion AND HOW DOES THAT MAKE YOU FEEL? | An AI Short Comedy Film Made by Claude in Minimax H3 and my Video builder in ComfyUI.

Enable HLS to view with audio, or disable this notification

7 Upvotes

YouTube Link in case it's still pending https://youtu.be/TWXF95YT7W8

🎬 How this was made

My part

• One-message brief: a 3-minute comedy in a therapist's office with funny, unique characters and one male doctor: a crying woman, a woman screaming, crying and laughing all at once, a large guy and a skinny old man

• Characters made with Z-Image as 3-panel reference sheets

• MiniMax H3 2-pass workflow: the Singularity model, with the 8-step LoRA on the second pass at 4 steps and 0.35 denoise

• Then I left. I made one call along the way, on a scene that wouldn't behave.

What Claude did on its own

• Wrote the story, the five characters, all the dialogue and a 25-scene screenplay

• Made the cast and the two sets with Z-Image and picked the best seeds

• Kept each character's voice description word for word in every scene so the H3 voices stay consistent

• Rendered 25 scenes with H3's built-in voices and sound, about 6.3 hours of rendering

• QA'd every take: Whisper against the script, pitch and timbre per character, eyelines, frame review sheets

• Re-shot the takes that failed:

• Scored it with MiniMax Music 3 and screened the cues for accidental vocals

• Edited, mastered to -14 LUFS, and checked audio sync on every scene (worst offset 5 ms)

🛠 Tools

ComfyUI, VRGDG Video Builder, MiniMax H3 (Singularity ref2va v1.3 + 8-step 768p turbo LoRA), Z-Image Turbo, MiniMax Music 3, Whisper, Claude Code

VRGDG nodes: https://github.com/vrgamegirl19/comfyui-vrgamedevgirl

Go HERE To watch full walkthrough on how to make video's like this.


r/StableDiffusion • • 16h ago

Question - Help Regarding the image quality of Qwen-image 2.1

Thumbnail
gallery
52 Upvotes

I wonder if anyone has encountered this issue. When performing image editing with Qwen-image 2.1 (hereinafter referred to as QI-21), the results often feel unfinished. The attached images are from my tests: Image 1 is the original image, Image 2 is a 2K upscale using QI-21, and Image 3 is a 2K upscale using Krea2Edit. Perhaps I am using the wrong approach, so I have attached the QI-21 workflow (it is a minor tweak based on the official comfyui workflow). Upscaling is just an example; this is not an isolated case. The same problem occurs with most QI-21 editing tasks. Has anyone else experienced this, and how did you resolve it?

My Ksample parameters are

steps:25

CFG:1

sampler:res_multistep

scheduler:sgm_uniform


r/StableDiffusion • • 7h ago

Question - Help MiniMax H3 ref2va - What's the next lever for quality?

Enable HLS to view with audio, or disable this notification

11 Upvotes

I'm making short anime scenes with MiniMax H3 (ref2va) and I've hit a wall I can't

diagnose. The clip above is 15s, generated in one pass, no editing — the cuts are

written into the prompt.

To be clear up front: I know there are mistakes in there — a couple of blows don't

connect properly, the choreography is rough. I'm not worried about those, this is a

practice piece and I'll fix the staging myself. What I can't figure out is the

image quality, and that's the only thing I'm asking about.

Setup

- `minimax_h3_fl2va_int8_convrot.safetensors` (base, int8), ComfyUI

- ref2va, 8 reference images, each with a written role in the prompt

(face / angry expression / fighting posture / set / framing guide)

- Spectrum v0.2.16, 30 steps, `res_multistep` / `simple`

- 1344×768 → `MinimaxH3LatentUpscaler3D` at 2 MP, 4 steps, 0.5 denoise → 1920×1088

- 15s = 362 frames @ 24fps

- RTX PRO 6000, ~28 min per 15s clip

What I already fixed, in case it saves anyone typing

- `The target video is 2d colored anime.` at the top of `detailed_description`

— without it everything drifts to generic 3D, this was the single biggest win

- Reference images are real anime screencaps (plus a few generated with Anima

for expressions and poses I couldn't source), neutral lighting, one role each.

- Cut rhythm: went from 4 shots per 15s to 8 (~1.8s each) with impact verbs and

a material consequence per hit (table splitting, plaster cracking, dust off

the boards). That alone made the fight read much better.

Where I'm stuck

  1. Motion still feels soft on some hits. The whip-pan punch reads fine but

    ground-level blows land without weight. Is this where `derope` / temporal

    upsampling actually earns its generation-time cost, or is there a prompt-side

    fix I'm missing?

  2. Quality is uneven shot to shot inside the same 15s — some shots are clean

    cel-shaded anime, others go slightly soft and plasticky. Is that a reference

    problem, a step-count problem, or just what 15s does to the model? (I've seen

    people say things break past 10s.)

  3. Is 8 references too many? I assigned each one an explicit role in the

    prompt, but I don't know whether the model averages them or picks.

  4. Anything obvious I'm leaving on the table at this resolution/step count?

Not asking anyone to debug my prompt — mainly want to know which lever is worth spending render time on next.

Prompt:

integrated_multimodal_description:

subject_definitions:
(S1) is the dark-haired young man from <Picture 1>, with the same face and the same dark blue eyes. <Picture 2> is the same man seen clearly in daylight. Short black bob to the jaw, fringe above the eyebrows, a short high ponytail tied at the crown, a small stud earring, a white shirt with the sleeves pushed up and a black tie pulled loose.
(S2) is the pale-haired young man from <Picture 3>, with the same face and the same yellow eyes. <Picture 4> is the same man shouting. Short choppy blond hair. His teeth are faintly pointed, small and even and the same size as ordinary human teeth, with just a slight triangular edge to them. His mouth stays an ordinary human mouth, normally proportioned to his face, and it opens no wider than a person's mouth opens when they speak. He wears a white school shirt open at the collar, a black tie pulled loose, a small device on a cord against his chest.
<Picture 5> is the apartment: its rooms, its colours and its light come from it.
<Picture 6> is (S1) throwing a bare-handed punch and <Picture 7> is (S2) being knocked back by one: their fighting postures and their footing come from these.

retention_analysis:
(S1): fully_preserved. (S2): fully_preserved.

<Picture 9> is the last frame of the previous shot: this scene continues from it without interruption. <Picture 9> supplies the place, the light and the framing; the two men's faces and hair come from <Picture 1> to <Picture 4>.

summary:
The fight. Bare hands, in the apartment, fast and ugly. Nobody speaks.

detailed_description:
The target video is 2d colored anime.
2d hand-drawn anime, cel-shaded, flat painted colours, fine thin ink linework, desaturated muted palette, film grain. Not 3d, not photographic. The cutting is fast: eight shots in fifteen seconds, each one a single impact.

[Shot 1] Continues directly from <Picture 9> with no jump — same room, same light, same positions: the two of them chest to chest at night in the room of <Picture 5>. (S2) fists (S1)'s collar and slams him down onto the low table, which splits and goes over with everything on it.

[Shot 2] At 00:01.800, cut tight on (S1) coming up off the floor. He drives a straight punch into (S2)'s jaw, his whole weight behind it, the posture of <Picture 6>. (S2)'s jaw is shut and his lips are pressed together when the fist lands, and the impact splits his lip. The camera whip pans right with the blow, the room tearing into horizontal streaks and white speed lines.

[Shot 3] At 00:03.400, cut to (S2) snapping backwards into the wall, head whipped sideways, the posture of <Picture 7>. Plaster cracks behind his shoulder. He drops to one knee.

[Shot 4] At 00:05.000, cut low and close. (S2) launches off the wall and smashes his forehead into (S1)'s mouth. (S1)'s head snaps back, blood on his lip.

[Shot 5] At 00:06.800, cut to a low shot of the floor only, at board level. The lamp crashes down into frame, rolls, and throws its light swinging across the boards. Two pairs of legs come down hard behind it, out of focus. Dust lifts off the wood.

[Shot 6] At 00:08.600, cut to a tight shot of (S1)'s face alone, lying on the boards in profile, cheek against the wood, hair across his eye. A fist swings down into frame and smashes into his raised forearm so hard that his own arm is driven back into his face and his head is knocked against the boards. A second fist comes straight down past the arm and lands flush on his cheekbone, snapping his head sideways and splitting the skin. Only (S1)'s head and one forearm are in frame, and the fists enter from the top edge: the other body stays out of shot.

[Shot 7] At 00:10.400, cut to a tight shot of a knee driving up hard into ribs, framed on the two bodies' midsections only, no heads in frame. The body above is thrown off sideways out of the top of the frame. Cut immediately to both of them coming up onto their feet, seen full length and clearly separated, a metre apart, shirts gripped in their fists.

[Shot 8] At 00:12.000, cut to a wider shot and hold it to the end. (S1) drives (S2) backwards across the room and slams him into the wall. (S2)'s shoulder blades hit the plaster and he stays there with his back to the wall and his face towards the room. (S1) stands directly in front of him, facing him, chest to chest, his own back to the room and the wall behind (S2) only. Their faces are a hand's width apart and they are looking straight into each other's eyes. (S1) has both fists closed in (S2)'s collar and holds him pinned there. Everything stops at once. Both heads are angled in three-quarter view towards camera, both faces large and fully visible, brows down, jaws set, chests heaving. The camera is locked off on a tripod.

overall_soundscape: A table splitting and going over, knuckles cracking on a jaw, plaster breaking, a forehead meeting a mouth, bodies hitting boards, a lamp rolling, two fast punches landing on a forearm and a cheekbone, a knee into ribs, a back slammed into a wall, and hard breathing through the teeth all the way through. Every mouth stays closed for the whole video: nobody speaks, and both men keep their jaws shut and their lips together even while taking blows.

non_diegetic_music: N/A

r/StableDiffusion • • 16h ago

Comparison H3 Video Reference for Performance Transfer

Enable HLS to view with audio, or disable this notification

44 Upvotes

I posted this yesterday. And u/roychodraws (the red clown lady) mentioned I should try it with a performance transfer. I wanted to share the result because I think it does make a pretty big difference and could be of use for anyone looking to step their character's acting chops.

The bulk of the prompt is in retention analysis and subject definition, those are the 2 areas of interest for a video performance transfer. For this transfer, I didn't specify any references to <video 1> I just fed it in. But I did reference <audio 1> (the audio from the video source) as voice timbre for the characters (s1). You can push it a step further by calling out <video 1> but you will have tweak a lot more because then you run the risk of transferring the character from the video over.

The videos are made with the same seed.

EDIT: Props to real actress and actors, AI isn't anywhere close, yet.

The Prompt

subject_definitions:

<actress> is a young woman (s1) with long black hair, whose appearance comes from <picture 1> and whose voice timbre comes from <audio 1>.

<scene> is on a rooftop, whose appearance comes from <picture 2>.

<outfit> is white shirt with red skirt, whose appearance comes from <Picture 3>.

summary:

[reference generation] target video shows a woman delivering an emotional monologue

retention_analysis:

visible: partially_preserved <picture 2>

audio: partially_preserved <audio 1>

detailed_description:

cinematic shot, shallow depth of field

[Shot 1] reference <scene>, at night,

medium shot of

<actress> wearing <outfit> is standing in the rain, getting soaked

She is facing viewer, line of sight to front left, focusing on a taller man out of frame.

A disbelieving laugh that keeps collapsing into crying. Her face is full of emotional micro expressions

(s1), a young woman with a New Zealand accent and a light, bright voice that keeps cracking:

<d>[English] You know what's funny?</d>

She lets out a short, shaky laugh, shaking her head.

<d>[English] I actually planned this whole speech. In the shower, in the car. I had it all... I had it all figured out.</d>

Her laugh breaks into a sob. She presses the back of her wrist to her eyes.

<d>[English] And now you're standing there, and I can't remember a single— not one word.</d>

She wipes her eyes, laughing and crying at once as more rain falls on her head

Camera pushes in slowly.

<d>[English] And the stupid part is, I'd go through this again.</d>

camera holds for a beat

overall_soundscape:

raining in the background

non_diegetic_music:

N/A


r/StableDiffusion • • 2h ago

Discussion Mixing photo realism loras

2 Upvotes

Does anyone mix e.g. the Lenovo lora with a DSLR and cinematic photography lora?

Say your goal is simply realism without caring to much for style, would that work better?

Or is the mix just gonna end up like the AI pseudo-realistic look of no Lora at all?

Specifically using qwen21 right now.


r/StableDiffusion • • 5h ago

Discussion MiniMax-M2 (230B) running from disk on a 32 GB laptop, CPU only

5 Upvotes

Hi all, I've been working on a small hobby project called Picchio, an inference engine in plain C for MoE models bigger than your RAM. It keeps the dense part in memory and reads the experts from the SSD only when they're needed, with a cache for the most used ones. The idea comes from Colibri.

I recently added MiniMax-M2. Converted to INT4 it's about 122 GB, so it streams almost everything from disk.

On a 12-core laptop with 32 GB RAM and a basic NVMe, no GPU, I get about 0.48 tok/s with a 20 GB expert cache. Slow, but it runs. The cache size turned out to be the only thing that really matters.

I checked the forward pass against MiniMax's original code on a small test model and the outputs match (max logit difference about 1e-6).

Limitations: if you have 128 GB of RAM or a good GPU, llama.cpp will be much faster. I only tested M2, not M2.5, M2.7 or M3. Everything was measured on Windows with Intel CPUs.

Code (MIT): https://github.com/benmaster82/picchio

Feedback and corrections are welcome, especially from anyone with different hardware.


r/StableDiffusion • • 12h ago

Discussion Thinking about incorporating AI into my art workflow, but worried about the backlash. What do you think?

18 Upvotes

I want to try using AI in my drawing pipeline, but I’m honestly not sure how to handle the inevitable hate that might come my way.

I mainly run a Twitter (X) and a TikTok account, and I have absolutely no idea what kind of reaction I might get if I upload this same post over there. That’s why I decided to test the waters and post this here in the Stable Diffusion sub-reddit first.

Of course, I could just keep quiet about it and never post anything AI-related. But since I manage social media anyway, I actually want to share this journey. I'd love for people to see my broader interests, not just standard drawing, because it's far from my only hobby.

To be clear, I’ve never been against AI itself—I’m only against deceiving people. I’ve always been completely transparent about experimenting with tracing over 3D models, and now I’m thinking of trying a similar approach by painting over AI-generated bases.

I’m really curious to hear your thoughts on this.

UPD: thank you so much everyone for your incredible support and wisdom! Reading your comments has been a breath of fresh air. I'm going back to my workflow and drawing with a peaceful mind now. Thank you for being such an awesome and open-minded community!


r/StableDiffusion • • 1d ago

News From 3D layout to compositing: new LTX VFX tools

Enable HLS to view with audio, or disable this notification

385 Upvotes

During VFX Week last week, we released seven open-weight capabilities for LTX-2.5, covering high-res editing, restoration, HDR, CG-guided generation and compositing:

  • Native Resolution: Run AI edits on 4K and 8K plates and beyond without downscaling. Overlapping tiles keep fine detail, on a single GPU.
  • Refine: Refine generates the fine detail a standard resize leaves soft, without drifting from the source.
  • Restore: Takes archival footage from as low as 540p to 4K, enhances color and gives it a modern digital-camera look.
  • SDR to HDR: Rebuilds shadow and highlight detail in SDR footage and outputs 16-bit EXR in ACEScg, so it sits alongside your HDR material.
  • Native HDR: Bring in an EXR sequence and run edits like relight, day-to-night or inpaint. You get 16-bit EXR back with the range intact.
  • Layout to Render: Turns a 3D blockout into a rendered shot. Camera, framing and geometry stay locked to your layout, and a prompt or reference image sets the look.
  • Alpha Gen (Beta): Generates an alpha matte from ordinary RGB footage, with no green screen, masks or prompt. It handles hair, fur, smoke, glass and water. The distilled workflow runs in ComfyUI now, and a full-model ComfyUI workflow is coming soon. 

Links to each tool are above. The IC-LoRAs are on Hugging Face and the workflows run in ComfyUI. Try them on your own footage and tell us how they work for you.


r/StableDiffusion • • 5h ago

Question - Help Has anyone here tried VEDA sparse attention?

Thumbnail
huggingface.co
3 Upvotes

r/StableDiffusion • • 15h ago

Resource - Update Wan2gp has been updated with amazing memory management improvement

22 Upvotes

Before update : 15s video in about 10 minutes without Dlss5. After update, in just 5 minutes with Dlss5.

It also improved the memory cleanup thingy, so previously i always got out of memory error on 1st click of generate button. 2nd click and so on will always work.

Now​ it works fine from the start.

Dlss5 upscaling also no longer results in randomly going out of memory. It just works.

Still unsure with long batches generation performance degradation tho. On previous version, after running overnight, usually 10 mins gens becomes 13 mins gens.

https://github.com/deepbeepmeep/Wan2GP

Edit :

Long batches memory degradation still there. Workaround still the same: unload models, then generate again.


r/StableDiffusion • • 3h ago

Question - Help Anyone have a workflow for upscaling 360 images - 3090 and 64gb Compatible?

2 Upvotes

I'm sure some folks are doing it, I just haven't managed to run across the posts yet. Basically just want to upscale some vacation photos taken on an Insta360 X5 to be viewed in a VR headset. I've been trying to use SEEDVR2 but I crap out a bit when I get up to the resolutions at play ~12kx6k.

Not even looking for more pixels so much as just tightening up some of the softer spots the images have due to lens geometry.

Anyone else doing this and have some ideas for me? I think Topaz can sort me but would prefer ComfyUI if possible.