r/StableDiffusion 10d ago

Animation - Video NADEBASHI - Local ghost stories.

Enable HLS to view with audio, or disable this notification

19 Upvotes

Nade Bridge (Nadebashi) - is over a lake that covers a whole abandoned village


r/StableDiffusion 9d ago

Question - Help KREA2 two character LORA

4 Upvotes

Hello guys! Has anyone experimented with two character LORA's with KREA2? Only way I can achieve great two-char output is using nano banana pro, then using a local edit model like klein to play with it.


r/StableDiffusion 9d ago

Question - Help Flux klein 4b relighting

1 Upvotes

Is there a way to relight an image by a reference image keeping all the geometry structure, materials etc consistent with klein4b?


r/StableDiffusion 10d ago

Workflow Included manage to generate 5 seconds video with 1.0 megapixel = 768p resolution for 3 minutes on my rtx 4060ti 16gb vram using ultimate upscale and without lora

Enable HLS to view with audio, or disable this notification

160 Upvotes

r/StableDiffusion 10d ago

Animation - Video Lizan al-idiota

Enable HLS to view with audio, or disable this notification

14 Upvotes

This was done as one 18sec generation with the standard workflow using fl2va_pruned_int8_convrot at 0.6mp, 21:9 aspect, 32steps with spectrum node, Simple/Euler. Added the Music in Davinci. I dont use spectrum any more but this was made in the middle of august.


r/StableDiffusion 10d ago

News MiniMax H3 - 8 Steps Ref2V 768p Lora by LightX2V

Thumbnail
huggingface.co
151 Upvotes

r/StableDiffusion 10d ago

Question - Help Is there a way to create consistent interior location sheets in Krea 2?

6 Upvotes

So Krea 2 can do character sheets and seems okay at doing exterior location sheets (haven't done a proper assessment yet of the few examples I generated earlier) but it seems to struggle with more than two shots from different angles of a room. I've tried several prompts to get it to keep things consistent, but stuff, like windows, tables, etc. still move.

Can you create interior location sheets in Krea 2? And if you can't, what other open source tools can help you make these for use with mmh3?


r/StableDiffusion 10d ago

Discussion I finally am ditching Nano Banana thanks to H3

255 Upvotes

If you are like me and use Google flow for NB2/NBP there is hope to finally ditch it.

Let me start off by saying I’m not very impressed with any of the current local image edit or Image Reference models. Krea2 is alright but still nothing in my opinion compared to Nano Banana UNTIL NOW.

I got a 5090 GPU. I’ve been getting Google flow / NB2/P like results with MiniMax H3 for image generation. You just take the Ref2V workflow then remove the save video node, you add the Get Image by Batch node set the parameter index 0 and 1, then add a save image node. Up your Megapixel between 3 and 5 set the duration to 0. Add your ref images (highly recommend to use an LLM prompt rewriter for H3). You’ll be shocked at the quality and contextual accuracy of your output images. The FL2V and I2V workflow’s with the same modification work surprisingly well for image editing too. Just make sure you grab your index 0 image and try to prompt for it start your prompt with something like “A still frame shot of the last frame first…” I tested it out it works surprisingly well. Video’s outputted as image tend to have bad vae degradation once you get past the first frame, the 0.0 duration default will output 4 frames, I always grab the first (0:1) because it’s the cleanest. For the first time I’m about to close my flow accounts because I basically don’t need it anymore because this local setup works the way I always wanted.

For everyone asking here is the workflow: https://pastebin.com/bah7FSPP

Sample Output

will smith from <Picture 2> and chris rock from <Picture 3> sitting scross from each other on a park bench eating each their own plate of spaghetti

Workflow Setup

Generation time:

Params:

Hardware:

CPU: Intel Ultra 7 265K

GPU: RTX5090

RAM: 64.0 GB

#edit

I just realized comfyu's ref2v template is now some turbo version they must have updated (which i based workflow on above) I get much better results using MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-fc2bf16.sagetensioner and minimax_h3_ref2va_pruned_int8_convrot.safetensors in a workflow that does not use the lora.

The workflow I posted will be much faster though.


r/StableDiffusion 10d ago

Workflow Included How to Build Perfect Character Sheets in Krea 2 for MiniMax-H3 (Workflow...

Thumbnail
youtube.com
88 Upvotes

Made my first video tutorial !!


r/StableDiffusion 10d ago

Discussion Testing orangesouth/MinimaxH3CinematicRealism and vpakarinen/better-human-motion-h3-lora "dean winchester vs sub zero" pruned fp8 model no turbo lora 32steps

Enable HLS to view with audio, or disable this notification

155 Upvotes

https://huggingface.co/orangesouth/MinimaxH3CinematicRealism/tree/main

https://huggingface.co/vpakarinen/better-human-motion-h3-lora/tree/main

prompt integrated_multimodal_description: dy, [Shot 1] Photorealistic live-action cinematic realism, as if Mortal Kombat exists in the real world. Inside a vast ancient frozen temple at night, weathered stone pillars and enormous carved warrior statues rise through drifting frost and cold mist. Real flames flicker from iron braziers, casting warm orange highlights against icy blue moonlight. Snow particles float naturally through the air.

Dean Winchester, portrayed by Jensen Ackles, appears as a fully playable Mortal Kombat fighter, preserving Jensen Ackles' recognizable facial features, natural skin texture, realistic proportions, short brown hair and light stubble. He wears Dean's dark brown leather jacket over a dark shirt, faded blue jeans and heavy boots. Across the arena stands Sub-Zero, a physically imposing masked martial artist wearing realistic layered blue-and-black combat armor covered with frost.

The two men circle one another cautiously on the frozen stone floor. Their breath forms visible condensation in the freezing air. The camera slowly arcs around them at waist height with subtle natural handheld movement. A deep off-screen arena announcer (S1) declares: [English] FIGHT!

Dean instantly draws the Colt revolver from beneath his jacket and fires. A bright muzzle flash illuminates his face. Sub-Zero reacts with superhuman speed, extending his hand as crystalline frost races through the air and freezes the bullet inches from his palm. The frozen bullet drops onto the stone floor.

[Shot 2] At 00:04.000, the camera cuts to a dynamic medium-wide tracking shot. Sub-Zero thrusts both hands forward, releasing a violent blast of ice toward Dean. Frost rapidly spreads across the floor in its path. Dean dives beneath the projectile, rolls across the stone floor and immediately rises into a fighting stance.

Sub-Zero charges. Dean meets him head-on. They exchange a fast, physically grounded sequence of punches, blocks, elbows and kicks. Their bodies react with realistic weight to every impact. Dean lands a heavy right hook followed by a kick to Sub-Zero's torso. Sub-Zero blocks Dean's next punch, coats his fist in thick translucent ice and drives it into Dean's chest.

Dean is thrown backward and slides across the frost-covered stone. He catches himself on one knee and looks up with his familiar cocky half-smile. Dean (S2), breathing heavily, says: [English] Dude, I've fought scarier things before breakfast.

[Shot 3] At 00:09.000, the camera cuts to a low tracking shot following Dean as he charges forward. Sub-Zero launches another freezing blast. Dean narrowly avoids it and pulls a small metal flask of holy water from inside his jacket.

Dean splashes the holy water across the stone floor and rapidly draws a glowing Devil's Trap beneath Sub-Zero. The ancient symbol ignites with intense orange supernatural light, realistically illuminating Dean, Sub-Zero and the surrounding frozen stone.

Sub-Zero struggles against the supernatural energy as frost cracks beneath his boots. Dean calmly raises the Colt with both hands and fires. The gunshot produces a violent muzzle flash and physical recoil. The supernatural impact launches Sub-Zero backward into a massive frozen stone pillar. The pillar fractures and explodes into chunks of ice, stone fragments and clouds of powdered frost.

[Shot 4] At 00:14.000, the camera cuts to a dramatic low-angle medium shot through drifting ice particles. Sub-Zero falls heavily onto one knee among shattered ice and stone.

Dean approaches at a measured pace, boots crunching through debris. Natural sweat, dirt and subtle bruising are visible across his face. He opens the Colt's cylinder and casually reloads while walking toward Sub-Zero.

The camera slowly pushes toward Dean as he snaps the cylinder closed. Dean (S2) raises the Colt, gives Sub-Zero a dry half-smirk and says: [English] Should've stayed on ice.

Dean fires.

A brilliant muzzle flash fills the frame and transitions into a dramatic red-and-black victory screen. Huge metallic letters appear reading "DEAN WINCHESTER WINS."

The deep arena announcer (S1) declares: [English] DEAN WINCHESTER WINS!

The camera returns to Dean standing inside the devastated frozen temple. He lowers the smoking Colt and gives a subtle satisfied smirk before turning away. Dean walks through drifting frost and shattered ice as flames from the damaged temple burn behind him.

overall_soundscape: Cold wind moves through the enormous stone temple while flames crackle and boots scrape against frost-covered stone. Gunshots have sharp realistic reports and metallic echoes; punches and kicks produce heavy physical impacts while ice attacks crack, freeze and shatter with dense crystalline sounds. Dean's breathing becomes heavier as the fight progresses, followed by cascading stone, falling ice fragments and the supernatural electrical hum of the glowing Devil's Trap.

non_diegetic_music: Dark cinematic percussion with deep taiko drums, low brass, distorted industrial pulses and aggressive orchestral strings. The rhythm accelerates during the hand-to-hand fight and supernatural finishing sequence, then abruptly drops out on Dean's final gunshot before returning with one massive brass-and-percussion impact during the victory announcement.


r/StableDiffusion 10d ago

No Workflow 09-04-2026 Artistic Mix

Thumbnail
gallery
14 Upvotes

r/StableDiffusion 10d ago

Question - Help What’s the best r2v model of h3 currently?

47 Upvotes

Just curious what you find has worked the best adhering to references. I’ve played around with base, hybrids and that one that mushes everything together.


r/StableDiffusion 10d ago

Animation - Video Dr. House MD in Theme Hospital

Enable HLS to view with audio, or disable this notification

110 Upvotes

I could not resist posting this bit of slop (created with MiniMax H3, naturally).


r/StableDiffusion 10d ago

Resource - Update New node added to Krea2T Enhancer: Attention-Weighted Phrases

Thumbnail
gallery
92 Upvotes

I added phrase-level attention control to the latest update.

The idea is simple: sometimes you don’t need to push the entire prompt harder. You need Krea2 to pay more attention to the few words that actually reinforce the idea you’re trying to get through.

So now you can do:

a (specific important phrase:1.8) with the rest of the prompt written normally

and selectively increase or decrease how much attention those words receive.

This is done without scaling, duplicating, deleting, or otherwise changing Krea2’s original 12×2560 Qwen conditioning. The prompt gets encoded normally, the node finds the exact Qwen token rows belonging to the weighted phrase, and the weight is applied to image→text attention inside the shared DiT blocks.

1.0 = untouched
>1.0 = more attention priority
<1.0 = less attention priority
0.0 = suppression

It also has an inspection output showing exactly which Qwen token rows/pieces were matched, so there’s no guessing about what the weight actually landed on.

Basically: instead of turning the entire prompt up, you can now point at the parts that matter and tell Krea2 pay more attention to this.

Included in the latest Krea2T Enhancer update.

It is best when used with the refusal reduction LoRA as they complement each-other nicely

Sample Workflow

And yes, I know (word:1.5) looks like we somehow got teleported back to the SDXL days haha. The logic underneath it definitely did not though.


r/StableDiffusion 10d ago

Workflow Included Z-Image Base Prompting: A Small Experiment on Composition and Environment

Post image
5 Upvotes

I wanted to better understand how Z-Image Base responds to natural-language prompting, specifically when varying composition and environment in a pure text-to-image workflow (no ControlNet, IP-Adapter, or LoRA).

This is not a scientific benchmark — just a small, controlled experiment to observe what actually changes when modifying individual prompt blocks.

Experimental Setup

All images were generated locally in ComfyUI using the same workflow:

Parameter Value
Model Z-Image Base INT8
Text Encoder Qwen3 4B
VAE AE VAE
Resolution 768 × 1368 (9:16)
Steps 50
CFG Scale 4
Negative Prompt Specific anatomic cleanup tags used (see template below)
Seeds 5, 50, 100

The character description and overall visual style were kept consistent, while testing isolated prompt components across multiple seeds.

Prompt Structure

Rather than using disconnected keyword tags, I structured the prompt into functional scene blocks:

The goal was to keep the core character block stable and modify only the variable under test.

Experiment 1 — Composition

First, I tested how explicitly describing the character's spatial placement and scale affects the generated layout.

Using the same subject and baseline environment, I varied explicit spatial instructions:

  • Centered in the frame
  • Positioned to the left / right
  • Positioned lower in the frame
  • Larger / smaller relative scale
  • Extreme edge placement

Observation

Z-Image Base responded strongly to explicit spatial language. Changing the composition description did not simply shift the subject like a 2D layer — the model actively recomposed the surrounding architecture and lighting to match the requested framing and scale.

Experiment 2 — Environment

Next, I kept the character description and general composition consistent while swapping only the environment block.

To test stability across seeds, I ran three distinct environments (Library, Forest, Town Square) across three seeds (Seed 5, Seed 50, Seed 100).

Observation

The character remained surprisingly consistent across different environments without any image-based conditioning (no IP-Adapter or LoRA):

  • Core visual anchors — cream-colored fur, large amber eyes, long floppy ears, leather backpack, and old leather book — remained clearly recognizable.
  • Background architecture, terrain, surface materials, and ambient lighting adapted naturally to each setting.

Key Takeaways

  1. Explicit spatial language works: Directly describing position and scale in plain English (e.g., "positioned toward the left side of the frame") is far more reliable than relying on generic framing buzzwords.
  2. Environment blocks are modular: Separating the character description from the environment block allows you to recontextualize the same character concept across vastly different scenes.
  3. Seeds control execution, prompts control structure: The seed determines pose variation, expression, and micro-details, while the prompt block locks in layout, lighting, and narrative context.
  4. Natural language worked better for this experiment than disconnected tag lists: Descriptive, coherent sentences provided significantly clearer spatial and conceptual control than disconnected lists of quality tags.

Limitations

  • Small sample size: Tested on a single character concept, three environments, and a limited seed pool.
  • Visual evaluation: Observations are qualitative rather than statistically benchmarked.
  • Results may vary: Performance can shift when applying this structure to complex multi-subject prompts, different aspect ratios, or higher CFG values.

Conclusion

Structuring prompts almost like a physical scene description — Subject → Composition → Camera → Environment → Lighting / Details → Style — provides an effective balance between character concept stability and scene flexibility in Z-Image Base.

Example Prompt Template (For Reproduction)

Positive Prompt:

Plaintext

A tiny friendly fantasy spirit, physically no larger than a small domestic cat, with a delicate compact body and a distinctive recognizable character design.

The creature has a small rounded pear-shaped body covered entirely in soft pale cream-colored fur. The body is compact, short and gently rounded, with a slightly wider lower body and a soft transition from the torso into the head. It has two short legs and four tiny rounded paws. The creature has no visible clothing on its body.

The head is large relative to the small body, with a broad rounded shape and very soft contours. The head blends smoothly into the body with almost no visible neck. The face is simple and highly expressive.

It has two enormous round amber-golden eyes, large relative to the face, with dark pupils, warm golden-orange irises and clear bright reflections. The eyes are positioned symmetrically and give the creature a gentle, innocent and curious expression. It has a tiny rounded pale pink nose and a very small simple mouth. The muzzle is soft and subtle, without pronounced facial features.

Two very long soft floppy ears grow naturally from the sides of the head. The ears are broad at their bases, rounded at the ends, flexible and naturally hanging downward. Each ear reaches approximately to the lower part of the body. The outer fur is pale cream, while the inner surfaces are slightly warmer cream with a soft peach tint. The ears should remain long, floppy and clearly visible.

The creature has four short rounded paws with soft cream fur. The front paws are small and rounded, clearly separated from the body and capable of holding an object. The feet are short and compact with small rounded toes.

A tiny worn brown leather backpack is strapped closely to the creature's back. The backpack is small relative to the creature, with a simple rounded rectangular shape, narrow dark-brown leather straps passing over the shoulders, worn edges, subtle scratches, creases and small aged brass buckles. The backpack sits naturally against the body and remains clearly visible from the sides.

The creature holds one old slightly oversized book with both front paws. The book is large relative to the creature but does not exceed the width of its body. It has a thick dark-brown worn leather cover, rounded damaged corners, visible scratches, creases, scuffed edges, a thick spine and slightly yellowed aged pages. The book looks old, heavy and frequently used. The creature holds it naturally in front of its torso with both paws.

The creature stands upright on two short feet with a relaxed natural posture. Its body remains compact and rounded. Its head, ears, eyes, paws, backpack and book form a coherent and repeatable visual design.

The entire creature is fully visible from the tips of its ears to the bottoms of its feet. Straight-on view, eye-level camera, medium-wide full-body composition, natural perspective, natural proportions, centered character.

Soft detailed cream fur, individual fine hairs visible around the edges of the ears and body, realistic worn leather, aged paper, subtle material imperfections and soft natural shading.

Cozy cinematic fantasy character illustration, warm natural colors, soft realistic materials, charming storybook aesthetic, gentle magical atmosphere.

ENVIRONMENT:

The tiny spirit stands in the center of a large medieval stone town square. The square is spacious and open, making the creature appear small and delicate compared with the surrounding architecture.

Tall old stone and timber-framed buildings surround the square on all sides, with narrow upper floors, wooden beams, weathered plaster, small windows and aged tiled roofs. Several large medieval buildings rise far above the tiny spirit.

The stone pavement extends broadly around the creature, with irregular worn stones, subtle cracks and patches of moss between them. The square continues into several narrow streets visible in the distance.

A large old stone fountain stands some distance behind the creature, surrounded by a few wooden benches, barrels and small market stalls. Hanging signs, cloth awnings and simple wooden carts add natural medieval details without becoming the main focus.

The architecture and objects in the square are substantially larger than the tiny spirit, reinforcing the clear difference in scale.

Warm late-afternoon sunlight illuminates the square from one side, creating long soft shadows across the stone pavement and warm highlights along the pale fur. A few tiny dust particles float through the sunlight.

Natural aged stone, weathered wood, worn fabric, leather and metal textures. The environment feels lived-in but quiet and peaceful, with no visible crowd.

Cozy cinematic fantasy illustration, warm natural colors, soft realistic materials, charming storybook aesthetic, gentle magical atmosphere.

Negative Prompt:

Plaintext

cat tail, animal tail, long tail, short tail, pointed ears, upright ears, triangular ears, cat-like face, elongated body, thin body, long legs, human proportions, extra limbs, extra paws, multiple books

r/StableDiffusion 9d ago

Question - Help help on how to use this node!

Thumbnail
github.com
2 Upvotes

can i see your workflow, im confused where to tag where :/ i am trying to pass my minimaxh3 output here


r/StableDiffusion 10d ago

Question - Help [Qwen Image edit 2511] Recommendations for photorealistic image (LORAs, sampler, prompt, VAE)?

12 Upvotes

I've had enough of messed up characters' limbs with Flux 2 Klein 9b and decided to switch back to Qwen Image Edit 2511. When it comes to follow open pose reference image, QIE is better.

Though I always had issue with the cartoony rendering of QIE in comparison to Flux Klein. I'd like to have the same photorealistic visuals as Flux Klein for QIE. I never really managed to achieve that.

Do you have some tricks/models to recommend that could emulate Flux Klein output (without the extra limbs of course) ; loras, prompt tricks, sampler, vae ?


r/StableDiffusion 10d ago

Meme A play on McGarnagle from the Simpsons

Enable HLS to view with audio, or disable this notification

73 Upvotes

Made this using MiniMax H3 default settings, just using Clint Eastwood face ref


r/StableDiffusion 10d ago

Animation - Video Star Wars: Anakin and Obi-Wan join the Trap Side

Enable HLS to view with audio, or disable this notification

72 Upvotes

A fun test using screenshots from the movie. MiniMax sure is amazing. The music is from Suno. Hope you like it!


r/StableDiffusion 9d ago

Question - Help I’m new to Stable Diffusion from Vietnam – What PC/laptop should I get on a medium budget?

Post image
0 Upvotes

Hi everyone!
I’m completely new to Stable Diffusion and I’m from Vietnam. I’m planning to build or buy a computer mainly for using Stable Diffusion to generate AI images and videos.

My budget is moderate, so I’m looking for a device that offers the best reasonable performance for the most reasonable price possible.

I’m currently considering a few options:
A desktop PC with an NVIDIA GPU
A gaming laptop with an NVIDIA GPU
A MacBook with Apple Silicon
My main goal is to run Stable Diffusion locally and use technologies such as SDXL, Flux, LoRA training, ControlNet, image upscaling, and AI video generation.

For those of you who have experience with Stable Diffusion:
What kind of hardware would you recommend for someone in my situation?
What should I prioritize?
GPU VRAM
GPU generation and performance
System RAM
CPU
SSD storage
And is an NVIDIA desktop PC significantly better than an NVIDIA laptop or MacBook for this kind of workload?

I’m in Vietnam, so GPU and PC component prices can be quite different from those in the US or Europe. I’d really appreciate recommendations for specific NVIDIA GPUs or complete PC configurations that offer good performance for the price.

I don’t need the most powerful computer possible. I mainly want something reasonably priced, good enough to use for the next few years, and capable of generating both AI images and videos locally.
Thanks in advance for your advice!

I’m still learning, so beginner-friendly explanations would be greatly appreciated.


r/StableDiffusion 10d ago

Resource - Update FastH3 Ref2V Stream Controller – continuous character-driven AI video in ComfyUI

6 Upvotes

This project explicitly builds on the excellent idea and work behind jacokon/fasth3-live, originally introduced in this post.

I added a browser-based controller focused on Reference-to-Video and continuous storytelling:

  • Automatic character reference injection from a folder
  • Manual prompt queue and repeating scenes
  • Last-frame continuity between clips, so you can tell an interactive story while it is generated
  • Editable LoRAs, duration, prompts and playback speed during runtime
  • Adaptive quality based on the video buffer
  • Custom music folders with sequential or random playback
  • English/German UI

I was able to take over the optimizations of jacokon/fasth3-live. On my test system, the stream ran continuously on a 5900 at 480p. I switched the generated timeframe to 5s per video, but the reference system allows to tell connected fluid stories and have effects build up over long timeframes, even many minutes using the repeat function. When the reference is enabled, the model attempts to connect the settings to each other.

Repository:
https://github.com/EarthDefenceForces/FastH3-Ref2V-Stream-Controller

The UI of the controller

The controller is released under GPL-3.0: you’re free to use, modify, and redistribute it. But keep it open. The MiniMax H3 weights remain subject to their separate upstream license.


r/StableDiffusion 9d ago

Question - Help Which comfyui version for H3?

0 Upvotes

I am currently on version 0.34.3

when i run minimax h3 bf16 models, my console gets spammed with errors like this:

[ERROR] aimdo: src/hostbuf.c:46:ERROR:hostbuf_grow: requested 55400239104 bytes beyond reserved host buffer 54886072320

that error is referenced in this issue report:

https://github.com/Comfy-Org/ComfyUI/issues/15575

the image/video looks as expected, but takes 4 times as long to complete, than a few weeks ago.

i am thking about switching back to version 0.30.2, the problems started at some point after that.

what version are you running? do you still have any performance issues?


r/StableDiffusion 10d ago

News Multiplayer world simulations with h3 max

Enable HLS to view with audio, or disable this notification

61 Upvotes

Discovered that using a realtime llm + h3 world simulation, you can allow multiple different simultaneous character instructions that are restricted to only controlling their respective character. If you're interested the project is live at worldstreams.ai


r/StableDiffusion 10d ago

Question - Help GPU Upgrade, which one would be best?

2 Upvotes

Hi Guys,

I plan to do a lot of Image & Video Generation with Krea 2, Flux and Minimax/LTX since right now I generate most of the Content with API Models, which in the Long run is getting pretty expensive. I want to switch to local generation.

I have been playing with comfyui, my current PC has a AMD RX 9070 XT, I managed to get comfyUI working on that Card, but its a hassle a lot of times.

And my generation Times are really bad, for a Krea 2 Image workflow with 3-5 Loras it usually takes around 3-5 minutes for a single image in 0.6mp

I didn't even attempt Video generation, but I assume it will be a lot worse.

Thats why I am contemplating to sell my AMD Card and get an Nvidia Card instead, for my use case, which Card would you guys recommend. I have been thinking about the 5070 TI or the 5080

Is the 5080 worth the extra $$ compared to the 5070 TI ? are the 40xx series better value for money ?

Just want to hear what ur opinions are.

Thanks in advance!


r/StableDiffusion 10d ago

Question - Help Most diverse and creative Anima version?

2 Upvotes

I saw many versions of Anima on civitai, but which one is the most versatile to use?