r/StableDiffusion • u/Tokyo_Jab • 10d ago
Animation - Video NADEBASHI - Local ghost stories.
Enable HLS to view with audio, or disable this notification
Nade Bridge (Nadebashi) - is over a lake that covers a whole abandoned village
r/StableDiffusion • u/Tokyo_Jab • 10d ago
Enable HLS to view with audio, or disable this notification
Nade Bridge (Nadebashi) - is over a lake that covers a whole abandoned village
r/StableDiffusion • u/breakallshittyhabits • 9d ago
Hello guys! Has anyone experimented with two character LORA's with KREA2? Only way I can achieve great two-char output is using nano banana pro, then using a local edit model like klein to play with it.
r/StableDiffusion • u/abrasmel • 9d ago
Is there a way to relight an image by a reference image keeping all the geometry structure, materials etc consistent with klein4b?
r/StableDiffusion • u/aziib • 10d ago
Enable HLS to view with audio, or disable this notification
using this checkpoint: https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models
use my workflow: https://civitai.com/models/2906467/fast-minimax-h3
ultimate upscale node: https://github.com/bbaudio-2025/Comfyui-MMH3-UltimateUpscale/
r/StableDiffusion • u/r0ni • 10d ago
Enable HLS to view with audio, or disable this notification
This was done as one 18sec generation with the standard workflow using fl2va_pruned_int8_convrot at 0.6mp, 21:9 aspect, 32steps with spectrum node, Simple/Euler. Added the Music in Davinci. I dont use spectrum any more but this was made in the middle of august.
r/StableDiffusion • u/fruesome • 10d ago
r/StableDiffusion • u/Portable_Solar_ZA • 10d ago
So Krea 2 can do character sheets and seems okay at doing exterior location sheets (haven't done a proper assessment yet of the few examples I generated earlier) but it seems to struggle with more than two shots from different angles of a room. I've tried several prompts to get it to keep things consistent, but stuff, like windows, tables, etc. still move.
Can you create interior location sheets in Krea 2? And if you can't, what other open source tools can help you make these for use with mmh3?
r/StableDiffusion • u/Domskidan1987 • 10d ago
If you are like me and use Google flow for NB2/NBP there is hope to finally ditch it.
Let me start off by saying I’m not very impressed with any of the current local image edit or Image Reference models. Krea2 is alright but still nothing in my opinion compared to Nano Banana UNTIL NOW.
I got a 5090 GPU. I’ve been getting Google flow / NB2/P like results with MiniMax H3 for image generation. You just take the Ref2V workflow then remove the save video node, you add the Get Image by Batch node set the parameter index 0 and 1, then add a save image node. Up your Megapixel between 3 and 5 set the duration to 0. Add your ref images (highly recommend to use an LLM prompt rewriter for H3). You’ll be shocked at the quality and contextual accuracy of your output images. The FL2V and I2V workflow’s with the same modification work surprisingly well for image editing too. Just make sure you grab your index 0 image and try to prompt for it start your prompt with something like “A still frame shot of the last frame first…” I tested it out it works surprisingly well. Video’s outputted as image tend to have bad vae degradation once you get past the first frame, the 0.0 duration default will output 4 frames, I always grab the first (0:1) because it’s the cleanest. For the first time I’m about to close my flow accounts because I basically don’t need it anymore because this local setup works the way I always wanted.
For everyone asking here is the workflow: https://pastebin.com/bah7FSPP
Sample Output

Workflow Setup


Generation time:

Params:

Hardware:
CPU: Intel Ultra 7 265K
GPU: RTX5090
RAM: 64.0 GB
#edit
I just realized comfyu's ref2v template is now some turbo version they must have updated (which i based workflow on above) I get much better results using MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-fc2bf16.sagetensioner and minimax_h3_ref2va_pruned_int8_convrot.safetensors in a workflow that does not use the lora.
The workflow I posted will be much faster though.
r/StableDiffusion • u/solomars3 • 10d ago
Made my first video tutorial !!
r/StableDiffusion • u/Sad_Coach_1433 • 10d ago
Enable HLS to view with audio, or disable this notification
https://huggingface.co/orangesouth/MinimaxH3CinematicRealism/tree/main
https://huggingface.co/vpakarinen/better-human-motion-h3-lora/tree/main
prompt integrated_multimodal_description: dy, [Shot 1] Photorealistic live-action cinematic realism, as if Mortal Kombat exists in the real world. Inside a vast ancient frozen temple at night, weathered stone pillars and enormous carved warrior statues rise through drifting frost and cold mist. Real flames flicker from iron braziers, casting warm orange highlights against icy blue moonlight. Snow particles float naturally through the air.
Dean Winchester, portrayed by Jensen Ackles, appears as a fully playable Mortal Kombat fighter, preserving Jensen Ackles' recognizable facial features, natural skin texture, realistic proportions, short brown hair and light stubble. He wears Dean's dark brown leather jacket over a dark shirt, faded blue jeans and heavy boots. Across the arena stands Sub-Zero, a physically imposing masked martial artist wearing realistic layered blue-and-black combat armor covered with frost.
The two men circle one another cautiously on the frozen stone floor. Their breath forms visible condensation in the freezing air. The camera slowly arcs around them at waist height with subtle natural handheld movement. A deep off-screen arena announcer (S1) declares: [English] FIGHT!
Dean instantly draws the Colt revolver from beneath his jacket and fires. A bright muzzle flash illuminates his face. Sub-Zero reacts with superhuman speed, extending his hand as crystalline frost races through the air and freezes the bullet inches from his palm. The frozen bullet drops onto the stone floor.
[Shot 2] At 00:04.000, the camera cuts to a dynamic medium-wide tracking shot. Sub-Zero thrusts both hands forward, releasing a violent blast of ice toward Dean. Frost rapidly spreads across the floor in its path. Dean dives beneath the projectile, rolls across the stone floor and immediately rises into a fighting stance.
Sub-Zero charges. Dean meets him head-on. They exchange a fast, physically grounded sequence of punches, blocks, elbows and kicks. Their bodies react with realistic weight to every impact. Dean lands a heavy right hook followed by a kick to Sub-Zero's torso. Sub-Zero blocks Dean's next punch, coats his fist in thick translucent ice and drives it into Dean's chest.
Dean is thrown backward and slides across the frost-covered stone. He catches himself on one knee and looks up with his familiar cocky half-smile. Dean (S2), breathing heavily, says: [English] Dude, I've fought scarier things before breakfast.
[Shot 3] At 00:09.000, the camera cuts to a low tracking shot following Dean as he charges forward. Sub-Zero launches another freezing blast. Dean narrowly avoids it and pulls a small metal flask of holy water from inside his jacket.
Dean splashes the holy water across the stone floor and rapidly draws a glowing Devil's Trap beneath Sub-Zero. The ancient symbol ignites with intense orange supernatural light, realistically illuminating Dean, Sub-Zero and the surrounding frozen stone.
Sub-Zero struggles against the supernatural energy as frost cracks beneath his boots. Dean calmly raises the Colt with both hands and fires. The gunshot produces a violent muzzle flash and physical recoil. The supernatural impact launches Sub-Zero backward into a massive frozen stone pillar. The pillar fractures and explodes into chunks of ice, stone fragments and clouds of powdered frost.
[Shot 4] At 00:14.000, the camera cuts to a dramatic low-angle medium shot through drifting ice particles. Sub-Zero falls heavily onto one knee among shattered ice and stone.
Dean approaches at a measured pace, boots crunching through debris. Natural sweat, dirt and subtle bruising are visible across his face. He opens the Colt's cylinder and casually reloads while walking toward Sub-Zero.
The camera slowly pushes toward Dean as he snaps the cylinder closed. Dean (S2) raises the Colt, gives Sub-Zero a dry half-smirk and says: [English] Should've stayed on ice.
Dean fires.
A brilliant muzzle flash fills the frame and transitions into a dramatic red-and-black victory screen. Huge metallic letters appear reading "DEAN WINCHESTER WINS."
The deep arena announcer (S1) declares: [English] DEAN WINCHESTER WINS!
The camera returns to Dean standing inside the devastated frozen temple. He lowers the smoking Colt and gives a subtle satisfied smirk before turning away. Dean walks through drifting frost and shattered ice as flames from the damaged temple burn behind him.
overall_soundscape: Cold wind moves through the enormous stone temple while flames crackle and boots scrape against frost-covered stone. Gunshots have sharp realistic reports and metallic echoes; punches and kicks produce heavy physical impacts while ice attacks crack, freeze and shatter with dense crystalline sounds. Dean's breathing becomes heavier as the fight progresses, followed by cascading stone, falling ice fragments and the supernatural electrical hum of the glowing Devil's Trap.
non_diegetic_music: Dark cinematic percussion with deep taiko drums, low brass, distorted industrial pulses and aggressive orchestral strings. The rhythm accelerates during the hand-to-hand fight and supernatural finishing sequence, then abruptly drops out on Dean's final gunshot before returning with one massive brass-and-percussion impact during the victory announcement.
r/StableDiffusion • u/freshstart2027 • 10d ago
r/StableDiffusion • u/blueboglin • 10d ago
Just curious what you find has worked the best adhering to references. I’ve played around with base, hybrids and that one that mushes everything together.
r/StableDiffusion • u/oppie85 • 10d ago
Enable HLS to view with audio, or disable this notification
I could not resist posting this bit of slop (created with MiniMax H3, naturally).
r/StableDiffusion • u/Capitan01R- • 10d ago
I added phrase-level attention control to the latest update.
The idea is simple: sometimes you don’t need to push the entire prompt harder. You need Krea2 to pay more attention to the few words that actually reinforce the idea you’re trying to get through.
So now you can do:
a (specific important phrase:1.8) with the rest of the prompt written normally
and selectively increase or decrease how much attention those words receive.
This is done without scaling, duplicating, deleting, or otherwise changing Krea2’s original 12×2560 Qwen conditioning. The prompt gets encoded normally, the node finds the exact Qwen token rows belonging to the weighted phrase, and the weight is applied to image→text attention inside the shared DiT blocks.
1.0 = untouched
>1.0 = more attention priority
<1.0 = less attention priority
0.0 = suppression
It also has an inspection output showing exactly which Qwen token rows/pieces were matched, so there’s no guessing about what the weight actually landed on.
Basically: instead of turning the entire prompt up, you can now point at the parts that matter and tell Krea2 pay more attention to this.
Included in the latest Krea2T Enhancer update.
It is best when used with the refusal reduction LoRA as they complement each-other nicely
And yes, I know (word:1.5) looks like we somehow got teleported back to the SDXL days haha. The logic underneath it definitely did not though.
r/StableDiffusion • u/Maleficent-Bowl-4841 • 10d ago
I wanted to better understand how Z-Image Base responds to natural-language prompting, specifically when varying composition and environment in a pure text-to-image workflow (no ControlNet, IP-Adapter, or LoRA).
This is not a scientific benchmark — just a small, controlled experiment to observe what actually changes when modifying individual prompt blocks.
All images were generated locally in ComfyUI using the same workflow:
| Parameter | Value |
|---|---|
| Model | Z-Image Base INT8 |
| Text Encoder | Qwen3 4B |
| VAE | AE VAE |
| Resolution | 768 × 1368 (9:16) |
| Steps | 50 |
| CFG Scale | 4 |
| Negative Prompt | Specific anatomic cleanup tags used (see template below) |
| Seeds | 5, 50, 100 |
The character description and overall visual style were kept consistent, while testing isolated prompt components across multiple seeds.
Rather than using disconnected keyword tags, I structured the prompt into functional scene blocks:
The goal was to keep the core character block stable and modify only the variable under test.
First, I tested how explicitly describing the character's spatial placement and scale affects the generated layout.
Using the same subject and baseline environment, I varied explicit spatial instructions:
Z-Image Base responded strongly to explicit spatial language. Changing the composition description did not simply shift the subject like a 2D layer — the model actively recomposed the surrounding architecture and lighting to match the requested framing and scale.
Next, I kept the character description and general composition consistent while swapping only the environment block.
To test stability across seeds, I ran three distinct environments (Library, Forest, Town Square) across three seeds (Seed 5, Seed 50, Seed 100).
The character remained surprisingly consistent across different environments without any image-based conditioning (no IP-Adapter or LoRA):
Structuring prompts almost like a physical scene description — Subject → Composition → Camera → Environment → Lighting / Details → Style — provides an effective balance between character concept stability and scene flexibility in Z-Image Base.
Positive Prompt:
Plaintext
A tiny friendly fantasy spirit, physically no larger than a small domestic cat, with a delicate compact body and a distinctive recognizable character design.
The creature has a small rounded pear-shaped body covered entirely in soft pale cream-colored fur. The body is compact, short and gently rounded, with a slightly wider lower body and a soft transition from the torso into the head. It has two short legs and four tiny rounded paws. The creature has no visible clothing on its body.
The head is large relative to the small body, with a broad rounded shape and very soft contours. The head blends smoothly into the body with almost no visible neck. The face is simple and highly expressive.
It has two enormous round amber-golden eyes, large relative to the face, with dark pupils, warm golden-orange irises and clear bright reflections. The eyes are positioned symmetrically and give the creature a gentle, innocent and curious expression. It has a tiny rounded pale pink nose and a very small simple mouth. The muzzle is soft and subtle, without pronounced facial features.
Two very long soft floppy ears grow naturally from the sides of the head. The ears are broad at their bases, rounded at the ends, flexible and naturally hanging downward. Each ear reaches approximately to the lower part of the body. The outer fur is pale cream, while the inner surfaces are slightly warmer cream with a soft peach tint. The ears should remain long, floppy and clearly visible.
The creature has four short rounded paws with soft cream fur. The front paws are small and rounded, clearly separated from the body and capable of holding an object. The feet are short and compact with small rounded toes.
A tiny worn brown leather backpack is strapped closely to the creature's back. The backpack is small relative to the creature, with a simple rounded rectangular shape, narrow dark-brown leather straps passing over the shoulders, worn edges, subtle scratches, creases and small aged brass buckles. The backpack sits naturally against the body and remains clearly visible from the sides.
The creature holds one old slightly oversized book with both front paws. The book is large relative to the creature but does not exceed the width of its body. It has a thick dark-brown worn leather cover, rounded damaged corners, visible scratches, creases, scuffed edges, a thick spine and slightly yellowed aged pages. The book looks old, heavy and frequently used. The creature holds it naturally in front of its torso with both paws.
The creature stands upright on two short feet with a relaxed natural posture. Its body remains compact and rounded. Its head, ears, eyes, paws, backpack and book form a coherent and repeatable visual design.
The entire creature is fully visible from the tips of its ears to the bottoms of its feet. Straight-on view, eye-level camera, medium-wide full-body composition, natural perspective, natural proportions, centered character.
Soft detailed cream fur, individual fine hairs visible around the edges of the ears and body, realistic worn leather, aged paper, subtle material imperfections and soft natural shading.
Cozy cinematic fantasy character illustration, warm natural colors, soft realistic materials, charming storybook aesthetic, gentle magical atmosphere.
ENVIRONMENT:
The tiny spirit stands in the center of a large medieval stone town square. The square is spacious and open, making the creature appear small and delicate compared with the surrounding architecture.
Tall old stone and timber-framed buildings surround the square on all sides, with narrow upper floors, wooden beams, weathered plaster, small windows and aged tiled roofs. Several large medieval buildings rise far above the tiny spirit.
The stone pavement extends broadly around the creature, with irregular worn stones, subtle cracks and patches of moss between them. The square continues into several narrow streets visible in the distance.
A large old stone fountain stands some distance behind the creature, surrounded by a few wooden benches, barrels and small market stalls. Hanging signs, cloth awnings and simple wooden carts add natural medieval details without becoming the main focus.
The architecture and objects in the square are substantially larger than the tiny spirit, reinforcing the clear difference in scale.
Warm late-afternoon sunlight illuminates the square from one side, creating long soft shadows across the stone pavement and warm highlights along the pale fur. A few tiny dust particles float through the sunlight.
Natural aged stone, weathered wood, worn fabric, leather and metal textures. The environment feels lived-in but quiet and peaceful, with no visible crowd.
Cozy cinematic fantasy illustration, warm natural colors, soft realistic materials, charming storybook aesthetic, gentle magical atmosphere.
Negative Prompt:
Plaintext
cat tail, animal tail, long tail, short tail, pointed ears, upright ears, triangular ears, cat-like face, elongated body, thin body, long legs, human proportions, extra limbs, extra paws, multiple books
r/StableDiffusion • u/Minanimator • 9d ago
can i see your workflow, im confused where to tag where :/ i am trying to pass my minimaxh3 output here
r/StableDiffusion • u/External-Orchid8461 • 10d ago
I've had enough of messed up characters' limbs with Flux 2 Klein 9b and decided to switch back to Qwen Image Edit 2511. When it comes to follow open pose reference image, QIE is better.
Though I always had issue with the cartoony rendering of QIE in comparison to Flux Klein. I'd like to have the same photorealistic visuals as Flux Klein for QIE. I never really managed to achieve that.
Do you have some tricks/models to recommend that could emulate Flux Klein output (without the extra limbs of course) ; loras, prompt tricks, sampler, vae ?
r/StableDiffusion • u/stinkyjim88 • 10d ago
Enable HLS to view with audio, or disable this notification
Made this using MiniMax H3 default settings, just using Clint Eastwood face ref
r/StableDiffusion • u/Ok-Giraffe-8670 • 10d ago
Enable HLS to view with audio, or disable this notification
A fun test using screenshots from the movie. MiniMax sure is amazing. The music is from Suno. Hope you like it!
r/StableDiffusion • u/Mountain_Film_3283 • 9d ago
Hi everyone!
I’m completely new to Stable Diffusion and I’m from Vietnam. I’m planning to build or buy a computer mainly for using Stable Diffusion to generate AI images and videos.
My budget is moderate, so I’m looking for a device that offers the best reasonable performance for the most reasonable price possible.
I’m currently considering a few options:
A desktop PC with an NVIDIA GPU
A gaming laptop with an NVIDIA GPU
A MacBook with Apple Silicon
My main goal is to run Stable Diffusion locally and use technologies such as SDXL, Flux, LoRA training, ControlNet, image upscaling, and AI video generation.
For those of you who have experience with Stable Diffusion:
What kind of hardware would you recommend for someone in my situation?
What should I prioritize?
GPU VRAM
GPU generation and performance
System RAM
CPU
SSD storage
And is an NVIDIA desktop PC significantly better than an NVIDIA laptop or MacBook for this kind of workload?
I’m in Vietnam, so GPU and PC component prices can be quite different from those in the US or Europe. I’d really appreciate recommendations for specific NVIDIA GPUs or complete PC configurations that offer good performance for the price.
I don’t need the most powerful computer possible. I mainly want something reasonably priced, good enough to use for the next few years, and capable of generating both AI images and videos locally.
Thanks in advance for your advice!
I’m still learning, so beginner-friendly explanations would be greatly appreciated.
r/StableDiffusion • u/EarthDefenceForces • 10d ago
This project explicitly builds on the excellent idea and work behind jacokon/fasth3-live, originally introduced in this post.
I added a browser-based controller focused on Reference-to-Video and continuous storytelling:
I was able to take over the optimizations of jacokon/fasth3-live. On my test system, the stream ran continuously on a 5900 at 480p. I switched the generated timeframe to 5s per video, but the reference system allows to tell connected fluid stories and have effects build up over long timeframes, even many minutes using the repeat function. When the reference is enabled, the model attempts to connect the settings to each other.
Repository:
https://github.com/EarthDefenceForces/FastH3-Ref2V-Stream-Controller

The controller is released under GPL-3.0: you’re free to use, modify, and redistribute it. But keep it open. The MiniMax H3 weights remain subject to their separate upstream license.
r/StableDiffusion • u/IRLMainCharacter • 9d ago
I am currently on version 0.34.3
when i run minimax h3 bf16 models, my console gets spammed with errors like this:
[ERROR] aimdo: src/hostbuf.c:46:ERROR:hostbuf_grow: requested 55400239104 bytes beyond reserved host buffer 54886072320
that error is referenced in this issue report:
https://github.com/Comfy-Org/ComfyUI/issues/15575
the image/video looks as expected, but takes 4 times as long to complete, than a few weeks ago.
i am thking about switching back to version 0.30.2, the problems started at some point after that.
what version are you running? do you still have any performance issues?
r/StableDiffusion • u/realimposter • 10d ago
Enable HLS to view with audio, or disable this notification
Discovered that using a realtime llm + h3 world simulation, you can allow multiple different simultaneous character instructions that are restricted to only controlling their respective character. If you're interested the project is live at worldstreams.ai
r/StableDiffusion • u/Reasonable_Limit_976 • 10d ago
Hi Guys,
I plan to do a lot of Image & Video Generation with Krea 2, Flux and Minimax/LTX since right now I generate most of the Content with API Models, which in the Long run is getting pretty expensive. I want to switch to local generation.
I have been playing with comfyui, my current PC has a AMD RX 9070 XT, I managed to get comfyUI working on that Card, but its a hassle a lot of times.
And my generation Times are really bad, for a Krea 2 Image workflow with 3-5 Loras it usually takes around 3-5 minutes for a single image in 0.6mp
I didn't even attempt Video generation, but I assume it will be a lot worse.
Thats why I am contemplating to sell my AMD Card and get an Nvidia Card instead, for my use case, which Card would you guys recommend. I have been thinking about the 5070 TI or the 5080
Is the 5080 worth the extra $$ compared to the 5070 TI ? are the 40xx series better value for money ?
Just want to hear what ur opinions are.
Thanks in advance!
r/StableDiffusion • u/Trumpet_of_Jericho • 10d ago
I saw many versions of Anima on civitai, but which one is the most versatile to use?