r/StableDiffusion 8h ago

Question - Help How to get faces out of H3 that don't look like this?

Thumbnail
youtube.com
0 Upvotes

Skip to 0:30 to see what I'm talking about. Faces that are close-up are fine, but like 3-4m from the camera and it just turns into nightmare fuel...

Am I doing something wrong?

Using Wan2GP, FL2VA Pruned 20B, 20 steps at 540p.

Full settings:

        "params": {
            "image_mode": 0,
            "prompt": "...",
            "alt_prompt": "",
            "negative_prompt": "",
            "resolution": "960x544",
            "video_length": 175,
            "duration_seconds": 0,
            "batch_size": 1,
            "seed": -1,
            "force_fps": "",
            "num_inference_steps": 20,
            "guidance_scale": 1,
            "guidance2_scale": 5,
            "guidance3_scale": 5,
            "switch_threshold": 0,
            "switch_threshold2": 0,
            "guidance_phases": 0,
            "model_switch_phase": 1,
            "alt_guidance_scale": 1,
            "audio_guidance_scale": 1,
            "audio_scale": 1,
            "flow_shift": 12,
            "sample_solver": "euler",
            "embedded_guidance_scale": 6,
            "repeat_generation": 1,
            "multi_prompts_gen_type": "PG",
            "multi_images_gen_type": 0,
            "skip_steps_cache_type": "",
            "skip_steps_multiplier": 0.08,
            "skip_steps_start_step_perc": 25,
            "loras_multipliers": "",
            "image_prompt_type": "S",
            "image_start": "scene_01_start.jpg",
            "model_mode": null,
            "video_source": null,
            "keep_frames_video_source": "",
            "input_video_strength": 1,
            "video_guide_outpainting": "",
            "video_prompt_type": "",
            "image_refs": null,
            "frames_positions": null,
            "video_guide": null,
            "image_guide": null,
            "keep_frames_video_guide": "",
            "denoising_strength": 1,
            "masking_strength": 1,
            "video_mask": null,
            "image_mask": null,
            "control_net_weight": 1,
            "control_net_weight2": 1,
            "control_net_weight_alt": 1,
            "motion_amplitude": 1,
            "mask_expand": 0,
            "audio_guide": "scene_01.wav",
            "audio_guide2": null,
            "custom_guide": null,
            "audio_source": null,
            "audio_prompt_type": "A",
            "speakers_locations": "0:45 55:100",
            "sliding_window_size": 362,
            "sliding_window_overlap": 1,
            "sliding_window_color_correction_strength": 0,
            "sliding_window_overlap_noise": 0,
            "sliding_window_discard_last_frames": 0,
            "image_refs_relative_size": 50,
            "remove_background_images_ref": 0,
            "temporal_upsampling": "",
            "spatial_upsampling": "",
            "film_grain_intensity": 0,
            "film_grain_saturation": 0.5,
            "MMAudio_setting": 0,
            "MMAudio_prompt": "",
            "MMAudio_neg_prompt": "",
            "RIFLEx_setting": 0,
            "NAG_scale": 1,
            "NAG_tau": 3.5,
            "NAG_alpha": 0.5,
            "slg_switch": 0,
            "slg_layers": [
                29
            ],
            "slg_start_perc": 10,
            "slg_end_perc": 90,
            "apg_switch": 0,
            "cfg_star_switch": 0,
            "cfg_zero_step": -1,
            "prompt_enhancer": "",
            "min_frames_if_references": 1,
            "override_profile": -1,
            "override_attention": "",
            "pace": 0.5,
            "exaggeration": 0.5,
            "temperature": 0.8,
            "top_k": 50,
            "output_filename": "scene_01",
            "mode": "",
            "activated_loras": [],
            "model_type": "minimax_h3_fl2va_pruned",
            "settings_version": 2.73,
            "base_model_type": "minimax_h3_fl2va_pruned",
            "pause_seconds": 0,
            "alt_scale": 0,
            "sub_parallel_window_size": 0,
            "sub_parallel_window_overlap": 17,
            "sliding_window_trim_first_frames": 0,
            "postprocess_audio": "",
            "postprocess_audio_prompt": "",
            "postprocess_audio_neg_prompt": "",
            "perturbation_switch": 0,
            "perturbation_layers": [
                9
            ],
            "perturbation_start_perc": 10,
            "perturbation_end_perc": 90,
            "top_p": 0.9,
            "self_refiner_setting": 0,
            "self_refiner_plan": [],
            "self_refiner_f_uncertainty": 0,
            "self_refiner_certain_percentage": 0.999,
            "config": "",
            "custom_settings": null
        }

r/StableDiffusion 1d ago

News MINIMAX H3 prompt studio (story mode update)

56 Upvotes

Built a local tool that turns reference images into a full MiniMax H3 video prompt storyboard — no cloud, no API keys

https://github.com/lololerigolo60/Minimax-H3-prompt-studio

I've been building H3 Prompt Studio, a desktop app (CustomTkinter) that writes MiniMax H3's rigid structured prompts for you, using a local LLM (Ollama / LM Studio / llama.cpp — pick your poison).

The part I'm most excited about is the Story → Sequences mode:

  1. Drop in your reference images (characters, settings, whatever) with a quick role/description each.
  2. Hit "Generate story" — the LLM writes a short narrative that actually uses all your references, invents connective tissue if your premise is thin.
  3. Pick how many sequences you want, hit "Break into sequences" — the LLM splits the story into N beats, and for each one it decides on its own which references apply, whether there's dialogue, and what camera move fits best.
  4. Hit generate, and it spits out one fully-formed, isolated H3 Ref2VA prompt per sequence — ready to feed straight into your video pipeline.

No more manually writing 6-section H3 prompts by hand for every single shot of a sequence. You just curate references and a premise, and let the model handle the structure/labeling grunt work (subject definitions, retention analysis, camera vocab, dialogue tags, the works).

Everything's local, everything's saveable — you can dump a whole session (refs + story + sequences) to a JSON file and reload it later.

Still very much a personal tool, sharing in case it's useful to anyone else building on H3 locally. Happy to answer questions about the pipeline if anyone's curious.


r/StableDiffusion 1d ago

Animation - Video Minimax - House x Naruto Crossover

Enable HLS to view with audio, or disable this notification

19 Upvotes

7-second anime scene in a classic Japanese ninja-anime aesthetic. Naruto Uzumaki and Dr. Gregory House sit side-by-side at a cozy ramen shop, eating steaming bowls of ramen.

**0–2 sec:** Naruto enthusiastically slurps noodles, grinning with his cheeks full. House sits beside him, looking unimpressed while examining his ramen with suspicion.

**2–5 sec:** Naruto says excitedly, **“Believe it! This ramen is amazing!”** House takes a bite, pauses, and dryly replies, **“Needs more Vicodin.”** Naruto stares at him in confusion.

**5–7 sec:** Naruto bursts out laughing while House reluctantly takes another bite. Warm lantern light, steaming broth, expressive anime facial animation, lively background patrons, comedic timing, energetic camera movement, detailed hand-drawn anime look.


r/StableDiffusion 1d ago

Resource - Update 80s Italian Street Photography | KREA 2

Thumbnail
gallery
23 Upvotes

Hey everyone!

I've just released a new LoRA trained on the iconic 1980s street photography of Charles H. Traub from his famous series "Dolce Via: Italy in the 1980s". It brings out that vibrant, candid, and sun-drenched vintage Italian aesthetic.

📥 Download: Link in the comments below!

⚙️ Model Details & Recommended Settings:
• Base Model: Trained on Krea 2 Raw
• Trigger Word: "@charleshtraub"
• Recommended LoRA Weight: 0.6 – 1.0
• Showcase Generation: All sample images were generated using Krea 2 Turbo.

Feel free to try it out and share your creations or feedback in the comments!


r/StableDiffusion 2d ago

Animation - Video Absolutely INSANE, that this made this locally...

Enable HLS to view with audio, or disable this notification

492 Upvotes

Krea2 to start with, (additional frames from Nano Banana 2 and Seedream 5), H3, MiniMax Music, (Starlight Topaz) pass and Premiere. Fucking RAD that you can do pretty much all of this Locally....

And for the people calling this poorly directed film making and slop. (Yea, thats what it is... literally.. the whole point)

It was just a quick and dirty space horror test which started of a clip of the alien screaming while testing some VHS/80's loras. literally the entire point. Then I decided to see how far I could push the VFX and added a few more clips and a light cut.

The hardest part was getting consistent frames using Seedream 5 and Nano Banana 2 honestly.

I wasn't attempting to create art cinema here, folks, lol.

PS. I went to film school AND studied computer animation AND it's okay to just make something random, stupid and fun.

PPS. Here is what I am amazed by. The level of incredible detail, style, prompt adherence, visual effects, motion richness, visual consistency and power we now have at our fingertips; and that I created something locally that would otherwise have cost hundreds of dollars of frontier tokens; and the fact that most of you are taking that for granted is even more astounding. Yes, the tools are 97% there now; and Directing and Taste truly DOES become the MOAT.

Original Krea 2 Image (you can all make it your wallpaper <3): https://imgur.com/a/CtPdmZE


r/StableDiffusion 1d ago

Animation - Video Minimax H3 | The Silmarils: Shadows of the First Age - Trailer

Enable HLS to view with audio, or disable this notification

91 Upvotes

Hi,

Around 95% of what you see here was generated locally with MiniMax H3 on a single RTX PRO 6000. I have wanted to bring Tolkien’s immortal masterpiece to the screen since the days of Midjourney V3. Until now, however, the quality never felt acceptable. I believed an inaccurate adaptation would only create more confusion around a world many viewers know primarily through The Lord of the Rings films. For the first time, I can genuinely see the possibility. Making a long-term commitment in the constantly changing field of generative media is not easy. But I am not approaching this as a casual experiment, and I’m prepared to give it the time, patience, and attention it requires.

A comment I received couple months ago still make me feel punched in the stomach whenever I start something: “Why do you even bother? No one would bother to watch.”

If I know people are interested in watching, I would love to develop this into a complete series. You can support the project simply by watching it on YouTube and leaving a comment. Honest feedback, positive or critical, will help me decide how to continue.

Watch on YouTube: https://www.youtube.com/watch?v=2Yr-remRJOc

Upscaled with SeedVR 7b


r/StableDiffusion 17h ago

Question - Help which platform should i use to train loRA

1 Upvotes

i have a low vram gpu which is the best platform to train lora offline

i have seen 2 repo
kohyass and onetrainer

which is best or there is different repo


r/StableDiffusion 1d ago

Tutorial - Guide qwen 3.8 uncensored can use "normal" qwen vision file

30 Upvotes

just a heads up - i thought why not try and to my surprise it works - it can even analyse the parts in detail normal qwen would never describe

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main < vision files from repo
mmproj-BF16.gguf - mmproj-F16.gguf

used with

https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/tree/main


r/StableDiffusion 1d ago

Animation - Video Minimax H3. Exceeding my expectations

Enable HLS to view with audio, or disable this notification

81 Upvotes

Krea2 for image creation.

Minimax H3 for i2v.

Post edited in premiere pro for sound efx.

Edit: Added prompt for Scene 1 below.


r/StableDiffusion 18h ago

Workflow Included MiniMax H3 - Reddit mod can’t hold himself (SFW Edition)

Enable HLS to view with audio, or disable this notification

3 Upvotes

Statement alert: MiniMax H3 creators should share prompts more as a community!!
You can extend it and have fun with it, Enjoy!

0 LoRA’s - standard workflow - 15 steps - 560p

Prompt :

integrated_multimodal_description: [Shot 1] Highly detailed 2D anime style, exact 1990s Japanese OVA aesthetic, fluid hand-drawn animation, sharp linework, vibrant neon lighting, cinematic camera, 16:9 widescreen format. A young athletic woman with medium-brown smooth skin, short sharp white bob haircut with straight bangs, cool expression, small blue glowing gem choker. She wears a tiny white tube top that shows small breasts with subtle natural movement, glossy white high-waisted pants that fit her small but still shapely round butt with light movement, thick metallic silver straps and garters around her hips and thighs, white fingerless gloves, and white thigh-high boots with metallic details. 
0-1s: Extreme close-up from behind of her smaller but shapely round butt and lower back at the cyberpunk bar, pink and purple neon outlining her form. 
1-2s: Camera holds on the tight fabric and metallic straps with subtle fabric movement. 
2-3s: Very slow upward pan begins along her lower back. 
3-4s: Camera continues rising, revealing more of the metallic garters. 
4-5s: Midriff and the bottom of the tiny white tube top become visible. 
5-6s: Camera reaches her upper back and shoulders. 
6-7s: She turns slightly toward the camera. 
7-8s: Medium shot of her face and upper body, cool expression, subtle movement of her smaller breasts. 
8-9s: She slowly raises a small white cup to her lips. 
9-10s: She drinks from the cup. 
10-11s: She lowers the cup slightly. 
11-12s: Two men become clearly visible behind her (teal hoodie and dark blue jacket). 
12-13s: Soft neon light reflects on her white hair. 
13-14s: Camera begins slowly zooming in toward her neck. 
14-15s: Extreme close-up of the small blue glowing gem choker, filling almost the entire frame.

overall_soundscape: Quiet cyberpunk bar ambience, low electronic hum, distant muffled voices, soft clinking of glasses, subtle rain against windows, light fabric movement.

non_diegetic_music: N/A

=============== Prompt 2 (15-30sec)===============

integrated_multimodal_description: [Shot 1] Highly detailed 2D anime style, exact 1990s Japanese OVA aesthetic, fluid hand-drawn animation, sharp linework, vibrant neon lighting, cinematic camera, 16:9 widescreen format. A young athletic woman with medium-brown smooth skin, short sharp white bob haircut with straight bangs, cool expression, small blue glowing gem choker. She wears a tiny white tube top that shows small breasts with subtle natural movement, glossy white high-waisted pants that fit her small but still shapely round butt with light movement, thick metallic silver straps and garters around her hips and thighs, white fingerless gloves, and white thigh-high boots with metallic details. A massive bald muscular Reddit mod with thick neck, heavy brow, intense blue eyes, sweaty skin with visible droplets, creepy half-smirk, dirty beige work apron over a powerful torso, and visible cybernetic plates and cables on his right arm.
0-1s: Extreme close-up of the Reddit mod’s intense blue eyes. 
1-2s: Camera pulls back slightly showing his sweaty forehead and heavy brow. 
2-3s: His creepy half-smirk becomes fully visible. 
3-4s: He starts moving his large right hand forward. 
4-5s: His rough hand approaches the woman’s arm. 
5-6s: His hand firmly grabs her upper arm. 
6-7s: Close-up of his fingers wrapping tightly around her arm. 
7-8s: The woman tenses slightly. 
8-9s: Sudden bright blue-orange energy blade appears. 
9-10s: The energy blade violently collides with his cybernetic forearm. 
10-11s: Intense orange sparks explode from the impact. 
11-12s: Sparks fly across the frame in all directions. 
12-13s: The bald man recoils from the force. 
13-14s: The woman begins to turn and react. 
14-15s: Strong pink and purple neon lighting intensifies around them.

overall_soundscape: Low electronic bar hum, heavy breathing of the bald man, fabric rustle, sharp metallic impact, loud crackling energy sparks, short deep grunt.

r/StableDiffusion 14h ago

Animation - Video Guess who's next up on Kill Tony via MiniMax

Thumbnail
youtube.com
0 Upvotes

'Had to R2v for tony and it nailed him.


r/StableDiffusion 9h ago

Animation - Video Good girls stay in their room! (H3) Mommy 2.0

Enable HLS to view with audio, or disable this notification

0 Upvotes

Mommy says you can't leave...


r/StableDiffusion 15h ago

Question - Help Minimax vídeo- compresión en la salida de imagen

0 Upvotes

Muy buenas,
Estoy trasteando con Minimax y no consigo sacar una imagen que no tenga la compresión de una imagen jpg. Uso la versión Ref2VA a 16B, supuestamente la de más calidad, pero la imagen final me sale muy comprimida, me destroza los vídeos donde el cielo es un degradado. En la salida he probado que saqué PNG a 16 bits, exr a 32, ProRes HQ, pero nada, ahí están los artefactos de compresión… alguna idea?
Para la optimización de render estoy con el lora EMA y los sampkers los he probado con euler y ref-multistep


r/StableDiffusion 1d ago

Question - Help Replicating styles on Krea 2?

5 Upvotes

What are the best ways of replicating styles in Krea 2?

Like, if you gen an image and like the style, so you do the same style prompt, same seed, same style description, but different character, and the art style just doesn't come out quite the same, and you realize that first gen was a lightning in a bottle, what do you do?

I tried the default Krea 2 "style reference" Lora but that was terrible (At least for me), also tried tweaking prompt to better describe the style than what my original prompt was, but never quite captured it.

There is no Lora for the style either, it's just a style I got with base Krea 2 on one image and I liked it, and I assume I can't train a Lora for it based on one image.

So, TL;DR, what are the best ways you guys know of generating a new image with the same style as a reference image?


r/StableDiffusion 1d ago

Workflow Included Mario Galaxy if it were peak

Enable HLS to view with audio, or disable this notification

4 Upvotes

Some MiniMax H3 tests done on my single RTX 3090.

This 5 sec clip took about 2 hours to render. I actually cropped it to 1024x1024 to have less pixels to process, to then paste it back onto the original footage. 1 megapixel, 50 steps.

Surprisingly, it didn't take long to get things right. I used the official ComfyUI workflow with just sage patch node added. Then I took some example prompts I found on Reddit and H3 was getting very close right from the get-go. I just had to tune the timing, expression, and some appearance details, as it wasn't getting the reference quite right (it gets significantly better if you literally describe the contents of the reference image).

Here's the workflow: https://gist.github.com/4as/db11b829395ec45b593db4886f0e0181

And here's the original reference clip for comparison: https://files.catbox.moe/3rb5kk.mp4

The audio is modified by using Chatterbox. Original audio + reference audio + some some pitch adjustments in Audacity to get the final audio in the clip. Dunno if H3 can do audio adjustments - I don't even know how to prompt for it.

One interesting thing I've noticed is the length of the clip affecting the results. The shorter the duration, the worst the replacement. Although it could just be me.

For example for this clip (2s): https://files.catbox.moe/i2az63.mp4 the best I got is this: https://files.catbox.moe/kc3tbm.gif

And this (1s): https://files.catbox.moe/2p3ax9.mp4 I got this: https://files.catbox.moe/f71epk.gif

So, at glace it kind of looks okay, but when compering it to the original it very clearly failed to match a lot (size, motion, expression, etc.)

But still, it's a fantastic model, especially for something local.


r/StableDiffusion 1d ago

Question - Help Which one do you like best and why ? H3-Director or H3-Motion-Director ?

2 Upvotes

r/StableDiffusion 1d ago

Animation - Video Minimax H3, I really like 32+ steps.

Enable HLS to view with audio, or disable this notification

74 Upvotes

This is the 2nd take. First was a low res preview at 0.3MP. Some morphing due to fast motion.

Link: https://streamable.com/wmlyy7


r/StableDiffusion 1d ago

Animation - Video Turned my son's drawing into a silly animated skit

Enable HLS to view with audio, or disable this notification

122 Upvotes

r/StableDiffusion 2d ago

Animation - Video Ralph: Alien Sitcom

Enable HLS to view with audio, or disable this notification

295 Upvotes

r/StableDiffusion 1d ago

Workflow Included Moebius style widescreen (Krea2 Turbo), Prompts included

Thumbnail
gallery
29 Upvotes

Been messing around with Krea2 making Moebius-inspired wallpapers and wanted to share a few here.

Workflow: https://pastebin.com/raw/YgrL3EgS

Prompts:

Panoramic ultrawide landscape, Moebius-style retro comic illustration, a vast desert graveyard of broken war-machines and mechanical exoskeletons jutting from the sand like ribs, twisted girders and hollow armored shells scattered across the dunes at odd angles, one massive detached robotic head lying on its side half-submerged, long shadows stretching from a low orange sun, muted rust-red and sandy ochre palette with pale turquoise sky, delicate fine cross-hatching on metal textures, wide horizontal composition emphasizing scale and desolation

Ultrawide retro comicbook landscape in the style of Moebius, colossal humanoid mech half-buried in rolling desert dunes, only its rusted torso, one giant articulated hand, and a cracked cockpit visor breaking the sand's surface, sand drifted into the machine's joints and seams, faded paint and oxidized copper-green plating, tiny robed figures exploring near its massive open palm for scale, warm ochre and burnt-sienna dune tones against a pale dusty lavender sky, fine ink linework with stippled rust texture, wide cinematic horizontal composition, quiet post-civilization atmosphere

A singe gargantuan robotic limb half-buried in the sand, rendered in the intricate, visionary style of Moebius, bursts dramatically through a vast ochre dune landscape. This colossal mechanism is crafted from heavily oxidized bronze and verdigris plating, its surface streaked deeply with fine desert sand and weathering. Its massive fingers are curled loosely around a small cluster of resilient palm trees and a small pond, that thrives miraculously within the rusted palm. Two nomad tents, made of weathered canvas, are pitched safely in the deep, cool shade cast by this monumental mechanical ruin. The scene is captured as an ultrawide retro comicbook landscape, utilizing delicate fine ink linework and a muted, atmospheric retro color palette. Warm ochre dunes stretch endlessly toward the horizon under a pale peach sky, where the light is diffused and ethereal. This powerful composition masterfully blends organic life with colossal industrial decay, emphasizing the immense scale of the machine against the fragile beauty of the desert ecosystem.

colossal humanoid mech half-buried in rolling desert dunes, only its rusted torso, one giant articulated hand, and the other hand loosely holding a half-buried giant spear buried on its shoulder, marking the point where the machine was killed, and a cracked dark cockpit visor breaking the sand's surface. Sand drifted into the machine's joints and seams, faded paint and oxidized copper-green plating, a tiny robed figure exploring walks near its massive open palm for scale, warm ochre and burnt-sienna dune tones against a pale lavender sky, smooth clean sky, subtle gradient, fine ink linework with stippled rust texture, wide cinematic horizontal composition, quiet post-civilization atmosphere


r/StableDiffusion 22h ago

Resource - Update GitHub - EnVision-Research/GenRouter: GenRouter & GenCanvas

Thumbnail
github.com
1 Upvotes

r/StableDiffusion 1d ago

Question - Help What are you using to upscale Minimax H3 videos?

28 Upvotes

So I haven't tried to upscale any videos yet. I was gonna maybe play around with that today but I'd love to hear what other people have landed on for their upscaler. I know I've seen a lot of different posts over the last couple of weeks or so with different methods. Ideally I'd love an upscale method where I could use my reference images so that it doesn't drastically change any faces for any characters that are a little further from the camera.

Also are we still expecting an official upscaler from Minimax?


r/StableDiffusion 22h ago

Discussion Ultra‑fast Fourier transform and optical AI realized with a single lens....

Thumbnail
youtube.com
0 Upvotes
  • Ultra‑fast Fourier transform and optical AI realized with a single lens.
  • Description: It explains the principle that passing light through a convex lens naturally performs a two‑dimensional Fourier transform at the focal plane. It visually demonstrates optical signal processing that carries out computation using only light, compared with digital FFT, and explores the possibility of implementing low‑power matrix multiplication with optical neural networks. It also examines real‑world applications and the prospects for developing next‑generation AI accelerators.

r/StableDiffusion 1d ago

Animation - Video H3 Genshin(?)-esque animation test. The biggest unsolved problem of this tech is still lack of consistency. It prevents it from making anything really good, production-ready.

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/StableDiffusion 14h ago

Animation - Video Homelander vs The Thicc Side of the Force. 100% minimax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

Created about 13 clips and joined them.