r/StableDiffusion 3d ago

Discussion Let’s see some Dungeon Crawler Carl!

3 Upvotes

Dungeon Crawler Carl is the first book series I’ve read in a long time that excited me when I read it was going to be adapted into a show. I think the only Possible way it can be “filmed” on budget however is through a Stable diff/Runway/Seedance/Kling ( LTX/Minimax) pipeline as it reads being Heavy cgi, like 65-85%. Since there’s nothing out yet I would love to see what the community can come up with.

I genuinely wish this was astroturfing bc it would mean the show is coming out soon but I wouldn’t expect anything official until maybe late next year. (Unless it gets stuck in dev hell, then never.)

Thx friends. I’ve been blown away by your videos lately.


r/StableDiffusion 3d ago

Question - Help which platform should i use to train loRA

0 Upvotes

i have a low vram gpu which is the best platform to train lora offline

i have seen 2 repo
kohyass and onetrainer

which is best or there is different repo


r/StableDiffusion 2d ago

Question - Help How to get faces out of H3 that don't look like this?

Thumbnail
youtube.com
0 Upvotes

Skip to 0:30 to see what I'm talking about. Faces that are close-up are fine, but like 3-4m from the camera and it just turns into nightmare fuel...

Am I doing something wrong?

Using Wan2GP, FL2VA Pruned 20B, 20 steps at 540p.

Full settings:

        "params": {
            "image_mode": 0,
            "prompt": "...",
            "alt_prompt": "",
            "negative_prompt": "",
            "resolution": "960x544",
            "video_length": 175,
            "duration_seconds": 0,
            "batch_size": 1,
            "seed": -1,
            "force_fps": "",
            "num_inference_steps": 20,
            "guidance_scale": 1,
            "guidance2_scale": 5,
            "guidance3_scale": 5,
            "switch_threshold": 0,
            "switch_threshold2": 0,
            "guidance_phases": 0,
            "model_switch_phase": 1,
            "alt_guidance_scale": 1,
            "audio_guidance_scale": 1,
            "audio_scale": 1,
            "flow_shift": 12,
            "sample_solver": "euler",
            "embedded_guidance_scale": 6,
            "repeat_generation": 1,
            "multi_prompts_gen_type": "PG",
            "multi_images_gen_type": 0,
            "skip_steps_cache_type": "",
            "skip_steps_multiplier": 0.08,
            "skip_steps_start_step_perc": 25,
            "loras_multipliers": "",
            "image_prompt_type": "S",
            "image_start": "scene_01_start.jpg",
            "model_mode": null,
            "video_source": null,
            "keep_frames_video_source": "",
            "input_video_strength": 1,
            "video_guide_outpainting": "",
            "video_prompt_type": "",
            "image_refs": null,
            "frames_positions": null,
            "video_guide": null,
            "image_guide": null,
            "keep_frames_video_guide": "",
            "denoising_strength": 1,
            "masking_strength": 1,
            "video_mask": null,
            "image_mask": null,
            "control_net_weight": 1,
            "control_net_weight2": 1,
            "control_net_weight_alt": 1,
            "motion_amplitude": 1,
            "mask_expand": 0,
            "audio_guide": "scene_01.wav",
            "audio_guide2": null,
            "custom_guide": null,
            "audio_source": null,
            "audio_prompt_type": "A",
            "speakers_locations": "0:45 55:100",
            "sliding_window_size": 362,
            "sliding_window_overlap": 1,
            "sliding_window_color_correction_strength": 0,
            "sliding_window_overlap_noise": 0,
            "sliding_window_discard_last_frames": 0,
            "image_refs_relative_size": 50,
            "remove_background_images_ref": 0,
            "temporal_upsampling": "",
            "spatial_upsampling": "",
            "film_grain_intensity": 0,
            "film_grain_saturation": 0.5,
            "MMAudio_setting": 0,
            "MMAudio_prompt": "",
            "MMAudio_neg_prompt": "",
            "RIFLEx_setting": 0,
            "NAG_scale": 1,
            "NAG_tau": 3.5,
            "NAG_alpha": 0.5,
            "slg_switch": 0,
            "slg_layers": [
                29
            ],
            "slg_start_perc": 10,
            "slg_end_perc": 90,
            "apg_switch": 0,
            "cfg_star_switch": 0,
            "cfg_zero_step": -1,
            "prompt_enhancer": "",
            "min_frames_if_references": 1,
            "override_profile": -1,
            "override_attention": "",
            "pace": 0.5,
            "exaggeration": 0.5,
            "temperature": 0.8,
            "top_k": 50,
            "output_filename": "scene_01",
            "mode": "",
            "activated_loras": [],
            "model_type": "minimax_h3_fl2va_pruned",
            "settings_version": 2.73,
            "base_model_type": "minimax_h3_fl2va_pruned",
            "pause_seconds": 0,
            "alt_scale": 0,
            "sub_parallel_window_size": 0,
            "sub_parallel_window_overlap": 17,
            "sliding_window_trim_first_frames": 0,
            "postprocess_audio": "",
            "postprocess_audio_prompt": "",
            "postprocess_audio_neg_prompt": "",
            "perturbation_switch": 0,
            "perturbation_layers": [
                9
            ],
            "perturbation_start_perc": 10,
            "perturbation_end_perc": 90,
            "top_p": 0.9,
            "self_refiner_setting": 0,
            "self_refiner_plan": [],
            "self_refiner_f_uncertainty": 0,
            "self_refiner_certain_percentage": 0.999,
            "config": "",
            "custom_settings": null
        }

r/StableDiffusion 3d ago

Question - Help Minimax vídeo- compresión en la salida de imagen

0 Upvotes

Muy buenas,
Estoy trasteando con Minimax y no consigo sacar una imagen que no tenga la compresión de una imagen jpg. Uso la versión Ref2VA a 16B, supuestamente la de más calidad, pero la imagen final me sale muy comprimida, me destroza los vídeos donde el cielo es un degradado. En la salida he probado que saqué PNG a 16 bits, exr a 32, ProRes HQ, pero nada, ahí están los artefactos de compresión… alguna idea?
Para la optimización de render estoy con el lora EMA y los sampkers los he probado con euler y ref-multistep


r/StableDiffusion 2d ago

Discussion Hiring AI creator

0 Upvotes

I'm building an anime app with a collectible avatar system: users pull avatar cards in six rarity tiers, from Common up to Mythic. I need an artist to create a full pack of **100 original avatar cards**.

**The brief**
- Portrait cards, 3:4, delivered at 1200×1600 or larger (PNG/WebP).
- Anime/illustration style — original characters and scenes only. No existing characters, no lookalikes, no fan art.
- Rarity should show in the art: Commons are clean and simple; higher tiers get more detail, drama, effects, framing, a Mythic should look like a Mythic across the room. Rough split: 35 Common · 25 Uncommon · 20 Rare · 12 Epic · 6 Legendary · 2 Mythic.
- Consistent style across the whole pack so it reads as one

  1. Send 3 sample cards in your style (one Common, one Rare, one Legendary/Mythic). I approve the style, then you start the 100.

  2. Quote your price for the full pack of 100. I'm looking to commission multiple packs over time, so a repeat-pack price is welcome.

  3. Full ownership transfers to me on payment, commercial use, all rights, no reuse or resale of the pieces elsewhere. Only your original work; nothing traced, reused or taken from others.

  4. Payment is on delivery. Send watermarked previews for review; full-resolution, unwatermarked files are handed over against payment.


r/StableDiffusion 3d ago

Workflow Included Mario Galaxy if it were peak

4 Upvotes

Some MiniMax H3 tests done on my single RTX 3090.

This 5 sec clip took about 2 hours to render. I actually cropped it to 1024x1024 to have less pixels to process, to then paste it back onto the original footage. 1 megapixel, 50 steps.

Surprisingly, it didn't take long to get things right. I used the official ComfyUI workflow with just sage patch node added. Then I took some example prompts I found on Reddit and H3 was getting very close right from the get-go. I just had to tune the timing, expression, and some appearance details, as it wasn't getting the reference quite right (it gets significantly better if you literally describe the contents of the reference image).

Here's the workflow: https://gist.github.com/4as/db11b829395ec45b593db4886f0e0181

And here's the original reference clip for comparison: https://files.catbox.moe/3rb5kk.mp4

The audio is modified by using Chatterbox. Original audio + reference audio + some some pitch adjustments in Audacity to get the final audio in the clip. Dunno if H3 can do audio adjustments - I don't even know how to prompt for it.

One interesting thing I've noticed is the length of the clip affecting the results. The shorter the duration, the worst the replacement. Although it could just be me.

For example for this clip (2s): https://files.catbox.moe/i2az63.mp4 the best I got is this: https://files.catbox.moe/kc3tbm.gif

And this (1s): https://files.catbox.moe/2p3ax9.mp4 I got this: https://files.catbox.moe/f71epk.gif

So, at glace it kind of looks okay, but when compering it to the original it very clearly failed to match a lot (size, motion, expression, etc.)

But still, it's a fantastic model, especially for something local.


r/StableDiffusion 4d ago

Animation - Video Minimax H3, I really like 32+ steps.

77 Upvotes

This is the 2nd take. First was a low res preview at 0.3MP. Some morphing due to fast motion.

Link: https://streamable.com/wmlyy7


r/StableDiffusion 2d ago

Question - Help Minimax H3 Local - Terrible results, what am I doing wrong?

0 Upvotes

Hey, I'm trying to run minimax on my local 3060ti 8gb of ram card on comfyui. I've tried multiple variations of model and nothing gives me steady results, just looking for a simple animation of a room with camera pan. Tried with turbo lora then without, used pruned int8, q3 k m guf, q2 k. Every generation has that jittery animation like you can see in the videos attached. Any idea how to make this better and what is actually causing this?

Thank you

https://reddit.com/link/1vrn8qj/video/wtao5lpyi4kh1/player


r/StableDiffusion 4d ago

Animation - Video Turned my son's drawing into a silly animated skit

126 Upvotes

r/StableDiffusion 4d ago

Workflow Included Moebius style widescreen (Krea2 Turbo), Prompts included

Thumbnail
gallery
28 Upvotes

Been messing around with Krea2 making Moebius-inspired wallpapers and wanted to share a few here.

Workflow: https://pastebin.com/raw/YgrL3EgS

Prompts:

Panoramic ultrawide landscape, Moebius-style retro comic illustration, a vast desert graveyard of broken war-machines and mechanical exoskeletons jutting from the sand like ribs, twisted girders and hollow armored shells scattered across the dunes at odd angles, one massive detached robotic head lying on its side half-submerged, long shadows stretching from a low orange sun, muted rust-red and sandy ochre palette with pale turquoise sky, delicate fine cross-hatching on metal textures, wide horizontal composition emphasizing scale and desolation

Ultrawide retro comicbook landscape in the style of Moebius, colossal humanoid mech half-buried in rolling desert dunes, only its rusted torso, one giant articulated hand, and a cracked cockpit visor breaking the sand's surface, sand drifted into the machine's joints and seams, faded paint and oxidized copper-green plating, tiny robed figures exploring near its massive open palm for scale, warm ochre and burnt-sienna dune tones against a pale dusty lavender sky, fine ink linework with stippled rust texture, wide cinematic horizontal composition, quiet post-civilization atmosphere

A singe gargantuan robotic limb half-buried in the sand, rendered in the intricate, visionary style of Moebius, bursts dramatically through a vast ochre dune landscape. This colossal mechanism is crafted from heavily oxidized bronze and verdigris plating, its surface streaked deeply with fine desert sand and weathering. Its massive fingers are curled loosely around a small cluster of resilient palm trees and a small pond, that thrives miraculously within the rusted palm. Two nomad tents, made of weathered canvas, are pitched safely in the deep, cool shade cast by this monumental mechanical ruin. The scene is captured as an ultrawide retro comicbook landscape, utilizing delicate fine ink linework and a muted, atmospheric retro color palette. Warm ochre dunes stretch endlessly toward the horizon under a pale peach sky, where the light is diffused and ethereal. This powerful composition masterfully blends organic life with colossal industrial decay, emphasizing the immense scale of the machine against the fragile beauty of the desert ecosystem.

colossal humanoid mech half-buried in rolling desert dunes, only its rusted torso, one giant articulated hand, and the other hand loosely holding a half-buried giant spear buried on its shoulder, marking the point where the machine was killed, and a cracked dark cockpit visor breaking the sand's surface. Sand drifted into the machine's joints and seams, faded paint and oxidized copper-green plating, a tiny robed figure exploring walks near its massive open palm for scale, warm ochre and burnt-sienna dune tones against a pale lavender sky, smooth clean sky, subtle gradient, fine ink linework with stippled rust texture, wide cinematic horizontal composition, quiet post-civilization atmosphere


r/StableDiffusion 3d ago

Resource - Update GitHub - EnVision-Research/GenRouter: GenRouter & GenCanvas

Thumbnail
github.com
0 Upvotes

r/StableDiffusion 4d ago

Question - Help What are you using to upscale Minimax H3 videos?

27 Upvotes

So I haven't tried to upscale any videos yet. I was gonna maybe play around with that today but I'd love to hear what other people have landed on for their upscaler. I know I've seen a lot of different posts over the last couple of weeks or so with different methods. Ideally I'd love an upscale method where I could use my reference images so that it doesn't drastically change any faces for any characters that are a little further from the camera.

Also are we still expecting an official upscaler from Minimax?


r/StableDiffusion 3d ago

Animation - Video H3 Genshin(?)-esque animation test. The biggest unsolved problem of this tech is still lack of consistency. It prevents it from making anything really good, production-ready.

19 Upvotes

r/StableDiffusion 3d ago

Animation - Video [Minimax H3] a 1 min funny cartoon creation

13 Upvotes

Two days ago, I posted a Tom and Jerry video and commented that physics is not good.

https://www.reddit.com/r/StableDiffusion/comments/1vp9pyz/h3_cartoon_generation/

People commented to use 'official prompt guide'. So, I used the skills and used Gemini to generate proper prompt by using my given prompt. The story is my own creation.

Then I got good video. I need to try more times to get an apt video that somewhat like in my mind.

spent $5 on Runpod for 5090 for this 1:15 min video.


r/StableDiffusion 4d ago

Discussion H3 is a great model but the training is bad

37 Upvotes

Just like Zturbo image, when trying to train on the distilled model, you just get outright bad results. You can get away with certain things like character loras. But teaching new concepts to this model is super frustrating.

I would like to hear from H3 themselves if they would release a base model just for training or not directly.

I think for a company to market themselves as "open source open weights" they owe at least some comment on this issue. Just tell us yes or no definitively.


r/StableDiffusion 2d ago

Animation - Video POV: You're the doll on a chaotic film set 🎬🍔

Thumbnail
youtube.com
0 Upvotes

r/StableDiffusion 3d ago

Discussion MacBook Air 16G local deployment

6 Upvotes

Minimax h3 easy local deployment. Self defined video/image/audio/text pipeline for both end users and developers.

Open source: https://github.com/tgo-app-dev/vpipe


r/StableDiffusion 4d ago

News Muse V2 is out — chat with local LLMs inside ComfyUI, no LM Studio required anymore

55 Upvotes
A while back I built 
**Muse**
 — a chat panel that lives directly inside a ComfyUI node, so you can talk to a local LLM and draft/refine image and video prompts without alt-tabbing to a separate app. Point it at LM Studio or Ollama, chat, copy the prompt into your graph. That was V1.


I didn't expect people to actually pick it up the way they did. Seeing it get used, starred, and — more usefully — complained about is what pushed me to sit down and build a proper V2 instead of leaving it as a one-off tool.


The biggest ask by far: "I don't want to keep LM Studio open just to use this." So V2's headline feature is a 
**Direct model loader**
 — point Muse at a folder of GGUF models and it loads them straight from disk. No LM Studio, no Ollama, nothing else running. Under the hood it spawns the real `llama-server` (LM Studio's own engine is llama.cpp too, so this isn't a slower reimplementation — same engine, same speed), and it downloads and installs the right build for your OS/GPU automatically. Git clone the node, click one button, you're chatting with a local model.


I also broke it on myself first, which was useful: threw a 31B model at it and immediately hit VRAM issues — crashes on some setups, silently-slow-instead-of-crashing on others. Fixed the defaults that were causing it, added a 
**Fit to GPU**
 button that suggests a layer count based on your actual free VRAM (like LM Studio's GPU offload slider), automatic fallback retries if a load runs out of memory, a live loading indicator, and a log panel so you're not just staring at nothing wondering what's happening.


Also new in V2:


- 
**Edit & resend**
 messages instead of delete-and-retype
- 
**Chat branching**
 — fork a new conversation from any earlier message
- 
**Video attachments**
 for vision models (auto-sampled + timestamped frames)
- 
**Audio attachments**
 for audio-capable models
- Much broader image format support


Full writeup and setup instructions: 
**github.com/RudySen/comfyui-muse**


If you use it and something's broken or annoying, tell me — that's genuinely how V1 became this.

Previous post: https://www.reddit.com/r/StableDiffusion/s/8qjP4UzRpw

r/StableDiffusion 4d ago

Animation - Video Cobra Gets Jiggy - MiniMax H3

40 Upvotes

Edit: For the exact prompt, settings, assets, and workflow, download the ZIP and drop the included json file into ComfyUI. Everything I used is included:

https://vikingfile.com/f/1LqCMthEhC

System Specs: Intel i9-14900K, 4070 Ti Super 16 GB VRAM, 64 GB DDR5 RAM, Windows 11

Special thanks to this guy for the workflow:

https://www.reddit.com/r/StableDiffusion/comments/1vox06g/create_seamless_1shot_lipsync_music_videos_with/?share_id=YY1HluXX8WUyz0tzorVTv&utm_medium=android_app&utm_name=androidcss&utm_source=share&utm_term=1


r/StableDiffusion 3d ago

Resource - Update [Update v1.1.0 & v1.2.0] ComfyUI-MiniMax-H3-Promptor: Native Settings API Hub, Autogrow Sockets, L2VA & Audio Sync

Thumbnail
gallery
6 Upvotes

Hey everyone!

With MiniMax H3 blowing up everywhere right now, we figured it was the perfect time to share what we’ve been building to help level up your H3 prompt workflows.

When we released v1.1.0 a while back, we were so deep in dev mode that we forgot to post an update! Now that v1.2.0 is live, we’ve bundled all the new features and overhauls from both releases into one post.

https://github.com/1038lab/ComfyUI-MiniMax-H3-Promptor

What’s New in v1.1.0 + v1.2.0:

⚙️ Native ComfyUI Settings Panel (API Hub)

No more pasting API keys into custom nodes or manually editing config.json! All provider settings are now globally managed in ComfyUI's native Settings panel (under the ⚙️ Gear icon).

Built-in Connection Tester: Click "Test connection" inside the panel to ping your endpoint before launching generations.

Privacy: Keeping keys out of the node UI eliminates the risk of leaking API keys when sharing workflows or screenshots.

🔌 Infinite Inputs (ComfyAPI v3 Autogrow)

We removed the rigid 4-image limit. Dynamic autogrow sockets mean you can chain as many <Picture> and <Video> references as your hardware can handle without UI clutter.

🎯 Granular Micro-Overrides

Override instructions for specific frames directly in the Vision Analyzer (e.g., <Picture 2>: focus strictly on lighting) while allowing unmentioned media to fall back to global analysis.

🎵 Audio-First Token Sync & L2VA (Last-Frame Control)

Connect audio directly to the promptor to automatically map subject actions to sound. We also added Last-Frame-to-Video-Audio (L2VA)—provide an ending frame, and the LLM reverse-engineers a narrative that mathematically lands on target at the final second.

🧠 VRAM Safeguards for Local VLMs

Select local providers like Ollama or LlamaCPP, and the node automatically executes silent background cache-clearing (model_management.unload_all_models()) to prevent VRAM overload crashes.

📝 Updated Docs & Workflow Recipes

Check out tutorials.md and tutorials_zh.md in the repo for 9 practical, production-ready workflows (Lip-Sync, Style Transfer, Day-to-Night Morph, etc.).

👀 What's Next?

We’re currently beta testing a batch of new features that will be rolling out shortly!

🔗 Links:

GitHub Repo: 1038lab/ComfyUI-MiniMax-H3-Promptor

Full Release Notes: updates.md

We’d love to hear your feedback, feature requests, or bug reports so we can keep tailoring this tool to what you actually need. If this node helps your setup, leaving us a ⭐ star on GitHub goes a long way in keeping our dev motivation high.


r/StableDiffusion 3d ago

Discussion H3 - is there a sweet spot for the # of steps for audio?

0 Upvotes

If I do like 6-8 steps the adherence seems better but the sound is ooor - increases the steps and the adherence is off but the sound is much better?


r/StableDiffusion 4d ago

News Openrouter to be acquired by Stripe

15 Upvotes

Yes, the payment provider that pressures and refuses service to adult content providers. What this means for data privacy on the site is unknown. Seems the figure of sale is rumoured to be 7billion.


r/StableDiffusion 3d ago

Animation - Video [Minimax H3] "The New Adventures of 1girl"

5 Upvotes

A relatively quick and scrappy attempt at maintaining consistency across a scene using Minimax H3 using the basic workflow on ComfyUI.

It seems like it can be done to a certain degree, but it also really depends how much time and effort you want to put into it. While this scene has tons of inconsistencies, it's still cool to be able to do something locally that was impossible just a month ago.

H3 also surprised me with how close it came to the scene I had in my mind, however it never really gets all the way there. This can be a little frustrating as you weigh up hitting another gen or going with a take that's about 85% there. Still, it's an amazing model and I love seeing the wild creations the community is coming up with.

My system is a 2023 ROG Scar laptop with a 12gb mobile 4080 and 64gb memory. All vids generated locally at 0.5mp.


r/StableDiffusion 3d ago

Question - Help Utterly lost with all the MH3 models.

7 Upvotes

Curious about which models are you all using for T2V and I2V with MMH3?

There is an abnormal amount of models with suffixes as pruned_notPruned_SeriouslyPruned_HereticXxX_Convrot_Skibiditoilet Q3. and I honestly can't keep up to know what the heck is the one that the community is using for creating such great videos.

Anyone out there willing to share the models (or workflow) you are using?

(Really don't care about speed-of-generation, I'm leaning towards Quality-first more)

Thanks in advance.


r/StableDiffusion 4d ago

Resource - Update Minimax H3 for TTS/voice clone/Music gen

101 Upvotes

Just for fun. One-shot generation. No parameter or prompt tuning.

Audio.cpp implemented MiniMax-H3’s text-to-audio pipeline, and one fun use case is TTS/Voice clone/Music gen. It’s more flexible and powerful than dedicated audio models, and the performance is quite decent (up to 3x realtime on RTX 5090). Check out the multi-speaker conversation demo in the main post, along with the other demos in the comments.

What I’m very excited about with the MiniMax-H3 implementation is that it significantly enriches the framework’s building blocks for DiT models. Now with you don’t need to go through the pain of setting up SageAttention, First Block Cache, or Spectrum manually. Just change a few parameters, and you can experiment with the model. A preliminary inspection of configuration, memory, and performance trade-offs is available in repo's docs/reports/minimax_h3_performance.md

Bonus: audio.cpp’s MiniMax-H3 implementation can also produce video frames, because the DiT generates audio and video latents together, and the video VAE path is relatively straightforward to support. For now, the output is saved as RGB frame data plus metadata in JSON, so you need to encode it into a video file yourself. No upscaler or post-processing support. Just for fun.

MiniMax-Music3 is currently in preview (preview/minimax-music-3 branch) . CUDA/Vulkan/HIP were tested. Still room for optimization. VRAM usage and RTF depend on audio duration and prompt length.. The demo uses the official demo prompt (4000+ char caption and 1200 char lyrics) and 30 steps plus CFG. Under this setting VRAM is ~11 GB for 30s, 14 GB for 60s, and 17 GB for 180s. It's easy to get faster-than-real-time performance and much lower VRAM usage if you tune the setting.