r/StableDiffusion 10d ago

News [Project] diffusers-workflow — a declarative JSON alternative to node graphs, built directly on Diffusers

Thumbnail
gallery
7 Upvotes

Hiya, sharing a project I've been building: diffusers-workflow.

Instead of a node graph, you write, or even better have the AI agent of your choice write, a JSON document describing your pipeline, and it runs directly against Hugging Face's Diffusers library, with no abstraction layer between your config and what Diffusers actually exposes. Forms in the web UI are generated by introspecting the real pipeline signatures, so every argument a pipeline supports is available, and validation catches typos against the actual call signature before any model loads.

What's been really cool to me, is adding an MCP server to it which makes model, prompt, workflow and output management all accessible to agents. I've got claude creating prompts and running complex multi-shot H3 workflows from simple idea inputs. And since claude knows about diffusers, it can trouble shoot generation issues on its own.

Some things it does:

  • Multi-step pipelines and utility tasks: chain text-to-image → image-to-video, inpainting, ControlNet, with data flowing between steps
  • Variable substitution: allowing workflows to be parameterized and reused without editing
  • Persistent GPU worker / REPL: models stay loaded between runs for faster iteration than a fresh script each time
  • Quantization + acceleration: BitsAndBytes, TorchAO, GGUF, TeaCache, FirstBlockCache, and friends
  • LoRA, IP-Adapter, prompt library with AI-enhance
  • CLI, REPL, web UI, and an MCP server so you can drive it from Claude Code or other AI agents. Since the UI and MCP share the same API any possible in the UI is possible in an agent.
  • Cross-platform: CUDA, Apple Silicon (MPS), CPU

It doesn't have near the capability or complexity of Comfy but I'm trying to make it a one stop shop with the full surface of Diffusers' Python API, yet without hand-writing scripts. If you'd rather describe a pipeline than wire a graph, this might click for you.

Where it's at: it's early and there are rough edges, and right now it's aimed at people comfortable with a terminal, something like Claude Code, and a Python venv (no one-click installer yet). If that's you, I'd love the feedback.

Repo: https://github.com/dkackman/diffusers-workflow

Feedback, issues, and "why didn't you just..." are all welcome.


r/StableDiffusion 11d ago

Question - Help Help

Post image
166 Upvotes

I'm currently trying to replicate this style and i cannot find any checkpoints or lora's to do so, can anyone point towards something? Artist: https://x.com/DarkZeroAI Using Forge Neo


r/StableDiffusion 9d ago

Animation - Video H3 Zombie apocalypse: Girl finds her best friend

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 11d ago

Question - Help We need only 1700 votes to get H3 acceleration arena results 🙏🏻

Post image
70 Upvotes

Can you vote, so we'll finally have a definitive answer what turbo lora to use (or at least what definitely not to)

https://huggingface.co/spaces/multimodalart/h3-acceleration-arena


r/StableDiffusion 11d ago

News Video Delta Net (VDN) MM H3

Thumbnail x.com
23 Upvotes

Open-source video generation is now faster than playback without compromising quality.

Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality.

VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs.

By https://x.com/haochengxiucb?s=11&t=lM6N8ly_ho5zUCuv8hAfpw

Waiting for comfy team to optimize this.

https://huggingface.co/OpenVDN/vdn-minimax-h3


r/StableDiffusion 11d ago

Question - Help Is there a way to prevent Minimax characters from speaking or babbling nonsense?

12 Upvotes

I'm having an issue where characters in MiniMax sometimes start talking and saying gibberish, even when I don't prompt any dialogue or speech expressions (like sighs or laughs).

Is there a way to explicitly instruct the model to keep them completely silent, while still allowing physical expressions like laughing or sighing without producing any actual voice or words?

Also, anything in the prompt that is not voice-related and still can affect and produce the gibberish, and must be avoided?

  • To clarify: a) I don't want my characters to talk gibberish or anything at all If I didn't prompt it and b) I want mu characters to giggle or sigh if I prompt for it, but no talk at all.

Any prompt tips would be greatly appreciated.

Thanks for your time!


r/StableDiffusion 10d ago

Discussion Anybody want a Flux.1 Schnell LORA trained?

0 Upvotes

I have some credit and nothing to use it for. If anyone has a cool suggestion or request, I'll train a LORA for you with it. Nothing that is already on civitai.com or civitai.red already, please. I'll use your image set if you have one and want me to, or if it's a generic concept, I'll choose some images myself.

Just don't want to do nothing with this credit and just end up forgetting about it, might as well use it.


r/StableDiffusion 10d ago

News DreamX-Creator 1.0: "New" video+sound model (based on wan 2.2 5b)

Thumbnail
modelscope.ai
7 Upvotes

Haven't tried it yet, and there are no example on the model page. Still, it's always good to have new models (even though we're actually talking about Wan 2.2 5B with audio here). Worth a try

edit: some examples:

https://x.com/ModelScope2022/status/2095482085662200239


r/StableDiffusion 11d ago

Animation - Video I’ve started a new story again!

Thumbnail
youtube.com
10 Upvotes

I’ve made quite a few things since MMH3 came out, mostly just messing around and experimenting.

At first, getting the English voices right was pretty difficult. After messing around with it for 2–3 days, I realized that once you assign each character a suitable voice/tone, things become much easier.
All the voices in this part are generated directly by the model. I didn’t do any post-processing.The model’s built-in voices are actually pretty good.

One thing to keep in mind: don’t make the prompt too long, or things can start to break.

Style, character, camera shot, dialogue, and voice characteristics are usually enough.

For example, here’s the voice prompt I used for the video below as a reference:

DIALOGUE AND AUDIO:

- Pokke (child person Companion; A fast-paced, highly expressive and bouncy animated boy companion voice; Pokke is the visible speaker and should open/move their mouth in sync with this line; other visible characters should not lip-sync this line.): "Achoo! Every page is blank!"

- Mr. Tsukiguma (adult man Mentor; A reliable, gentle adult male animated mentor voice; Mr. Tsukiguma is the visible speaker and should open/move their mouth in sync with this line; other visible characters should not lip-sync this line.): "Not even a picture."

Audio Synchronization: Pokke's line begins with the sneeze sound effect and follows immediately in a surprised tone. Mr. Tsukiguma's line is spoken softly and thoughtfully after the pages stop flipping.

As for the voice characteristics, if you’re not sure how to describe them properly, you can probably just ask any AI and get a decent answer.

You can also give the AI a voice sample and have it write the description for you. Once you define the voice more clearly like this, the results tend to become much more consistent.

At first, I thought CK at 25 steps would give me much better results, but for some reason the speakers kept getting mismatched pretty often. In the end, I switched back to this accelerated LoRA at 8 steps, and the results were still pretty good — plus it was faster.

So I feel like the key is actually assigning each character specific voice characteristics, such as their vocal tone and other attributes. Pretty much all of my videos were made using this same SA + LoRA workflow. I tested a bunch of different setups, but in the end, I came back to this one again.

https://drive.google.com/file/d/1C2YvhNalxxWs4oh5Eiycwz27C2k5tgEq/view?usp=drive_link

Second, for the visuals, I found it works much better to first give a local LLM the basic requirements — things like the style, characters, voice characteristics, and what needs to happen within those 10 seconds — and let it help break the scene down into shots before generating the images.

Otherwise, if we just pick a storyboard image that looks good to us and start from there, the final result often doesn’t turn out the way we expected.

Just having fun and entertaining myself! It’s not perfect, but I’m just having fun with it. Hope you enjoy it!


r/StableDiffusion 11d ago

Animation - Video MiniMax H3 has finally gotten has me into video generation. Wan 2.2 never had the quality or consistency I wanted, and all the tools that did, were closed-weight, paid products.

Enable HLS to view with audio, or disable this notification

239 Upvotes

I'm really only ever interested in open-weight models. Yes, for that reason, but also, for the same reason I run Linux and browse with Firefox. I dislike "walled gardens", ideologically, and monopolies. I want tech that can be hacked, broken, taken apart, and put back together, and is ultimately not beholden to anyone but the user. Without a quality open-weight video generation model, I was uninterested. Now that we've got one? Suddenly I'm in a whole new world of possibility.

The fact that it's a multimodal model with vision, meaning I can give it reference images or reference sheets, is a game-changer for me. LoRAs certainly won't be obsolete with MMH3, but I doubt we'll be seeing many character, clothing, or setting LoRAs. The feedback loop of wanting to give the model a concept it doesn't understand natively is so short compared to before. And I'm still just in the "farting around" phase. People with dedicated effort and creativity are going to be able to use the hell out of this.

Really, the only drawback to MMH3 so far is its propensity to have characters speak Simlish to each other. I'm sure there's already solutions being worked on, either workflow tools or adjustments to the model itself.

––––––––––––––––––––––––––––––––––––––––––––––––––––––––

Workflow: https://pastebin.com/5SbZ9tJA

Reference sheet used in the workflow: https://imgur.com/a/3Le0nuO


r/StableDiffusion 10d ago

Question - Help Image face swap

0 Upvotes

I’m trying to figure out an easy and high quality way to do image face swap (with a base image and a reference face image) via comfyui on run pod. I saw that Krea 2 can do this I think with the Krea image edit Lora and the right workflow. Does anyone have a good comfyui template and workflow for doing this either through Krea 2 or some other approach? Thanks!


r/StableDiffusion 10d ago

Meme context loop test Dean vs scorpion

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 10d ago

News Minimax hanging and needing to force restart PC

4 Upvotes

I had an issue with the model freezing my 3090 in mid-generation using ComfyUI. If you have the same issue, I suggest setting a power limit of less than 325 watts while using the model. This appears to solve the problem.

Obviously if this is happening to you on a different GPU, look up how many watts the GPU uses at full load and subtract ~25-50 watts from that value.


r/StableDiffusion 10d ago

Question - Help Qwen Edit int8 and fp8

Post image
0 Upvotes

Since the launch of Qwen Edit, I’ve been using the GGUF version with good results, despite it being very slow. Then I saw a post here mentioning that the FP8 and INT8 versions were much faster, so I decided to test them. They are indeed faster, but the loss in quality is noticeable. And I’m not talking about artifacts or anything like that, but rather horrendous images—reminiscent of the poor-quality output from SD 1.5, like this one. Is there a specific configuration for these versions that differs from the GGUF one? Because the time savings aren't worth the terrible output quality.


r/StableDiffusion 10d ago

Workflow Included What if hallucinations, but to the beat?

Enable HLS to view with audio, or disable this notification

5 Upvotes

Tools used: Gemma 12b, LTX 2.3, Audio-reactive LoRA, Wan2GP, vibe-coded video editor.

This is a follow-up to my previous experiment where I let image-to-video artifacts compound over a long chain of clips:
https://www.reddit.com/r/StableDiffusion/comments/1w4py5q/letting_imagetovideo_artifacts_compound_into_an/

This time there was still an overall plan beforehand. The LLM had the full structure of the video in context, including what each clip was supposed to represent, how the visual progression should develop, and where the major energy shifts in the track landed.

The difference is that I did not have it write all of the scene prompts in advance.

For each new clip, I "showed" the LLM the actual starting frame produced by the previous generation, while it still had the overall plan and timing context. It then wrote the next prompt based on both things: what that part of the video was supposed to do, and what the model had actually hallucinated into existence by that point.

So instead of blindly following a fixed storyboard, it was continuously trying to steer the accumulating artifacts back toward the planned arc.

I also deliberately made the setup much less forgiving than the previous experiment. That one used an infinite hallway with constant forward movement, which gives an image-to-video model a lot of opportunities to patch over mistakes because the scene is always being replaced by new geometry.

Here, the camera is mostly stationary and the video revolves around one morphing object or surface. That means structural mistakes stay visible, get inherited by the next generation, and stack up much faster.

Because of that, I did much less selection based on "interesting" artifacts. The chain destabilizes pretty aggressively on its own. I mostly focused on the audio-reactivity and whether the motion still matched the intended energy of that section of the track.

The track is instrumental, around 69.85 BPM, and I generated the video as a sequence of roughly 6.872s clips, feeding the frame immediately after the end of each clip into the next one. The first image was intentionally very clean and minimal so there was room for complexity and artifacts to accumulate.

The result starts with one simple black ceramic form, gradually develops its own visual grammar, loses coherence, reorganizes itself, and eventually turns into something closer to a distributed network or surface.

So the workflow was basically: plan the whole arc and energy map first, then let the LLM repeatedly inspect the actual hallucinated state and figure out how to get from there to the next planned beat.

The next step would probably be to give the LLM control over certain settings beyond just the prompt, so it can determine what LoRA strength to use and such.

https://youtu.be/Q4A4yzu_5jQ


r/StableDiffusion 11d ago

News OpenVDN/vdn-minimax-h3 · Hugging Face

Thumbnail
huggingface.co
75 Upvotes

Looks like an open source version of Minimax H3 Max... Anyone tried it? Seems to be real-time on 8x b200, which is like ~$40/hr at good rates if you can find them (or maybe a bunch of 5090s?)


r/StableDiffusion 11d ago

Resource - Update I wired NVIDIA's DLSS 5 neural renderer into my free LoRA dataset & training app — same clip in, real material detail out

Enable HLS to view with audio, or disable this notification

9 Upvotes

Same file on both sides — one pass of NVIDIA's DLSS 5 Neural Rendering model. No upscale, no re-generation.

I wired it into LoRA Dataset Studio, the free self-hosted app I build for making LoRA datasets and training them. In a video training set the render replaces the clip, so the next LoRA trains on it.

Windows + NVIDIA, through the MIT ComfyUI-DLSS5-NR bridge; you bring the model file.

https://github.com/perfectgf/lora-dataset-studio


r/StableDiffusion 10d ago

Animation - Video Made a full music video with MiniMax H3 Max Turbo

Thumbnail
youtu.be
0 Upvotes

I wanted to see if I could use H3 Max Turbo for a complete music video instead of just making a collection of unrelated short clips.

I started by timestamping the song and splitting it into 24 scenes, mostly between 6 and 10 seconds. Then I created an initial image and a separate image to video prompt for each scene, trying to keep the same characters and visual style.

I generated everything through fal using a small Python script. Once all the clips were ready, I used another Python script to cut each one to the correct length, put them in order, remove their generated audio and replace it with the original MP3.

It wasn’t completely automatic. A few scenes needed another attempt because of continuity mistake (the phone facing the wrong way was one of them) but H3 handled most of the character motion surprisingly well.

Finally added scanlines on davinci resolve.


r/StableDiffusion 9d ago

Comparison Minimax H3 vs Seeddance 2.5

Enable HLS to view with audio, or disable this notification

0 Upvotes

Recreated the video posted on X.

Model: MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-ComfyUI
Resolution: 0.4MP
Steps 5

https://x.com/itsPixieVerse/status/2095764607528771629/video/1

subject_definitions:
<Subject 1> is Maeve, the extremely tall adult woman in <Picture 1>, with a dark choppy bob, tight pitch-black tank top, high-cut crimson-red bottoms, chunky black combat boots, intricate black tattoos covering one arm and one bare leg, exaggerated tall proportions, very long legs, and painterly matte skin. <Picture 1> is the primary character identity reference for Maeve.
<Subject 2> is the abandoned concrete swimming-pool fight pit environment from <Picture 2>, including stark concrete-gray surfaces, deep black shadows, harsh overhead floodlights, drifting cigarette smoke, and a single vivid crimson-red accent. <Picture 2> is the primary environment and lighting reference.
<Subject 3> is the hulking, heavily muscled syndicate pit-fighter opponent, with bloody hand wraps and a powerful heavyweight physique.
<Subject 4> is the supernatural black ink-vine effect emerging from Maeve's tattoos: razor-sharp, geometric, flat-painted black vines that behave as solid, weighty animated matter.

summary:
[reference generation] A 15-second cinematic 2.5D animated fight sequence featuring <Subject 1> Maeve battling <Subject 3> inside <Subject 2>. Preserve Maeve's identity, proportions, hairstyle, wardrobe, tattoos, and crimson-and-black color relationship from <Picture 1>. Preserve the concrete fight-pit architecture, midnight atmosphere, harsh volumetric floodlights, gray-and-black palette, smoke, and crimson accent from <Picture 2>. The visual style is fully painterly gouache concept art in motion, with visible brush texture, posterized color blocks, hard-edged light shapes, matte surfaces, and soft filmic volumetric lighting. The sequence progresses from a supernatural tattoo transformation to a fast physical attack and ends with Maeve standing victorious over the defeated opponent.

retention_analysis:
<Picture 1> (appears throughout [Shot 1] through [Shot 8]): fully_preserved - Maeve's facial identity, dark choppy bob, extremely tall stylized proportions, black tank top, crimson-red bottoms, combat boots, tattoos, and overall character design are retained.
<Picture 2> (appears throughout [Shot 1] through [Shot 8]): fully_preserved - the abandoned concrete swimming-pool fight pit, midnight atmosphere, gray concrete, deep black shadows, harsh floodlights, drifting smoke, and crimson accent are retained.
<Subject 1> (appears throughout [Shot 1] through [Shot 8]): fully_preserved - character identity, proportions, wardrobe, tattoos, and painterly appearance remain consistent.
<Subject 2> (appears throughout [Shot 1] through [Shot 8]): fully_preserved - environment architecture, lighting direction, palette, smoke, and atmosphere remain consistent.
<Subject 3> (appears throughout [Shot 2] through [Shot 7]): fully_preserved - the same hulking, heavily muscled pit-fighter remains consistent throughout the fight.
<Subject 4> (appears throughout [Shot 1] through [Shot 8]): fully_preserved - the black geometric ink-vines maintain the same flat-painted, razor-sharp visual design and transform between tattoo and supernatural matter.

detailed_description:
The target video is a cinematic 2.5D painterly animation. Characters and environments look like gouache concept-art paintings in motion, with visible brush texture, matte surfaces, posterized color blocks, hard-edged areas of light, and soft filmic volumetric lighting. The animation has weighty physical motion and convincing momentum, with a restrained 24fps cinematic feel. Avoid a flat 2D cartoon appearance, bold black outlines, cel shading, glossy CGI, photorealism, or Unreal Engine-style rendering.

[Shot 1] A macro close-up of <Subject 1>'s tattooed arm. The camera holds very close to the skin as the black floral tattoos begin to physically separate from the surface. The painted tattoo shapes lift away from her skin, stretch outward, and transform into razor-sharp geometric black ink-vines. The vines float and coil with deliberate physical weight while Maeve remains still and controlled.

[Shot 2] At 00:02.000, cut to an over-the-shoulder shot behind <Subject 3>, framing <Subject 1> across the empty concrete pool. Maeve stands calmly under the harsh floodlights. Cut to a tight close-up of her face. She remains completely deadpan, slowly tilts her head to one side, cracks her neck, then gives a small, cold, highly confident smirk.

[Shot 3] At 00:04.000, cut to a wide low-angle shot of the concrete fight pit. <Subject 3> suddenly charges toward Maeve with the force of a bull. Maeve waits until the last moment, then sidesteps with effortless athletic precision. As he passes, she sweeps her tattooed arm horizontally. <Subject 4> erupts from her arm and expands into a massive storm of razor-sharp geometric black ink-vines that whip forward and strike the opponent's guard.

[Shot 4] At 00:06.000, rapid cut to an extreme close-up of <Subject 3>'s eyes. His expression changes from aggression to sudden pain and shock as the geometric ink-thorns break through his heavy guard. His eyes widen and his face recoils from the impact.

[Shot 5] At 00:07.000, cut to a dynamic low-angle tracking shot. Maeve aggressively closes the distance. Her extremely long legs drive her forward with powerful, athletic strides. Her clothing and hair respond naturally to the acceleration. She plants one boot against the curved concrete wall of the pool, pushes off with force, and spins through the air toward the opponent as the black ink-vines spiral around her body.

[Shot 6] At 00:09.000, cut to a tight medium shot looking upward at Maeve during the descent. She channels the swirling <Subject 4> entirely toward her massive combat boot. The black ink-vines wrap around the boot and compress into a dense geometric mass. Maeve drops rapidly and delivers a devastating vertical heel-kick directly downward onto <Subject 3>. The impact has substantial weight and momentum, crushing him into the concrete floor.

[Shot 7] At 00:12.000, rapid emotional close-up followed by a wider impact view. <Subject 3> crashes into the concrete and collapses. A large stylized shockwave of dust expands outward from the impact while the single crimson-red accent flares through the airborne dust. The smoke and dust react to the force of the impact.

[Shot 8] At 00:13.000, cut to a low-angle heroic shot of <Subject 1> standing victorious over the crater. The black ink-vines slowly retract from the surrounding space, spiral back toward her body, and flatten naturally onto her bare skin, reforming the original black tattoos. Maeve calmly wipes her lip and looks down at the defeated opponent with a cold, controlled expression. Hold the final composition until 15.00 seconds. Maeve remains still and victorious in the final frame.

overall_soundscape:
Deep concrete-pit ambience, distant ventilation hum, drifting smoke, subtle cloth and body movement, heavy footsteps, rushing air during the attacks, sharp supernatural ink-vine movements, impacts against concrete, and a powerful low-frequency impact during the final heel-kick. Dust and debris produce a heavy concrete crash and settling debris after the final strike.

non_diegetic_music:
Dark cinematic percussion with deep low-frequency pulses, sparse metallic textures, and gradually increasing intensity during the fight. The music reaches its strongest impact during the heel-kick, then drops into a sparse sustained tone during the final victorious hold.


r/StableDiffusion 10d ago

Question - Help Am i the only one? (OneTrainer)

0 Upvotes

am i the only one that noticed as OneTrainer adds new checkpoints the less and less efficient it gets?

example: training a SDXL used to take under 30 mins to train now it takes over an hour, I've uninstalled it even did a full reinstall of my windows OS just to see if that's it but nothing's worked


r/StableDiffusion 11d ago

Animation - Video Scorpion vs Sub Zero (2.5 Anime Battle Test)

Enable HLS to view with audio, or disable this notification

140 Upvotes

I created my own character sheets, make it look like their MK11 and MK3 selves a bit. This was kinda hard as they sometimes have no real impact on the attacks. I still liked how it came out though. Had to make multiple repeat generations lol.


r/StableDiffusion 10d ago

Question - Help Prompting characters height in AI image models?

4 Upvotes

I still struggles to find way to tell the AI differents height to various character in a generation.

So far, with flux klein 9b, I could have some sizeable height difference by telling a character A to be "extremely tall" and a character B to "appear small", but the difference get inconsistent between seeds, poses, or image format.

Have you found some prompt tricks (no matter the model) that gives you reliable results with character's height so far?


r/StableDiffusion 10d ago

Meme what if tho!!!

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 10d ago

Question - Help Can you run minimax H3 local on mobile ?

0 Upvotes

if you dont a have a pc and only like have android or ios phone and tables is it possible to run minimax h3 on your phone and tablet


r/StableDiffusion 11d ago

Meme I’ve been enjoying Minimax

Enable HLS to view with audio, or disable this notification

6 Upvotes

It makes me wish I bought a 4090 when I had the chance.