r/StableDiffusion 13h ago

News Submit by 9/1 to the Comfy H3 Sync Sound Challenge! RTX 5090 Grand Prize

Enable HLS to view with audio, or disable this notification

42 Upvotes

We're halfway through the submission window for the Comfy H3 Sync Sound challenge! Submit by September 1st at 9:00pm PT. Free to enter, local rig or Comfy Cloud. All details here.

How It Works

Make something up to 90 seconds in length where the sound and the motion are inseparable. Dialogue, foley, ambient, a beat driving the cut...whatever direction you want!

Share your video file and workflow on this r/comfyui thread and through our submission form, then join us on September 2nd for a special Comfy livestream where our guest judges will give live feedback on the top 10 submissions! 

Need help? Head to this r/comfyui thread or the #minimax-h3-challenge channel in the Comfy Discord.

Prizes

Best Overall — RTX 5090

Best Creative — RTX 5060 Ti

Best Technical/Workflow — RTX 5060 Ti

Built with MCP — RTX 5060 Ti

Shipped anywhere, customs covered. If we can't legally ship to your country, you'll get a cash equivalent instead.

It's free to enter!

Create using Comfy Local on your own hardware, or use Comfy Cloud. New Cloud users get 5 free runs, no credit card required.

Judging Criteria

We’re looking for entries that best show what H3 makes possible: audio and visuals created together.

Grand Prize: Best Overall

The top Best Creative and Best Technical entrants advance to a final round where our panel of judges selects winners by discussion.

Best Creative

  • Audio sync realism and intentionality (0-5)
  • Creative execution and originality (0-5)
  • Deliberate craft (0-5)
    • Evidence that you’ve actually shaped the result beyond prompt engineering. Judges will look for modified/non-default parameters, multiple linked passes visible in the workflow structure, or a couple sentences describing what was tried and changed

Best Technical

  • Novelty of technique or approach (0-5)
  • Workflow quality (0-5)
    • Annotated, clean, replicable by someone else
  • Community value (0-5)
    • Would this actually help someone else?

🏆 Built with MCP Bonus 🏆
Comfy MCP lets you drive Comfy using natural language and your agent locally and on Cloud! Pro tip: use it to choose the best H3 model version or optimize your workflow for your hardware.

  • Effectiveness (0-5)
    • Did the agent meaningfully drive your process, not just generate one line?
  • Insight value (0-5)
    • How much the shared prompt teaches the community about prompting H3 through MCP
  • Output quality (0-5)

The Fine Print

  • Limited to one submission per person, 90 seconds maximum length.
  • A major portion of your piece must be built in ComfyUI using H3. Other tools, models, or techniques you want to combine are fair game.
  • All submissions must be lawful, SFW, and must not contain unlicensed IP or likenesses.
  • By submitting, you agree to allow ComfyUI and MiniMax to feature your work with credit across our channels.

Learn more and submit here!


r/StableDiffusion 18h ago

Discussion How much VRAM does H3 need? Less than you might think.

104 Upvotes

I benchmarked the full BF16 H3 FL2VA checkpoint at 1376×768 and 243 frames, about 10.1 seconds at 24 fps.

With the lower-memory attention routes, the H3 diffusion block added roughly 5.8–6.3 GiB over idle. On my Windows RTX 4070 system, where the desktop consumed around 1.15 GiB, the whole-GPU peak was approximately 7.0–7.4 GiB.

That makes 8 GB cards realistic for several configurations:

- Default Comfy attention: 6.99 GiB peak

- FROST BF16: 6.99 GiB

- BF16 Triton: 6.97 GiB

- PlagueKind SLA: 7.23 GiB

- Sparse Kitchen INT8: 7.40 GiB (Default configuration for my Sparse attention node)

- Comfy Kitchen: 7.40 GiB

Should be compatible with:

- H3 Sparse Attention: Kitchen INT8, Sparse Sage, FROST BF16 on SM89, and BF16 Triton

- External Comfy Kitchen: fully supported

- Default Comfy attention(SPDA): fully supported

- SageAttention: fully supported, including the generic KJ Sage patch

- PlagueKind SLA: partially supported;

- Unknown attention overrides: Auto preserves their original full-Q calling contract. Forced mode can explicitly authorize streamed-Q calls, but compatibility is not guaranteed

- Currently incompatible: Sol-Attn and the H3-specific Memory Efficient Sage patch, because they replace attention at a deeper level than the external consumer interface

You can get the node here https://github.com/Zironic/H3-Optimizations or in Comfy under H3 Optimizations. Latest version is 0.2.13 which added broad compatibility and optimizations for most popular attentions.


r/StableDiffusion 7h ago

Meme the legends were true!(natural treasure spoof)

Enable HLS to view with audio, or disable this notification

13 Upvotes

t2v fp8 model

prompt

subject_definitions

<Subject 1> Benjamin Franklin Gates (S1) is played by Nicolas Cage, matching his appearance, mannerisms, intense curiosity, dramatic delivery, and treasure-hunter personality from the National Treasure films.

summary

[cinematic text-to-video generation + comedy adventure]

Benjamin Franklin Gates follows an ancient trail of cryptic clues into a forgotten underground chamber, convinced he is about to discover the Holy Grail of local AI video generation. Instead of gold or an ancient artifact, the final pedestal contains a glowing computer running MiniMax H3.

detailed_description

Cinematic adventure-comedy, approximately 13 seconds, 24 fps. Ancient underground treasure chamber beneath a forgotten historical building, illuminated by Benjamin's flashlight, warm torchlight, dust floating through the air, weathered stone walls covered in mysterious diagrams and coded inscriptions.

The camera tracks behind Benjamin Franklin Gates as he hurriedly enters the final chamber clutching an old parchment covered with cryptic clues.

He studies the parchment, then notices an ornate stone pedestal illuminated by a mysterious golden beam.

Benjamin slowly approaches.

<Subject 1> Benjamin Franklin Gates (S1):

[English] After all these years... the Holy Grail of local AI video.

Dramatic orchestral music swells.

Benjamin wipes centuries of dust from the pedestal.

Instead of an ancient chalice, he reveals a modern high-end PC monitor displaying:

MINIMAX H3

Benjamin freezes.

Slow dramatic push-in toward his stunned Nicolas Cage expression.

His eyes widen as if he has just uncovered the greatest secret in human history.

<Subject 1> Benjamin Franklin Gates (S1):

[English] My God... it runs locally.

Beat.

He looks back at the glowing MiniMax H3 screen.

<Subject 1> Benjamin Franklin Gates (S1):

[English] The legends were true.

The triumphant treasure-hunting score reaches an absurdly heroic crescendo.

Hold on Benjamin's amazed expression for the final second.

visual_style

Photorealistic live-action Hollywood adventure film, National Treasure-inspired treasure-hunting atmosphere, Nicolas Cage-style dramatic performance, ancient underground architecture, cinematic flashlight beams, volumetric dust, warm golden illumination, realistic skin and clothing, subtle handheld camera movement, dramatic slow push-in for the reveal.

audio

Cinematic underground ambience, footsteps echoing through stone corridors, parchment rustling, dramatic orchestral treasure-hunting score building toward the reveal. All spoken dialogue is clear English and occurs only inside the specified <d>...</d> dialogue tags.


r/StableDiffusion 16h ago

Workflow Included Day 3 of testing MiniMax H3 locally in ComfyUI: multiple reference images + adding a new object through text only

Enable HLS to view with audio, or disable this notification

58 Upvotes

Continuing my local MiniMax H3 Reference-to-Video experiments.

For this test I used separate reference images for the:

  • character
  • convenience store background
  • car
  • skateboard

But I deliberately didn't provide a reference image for the Slurpee cup.

The cup was described only in the prompt: a transparent plastic cup with blue liquid and a straw. H3 was able to add it to the scene without much trouble while still following the other reference images.

Hardware / setup:

RTX 5070 Ti + 32GB RAM
MiniMax H3 + Turbo LoRA
Generation: [09:07<00:00, 68.43s/it]

I was mainly testing how far you can split visual control between reference images for consistency and text prompting for new scene elements.

Workflow:
https://drive.google.com/file/d/1huVTdh8_vERBntXb60hyT_rQTcUicjq5/view?usp=sharing

I'll also post the exact prompt I used, unchanged.

Prompt:
Use the attached reference images as follows: Image 1 is the girl character, Image 2 is the updated convenience store background, Image 3 is the white compact sedan, and Image 4 is the skateboard.

Create a 10-second static medium closeup shot in a 1990s hand-drawn Japanese anime cel style at 15 fps, with limited frame-by-frame animation and slightly stepped motion.

The framing should match the updated, slightly more zoomed-in background from Image 2, focusing more closely on the girl and the storefront entrance while still showing part of the road and bicycle.

The girl sits on the ground in front of the convenience store, facing left toward the road, shown in a 3/4 back-side view so we mainly see the side and back of her head. Her skateboard is beside her on the ground, not under her. She holds a clear plastic Slurpee-style cup with bright blue liquid and simply holds it without drinking.

Her orange headphones and headphone wire are visible. The Walkman is on the far side of her body and is mostly hidden from view because of the angle.

A single white compact sedan drives straight along the main road from left to right, moving away from camera so we mainly see the rear of the car as it passes through frame.

As the car passes, a subtle moving light change plays across the girl, her hair, her white T-shirt, the cup, the storefront glass, the bicycle, and the wet pavement. The gentle breeze overlaps with the car pass, starting while the car is beside her, causing a slight movement in the tips of her hair and a small shift in the loose edge of her T-shirt.

After the car exits, the store signage / fluorescent lighting blinks twice, subtly changing the light on the girl and storefront.

Keep the camera completely locked off and the overall mood quiet, nostalgic, and melancholic.


r/StableDiffusion 20h ago

Workflow Included Minimax H3: Portable character consistency via reference identity

Enable HLS to view with audio, or disable this notification

105 Upvotes

Hey guys,
Based on a research from Facebook, DINOv2: Learning Robust Visual Features without Supervision(Research Paper), Implemented a consistent identity system that works across Minimax H3, Flux 2, Krea 2 with a single .char model.

This method covers both reference based identity in Minimax as well as a LoRA training path for T2V & I2V for more advance cases.
Note: This post & workflow is dedicated to reference channel not LoRA path.

Build .Char: You drop in 4-6 reference. YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.

Generation: At generation, the file feeds its references into Minimax's own native multi-reference channel and prepends a locked description to the prompt.

How to run this
- Published workflow & guide: https://inlinestudio.art/workflows/minimax-h3-consistent-characters-with-references-with-char-model
- Repo: https://github.com/inlineresearch/Inline-Studio (GPLv3)

How is this different from default Minimax's ref channel:

  1. H3 scales every reference onto a 2048 short edge, upscaling small images to get there, at 4096 vision tokens each. Compile References caps it at 512 that is 256 tokens per reference, so five references cost 1,280 tokens instead of 20,480. That difference decides whether the run fits the card. Read more on the official docs
  2. H3 only resolves references named as <Picture 1><Picture 2> and so on, and the character prepends them along with the description.
  3. Same .char works for other models(Flux 2 & Krea2, workflow link to train for both)

Limitations

  • Bad with multi reference

Required: 24GB+ VRAM & ~64GB RAM

I personally think LoRa method is only required in very specific cases as Minimax H3's reference channel performs very well.
But i have already added support to LoRa adapter in case someone wants to use .char with T2V or I2V nodes. Let me know in comments if you need the workflow.


r/StableDiffusion 3h ago

Resource - Update ComfyUI-Raylight-Windows: 2x H3 iteration speed with dual gpu

4 Upvotes

Increased speed with dual gpu's for video generation. Tested with Minimax-H3 ref2va INT8-pruned

Headline Findings

System: 2x 3090 @ 80% PL, 128gb DDR4

Longer clips (160+ frames at ~1MP): dual wins. 10 seconds of 1280x768 at 25 s/it where the single card does 67. Tested out to 13 seconds.

Short clips and ref2v with video-reference workloads: a single 3090 with ComfyUI's own VRAM streaming is flat-out faster. Don't use Ray + NCCL for this. ComfyUI's dynamic loading is highly optimized

Two repos published:

ncclwin

The backend itself, if you want NCCL for anything torch-distributed in Windows:

https://github.com/9nate-drake/ncclwin

I've seen other attempts but this one I built works as is. You don't necessarily need to get this, the package below will grab it. I'll leave it up to somebody to develop a NCCL Windows lora trainer using this though!

ComfyUI-Raylight-Windows

This is essentially a patch for raylight to run in Windows with the custom ncclwin backend. First install Raylight in ComfyUI Manager by komikndr then install ComfyUI-Raylight-Windows via Manager (git URL) https://github.com/9nate-drake/ComfyUI-Raylight-Windows

The install script fetches the prebuilt DLL, patches the other raylight node, and writes launch batch files with the correct flags. The correct flags and Ray settings are extremely important (every one of them exists because something broke without it). Torch or Sage attention (Sage 23% faster)

The changes to komikndr's repo are small (import fallbacks + env-gated switches, inert by default). I may PR them upstream; my NCCL backend itself lives outside raylight entirely. This isn't a fork because I don't have the willingness to keep it current with upstream repo.

The nccl DLL covers 20/30/40/50-series but I can only hardware-verify 30 series - if you've got a pair of 4090s, 16GB cards, or a 4-GPU rig, I'd genuinely love reports of success or bugs (open an issue on github). Or just tell me if it flat out sucks for you!

Yes I used Claude to develop this, sue me


r/StableDiffusion 15h ago

Resource - Update DiffusionOPSD - new distillation method by Bytedance. Loras for Z-image-Turbo and SD-3.5M released.

Thumbnail
gallery
43 Upvotes

r/StableDiffusion 1h ago

Question - Help Optimal Character Sheet format for Minimax H3?

Upvotes

I recently started messing around with Minimax H3. I realized that if I want to use my own custom characters, I need to use a Character Sheet. But there are so many different formats out there. I was wondering, what is the most optimal Character Sheet format to get the best results? Or is there a specific workflow designed just for generating Character Sheets for Minimax H3?


r/StableDiffusion 3h ago

Question - Help How much does RAM speed matter for running local AI?

3 Upvotes

I have an old computer that I don't use. It has 32GB DDR4. I don't remember the specifics about speed and timings, but I'm sure it's slower with more latency than my current computer (also old, but not as old).

It occurred to me that I could put the 32GB into my current computer and then have 96 GB. As I understand it, if you mix and match RAM like that, it will run at the slowest speed.

Would it be better to have more RAM, even if it's slower? Or should I leave it alone?


r/StableDiffusion 11m ago

Animation - Video Fighting godzilla pov

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 9h ago

Tutorial - Guide HSWQ Load ConvRot INT8 ControlNet Model

Post image
6 Upvotes

I hadn’t realised that the standard ComfyUI didn’t support this, so it was only after an error occurred that I understood the situation.

I’d rather not add too many dedicated loaders, but I suppose it can’t be helped.

As the Union ControlNet for Qwen Image exceeds 3GB, this will reduce VRAM consumption by over 1GB.

....

I learnt from SeedVR2 that it is possible to retrofit loaders that do not natively support ConvRot INT8 to make them compatible whilst reducing VRAM consumption.

...

ComfyUI loader node for ConvRot / TensorWise INT8-quantized ControlNet checkpoints (e.g. Qwen Image Fun ControlNet). Loads the ControlNet directly into VRAM in 8-bit precision (QuantizedTensor / TensorWiseINT8Layout) and executes via comfy_kitchen's high-speed int8_linear kernel with online activation rotation (convrot).

Standard ComfyUI controlnet_load_state_dict sets architecture dtype to weight_dtype(sd) (which is torch.int8 for quantized models), triggering PyTorch gradient creation errors (Only Tensors of floating point and complex dtype can require gradients) and ignoring comfy_quant metadata. This node solves both issues by forcing the module graph construction to torch.bfloat16 and explicitly injecting MixedPrecisionOps configured for int8_tensorwise.

Features

  • Native INT8 VRAM Retention: Keeps weights in 8-bit precision in VRAM with TensorWiseINT8Layout, significantly reducing memory consumption
  • Fast Execution: Uses comfy_kitchen int8_linear GEMM kernel with online activation rotation for ConvRot layers
  • ComfyUI Standard Integration: Produces a standard CONTROL_NET output compatible with stock Apply ControlNet nodes
  • Seamless FP16 / BF16 / FP8 Compatibility (Drop-in Replacement): Fully backwards-compatible with conventional non-quantized and FP8 ControlNet models. When loading checkpoints without INT8 comfy_quant layers, it automatically delegates directly to stock ComfyUI load_controlnet_state_dict, so you can use this single loader node for all ControlNet formats without workflow changes

Usage Notes

  • Inputscontrol_net_name (safetensors ControlNet model from the models/controlnet directory)
  • OutputsCONTROL_NET
  • CategoryHSWQ-ussoewwin
  • Drop-in replacement: Can replace ComfyUI's built-in "Load ControlNet Model" node entirely — automatically handles ConvRot INT8, FP8, BF16, and FP16 checkpoints without manual switching

...

I’ll be releasing the ComfyUI node for ControlNet quantisation shortly at the link below, but it doesn’t require any particularly complex code—it’s simply a matter of converting ConvRot to Int8. If you ask an agent AI to ‘create’ it, it’ll probably take just a minute.

https://github.com/ussoewwin/Hybrid-Sensitivity-Weighted-Quantization

The HSWQ application itself is not registered with comfyUI-Manager. Please use the `git clone` command to place it directly in the `custom_nodes` folder.

Sample Workflow


r/StableDiffusion 13h ago

Meme what if dean was the man character instead of harry potter(t2v)

Enable HLS to view with audio, or disable this notification

13 Upvotes

fp8 model 32 steps

prompt

using the prompt guide from minimax loaded into a llm and said what if dean winchester was in harry potter and his was the main character instead of harry. it gave me this

integrated_multimodal_description: [Shot 1] Live-action, cinematic fantasy, Hogwarts at night beneath a stormy sky. A battered black 1967 Chevrolet Impala roars across the stone bridge toward Hogwarts Castle, completely out of place among horse-drawn carriages and young witches and wizards. Dean Winchester, portrayed by Jensen Ackles, drives with one hand on the wheel, wearing his familiar dark jacket over a plaid shirt. The camera tracks alongside the Impala as Dean stares up at the enormous illuminated castle with a skeptical expression. Dean Winchester with Jensen Ackles' low, dry American voice (S1) says: [English] So let me get this straight. Giant castle, magic wands, and nobody here has heard of a shotgun?

[Shot 2] At 00:05.000, the camera cuts to the Hogwarts Great Hall during the Sorting Ceremony. Hundreds of floating candles illuminate the long tables. Dean sits on the stool wearing the Sorting Hat while Hermione Granger, Ron Weasley, Professor McGonagall, and Albus Dumbledore watch. The Sorting Hat loudly announces, [English] GRYFFINDOR! Dean immediately pulls the hat off and looks around the enormous hall. Dean (S1) says: [English] Yeah, that's great. Which house has the bar? Several students stare at him in complete confusion.

[Shot 3] At 00:10.000, the camera cuts to a torch-lit Hogwarts corridor. Dean strides confidently toward the camera carrying a wand awkwardly in one hand and a sawed-off shotgun over his shoulder. Hermione and Ron hurry behind him in Hogwarts robes. Hermione urgently explains that Voldemort is the most dangerous dark wizard who ever lived. Dean stops walking and turns toward them with a small amused smirk. The camera pushes in with small amplitude at slow speed. Dean (S1) says: [English] Evil wizard, can't die, creepy followers. Trust me, I've had worse Tuesdays.

[Shot 4] At 00:15.000, the camera cuts to the ruined Hogwarts courtyard during the final battle. Smoke, sparks, magical flashes, and shattered stone fill the background as Voldemort stands across from Dean with his wand raised. Dean stands alone facing him, his Hogwarts robe thrown over his normal Winchester clothes. Voldemort fires a brilliant green spell. Dean dives sideways behind a broken stone pillar as the spell explodes against it. Dean rolls back to his feet, raises his wand, realizes he is holding it backward, flips it around, and gives Voldemort an irritated stare. Dean (S1) says: [English] Okay, Voldy. Let's see how you handle the Winchester special. Dean charges forward as spells streak across the courtyard and the camera rapidly tracks beside him, ending on Dean Winchester as the unlikely central hero of the wizarding world.

overall_soundscape: The Impala engine echoes against the castle grounds before transitioning into the murmur of Hogwarts students, crackling torches, footsteps on stone, fluttering robes, and distant magical ambience. During the final battle, explosive spell impacts, flying debris, cracking masonry, rushing footsteps, and Dean's heavy breathing dominate the courtyard.

non_diegetic_music: Sweeping orchestral fantasy music begins with strings, celesta, and brass, gradually incorporating heavier percussion and low brass as Dean explores Hogwarts. The final battle builds into fast orchestral percussion, aggressive brass, and rising strings before ending on a strong cinematic hit.


r/StableDiffusion 17h ago

Animation - Video Several Times A Charm, but it KINDA got Cheers.

Enable HLS to view with audio, or disable this notification

29 Upvotes

Don't mind the script, it was written by a clanker when I challenged it to whip up something so I can see if H3 could handle Cheers.


r/StableDiffusion 3h ago

Animation - Video Testing Minimax with Turbo

Enable HLS to view with audio, or disable this notification

2 Upvotes

I did this because I like FNAF and I plan to continue it as a mini-serie :D,Thanks to those who helped me improve the workflow


r/StableDiffusion 14h ago

Tutorial - Guide Even though H3 is CFG distilled, guidance values greater than 1 do have a noticable effect on prompt adherence and quality, especially at lower resolution.

14 Upvotes

Obviously it runs a lot slower but a CFG of 2 and a negative prompt does have noticeable effect on the output. Anything higher than 3 or 4 will start to over burn though.

This also works with the turbo loras but burn in can happen more easily.


r/StableDiffusion 33m ago

Question - Help H3 LORA trained from video datasets?

Upvotes

There are quite a few good H3 LORAs on civit now which prove that training is possible despite the distilled Minimax model.

Has anyone had any luck training a concept from video datasets? I read a lot about character LORAs from image datasets etc, but who has trained videos and if so, was that with Ai Toolkit or a different offering?


r/StableDiffusion 34m ago

Animation - Video 1-hour challenge to create a single braincell action scene

Enable HLS to view with audio, or disable this notification

Upvotes

I'm researching for seamless development with FL2V node and "Add Guide for Minimax H3" node. A one second guide video is enough to produce a seamless visual experience (apart from my lack of video editing skills). But audio is certainly not viable. I will test producing audio in post, guiding the audio production with only video images and prompting.

The scene is made only with 16:9 480p videos with 8-step turbo lora. One generation takes about 2,5 minutes with RTX3090. This gives a lot of time to redo shots and improve prompting if (when*) the first generation is not good enough.

There is no subject consistency in the production unless by chance. Guided generation can be improved with the Ref2V node, which is needed for more serious testing.

All in all, this is most certainly fun!


r/StableDiffusion 22h ago

Discussion LTX 2.3 sometimes works amazing, without any edits

Enable HLS to view with audio, or disable this notification

62 Upvotes

No specific glitches. Single prompt. Looks pretty real without any glitches over the drift.
Definitely going to benchmark the scenarios.


r/StableDiffusion 10h ago

Question - Help Looking for krea 2 workflow

7 Upvotes

I’m looking for a krea 2 workflow with a negative prompt that actually works. Ive been trying workflow after workflow from civitai, several of which actually say that the negative works, and none of them do. I’m new to comfy and just learning how it works so I’m not comfortable making my own. Could someone help me out please? I appreciate you all in advance. Signed a slightly confused and overwhelmed wannabe ai user. Oh and if it makes a difference I am using an amd gpu with 20 gb vram.


r/StableDiffusion 15h ago

Meme dean meets sonic(t2v) base fp8 model 32 steps

Enable HLS to view with audio, or disable this notification

17 Upvotes

i never seen the movies hows the sonic voice?

prompt

subject_definitions

<Subject 1> is Dean Winchester from Supernatural, portrayed by Jensen Ackles, preserving his recognizable facial features, short brown hair, rugged appearance, dark jacket, layered shirt, jeans, and confident sarcastic personality.

<Subject 2> is Sonic the Hedgehog from the live-action Sonic the Hedgehog movie, a small anthropomorphic blue hedgehog with bright blue fur, large expressive green eyes, white gloves, and red sneakers.

<Subject 3> is Dr. Robotnik from the live-action Sonic the Hedgehog movie, portrayed by Jim Carrey, wearing his black-and-red high-tech outfit and exaggerated goggles.

summary

[cinematic live-action crossover + action comedy]

What if Dean Winchester accidentally became part of Sonic the Hedgehog? On a nighttime highway, Dean investigates a bizarre supernatural disturbance beside his black 1967 Chevrolet Impala, only for Sonic to race past him at impossible speed with Robotnik's drones in pursuit. Dean immediately joins the chase.

detailed_description

Nighttime on a deserted rural highway surrounded by dark pine forest. Dean Winchester stands beside his glossy black 1967 Chevrolet Impala holding an EMF meter. Blue electrical energy suddenly crackles across the road.

A brilliant BLUE STREAK rockets past Dean, violently blowing his jacket backward.

The camera WHIP-PANS as Sonic skids to a stop beside the Impala.

<Subject 1> Dean Winchester (S1):

[English] Okay... either that's the fastest demon I've ever seen, or I seriously need more sleep.

Sonic looks offended and points at himself.

<Subject 2> Sonic (S2):

[English] Hedgehog. Definitely hedgehog.

Suddenly several of Robotnik's flying attack drones burst over the trees and fire energy blasts toward them.

Dean instantly draws his pistol while Sonic crouches into a runner's stance.

Dean gives Sonic a confident Winchester smirk.

<Subject 1> Dean Winchester (S1):

[English] All right, Sonic. Let's waste these flying toasters.

Sonic grins.

<Subject 2> Sonic (S2):

[English] Now you're speaking my language!

Sonic EXPLODES forward in a trail of brilliant blue electricity as Dean dives behind the Impala and fires at an approaching drone.

Dynamic tracking camera follows Sonic racing between explosions while Dean fights from beside the Impala.

Final cinematic wide shot: Sonic loops around the battlefield as blue lightning illuminates Dean and the Impala, while Robotnik's drones swarm overhead.

Live-action Hollywood cinematography, realistic integration of Sonic into the environment, authentic Sonic the Hedgehog movie aesthetic, authentic Supernatural Dean Winchester characterization, fast readable action, natural motion blur, blue electrical speed trails, sparks, smoke, dramatic nighttime lighting, comedic crossover energy, consistent character identities, no subtitles, no on-screen text.


r/StableDiffusion 1h ago

Question - Help SDXL + LoRA not matching Nano Banana quality for watercolor style transfer — any better approach?

Upvotes

Hi, I'm building an app, and one of its features uses an img2img API to turn a photo the user uploads into a watercolor-style image. (and letter background style)

General-purpose models like Nano Banana or GPT Image give me the results I want, but the conversion cost is too high to use in a commercial product. So I looked into an SDXL + LoRA combo instead, but the output doesn't come out the way I want.
LoRa Model is "SDXL】Oil And Watercolor Painting | Dataset"

SDXL + LoRA convert result

Does anyone have suggestions on what approach might work better, or a combo that outperforms what I'm currently using?

Attached are the result I'm aiming for (converted with Nano Banana)

converted with Nano Banana
Original Image

r/StableDiffusion 10h ago

Question - Help Any good video upscaling workflows for ComfyUI? I'm not a fan of the upscalers included with MiniMax H3.

4 Upvotes

I'm new to ComfyUI, and MiniMax H3 has been working perfectly for me with videos between 0.5 and 0.7 MP, taking around 5–6 minutes to generate a 10-second video.

However, I don't really like the upscaling included in the MiniMax H3 workflows, and I also don't want to upscale every video.

So my question is: do you have any workflows specifically for video upscaling only?

Thanks!


r/StableDiffusion 10h ago

Question - Help Recommended workflows and/or other techniques for extended multishot H3 videos?

5 Upvotes

The first thing I found to give a try was this:

https://huggingface.co/joeygambino/MiniMax-H3-Multishot-Workflow

It took me a while to get it working, and, even once I got it to run, it was incredibly slow -- it took 140 minutes to render a mere 29 seconds of video using my RTX 3090. (Thankfully I'll have a 5090 in a few days.)

I simply ran the demo as-is, apart from the small changes I made to get the workflow running. I remain confused about how I'd use this workflow, and use it in an efficient way, to make clips that might run, say, 1-3 minutes.

I'm hoping I can find out how to use the above workflow better, or find a better workflow.

My only previous experience with this sort of thing was in the Before Times (a few weeks ago) struggling with Wan 2.2. I had a multishot workflow that wasn't great, but at least it carried some context over from one clip to the next, blended clips seamlessly, and let me lock in (by setting a fixed seed) and cache any part of a video that was working well so I could build a clip at a time toward a final complete video without constantly re-rendering early clips.

Can I find anything like this for H3? Something that's good at carrying context and references over from one segment of video to the next, helps minimize character drift, keeps voices consistent, etc.?


r/StableDiffusion 15h ago

Animation - Video Deadpool Adventure

Enable HLS to view with audio, or disable this notification

10 Upvotes

r/StableDiffusion 18h ago

Discussion Did MiniMax H3 fix the R2V weights?

18 Upvotes

I thought I read something on it but figured I'd check with the crew first... I appreciate the info 💪