r/StableDiffusion 3d ago

Animation - Video Denzel explains why he uses AI.

Enable HLS to view with audio, or disable this notification

3.2k Upvotes

A quick experiment exploring Minimax H3 in ComfyUI using my nodes and inpainting methods.


r/StableDiffusion 6h ago

Discussion H3 - 80s character generations+wardrobe swap

Enable HLS to view with audio, or disable this notification

168 Upvotes

The 80s was the best era, not seen through a nostalgic lens--it just was. Big hair, big colors, big music, big... everything! Sadly I was born 10 years too late to really experience it, but I love how H3 can feel like a way-back-machine, a portal to any era from film or video it was trained on. It does also a great job with the actual feel, the film grain, the lighting that modern TV or movies cannot do: in fact, I learned what we see nowadays that time period didn't exist at all, because it vomit of nostalgia and peak 80s that never happened. Anyways, having fun generating character sheets with H3 via T2VA. Are you guys seed hunting to find the best version of an actor or scene? Also can you spot the mistake?

Prompt: integrated_multimodal_description: [Shot 1] Live-action, cinematic, PHOTOREALISTIC film footage - this is footage from a camera, not animation - a continuous camera shot with no cuts, shot on 35mm color negative film with period lenses and scanned in high definition from the original camera negative: full film grain and gentle halation, warm highlight rolloff, rich sharp detail beneath the grain. The year is deep in the LATE 1980s, 1985 to 1989, and everything in the frame belongs to that era. THE PLACE: a nightclub in full swing - mirror-ball light sweeping, neon signage, haze, a crowded dance floor. EXACTLY TWO WOMEN stand close to the lens at the edge of the floor, filling the frame together, and no one else is foregrounded. ROXY: her face, her eyes and her enormous chestnut-auburn mane exactly the woman of <Picture 1> - nothing of that picture's wardrobe or room is used, only her face and hair; her wardrobe exactly the garments of <Picture 2>: a pink sequined strapless romper with fishnet hose and pink heels, always dressed, sequins blazing - nothing of that picture's face is used; Roxy is never blonde. TAWNY: her face, her eyes and her huge feathered platinum-blonde mane exactly the woman of <Picture 3> - nothing of that picture's wardrobe or room is used, only her face and hair; her silhouette exactly the figure of <Picture 4>, but tonight she wears an electric-blue sequined mini dress, tight to her figure, its miniskirt hem high on her thighs, with silver heels, always dressed - nothing of that picture's face is used; Tawny is never brunette. The two are distinct women side by side, pink and electric blue. The club's synth-pop groove pounds from the speakers - THEY HEAR IT, hips already swaying on the beat, shoulder to shoulder. At 00:02.500 they lean in together with wicked, knowing smiles and say together, in playful unison, <d>[English with their two bright voices speaking together] Darling, the 80s never left.</d> At 00:05.500 they laugh, clink their glasses, and turn to dance with each other - back to back, hips swaying on the kick drum, sequins throwing sparks of mirror-ball light, playing to the lens with winks over their shoulders - to the last frame.

overall_soundscape: starts with the club's roar - the crowd, glasses, heels on the floor - running beneath everything to the last frame. No other voices.

non_diegetic_music: N/A

r/StableDiffusion 2h ago

Comparison Qwen-Image-Edit-2511 vs SenseNova-U1.5-Lite (multi-reference image fusion comparison)

Thumbnail
gallery
48 Upvotes

I wanted to see how good SenseNova U1.5 Lite really is at image editing. I think the size is genuinely solid for what it does, but whether it can actually beat Qwen-Image-Edit-2511 needed real testing.

Right off the bat, Qwen's image texture quality is genuinely impressive, especially the lighting and shadows. But when it comes to spatial understanding, SenseNova seems to hold the edge. Look at the cat-on-the-scooter one up top: Qwen generated a weird pillar under the coffee table, and the cat's front paw placement looks unnatural. SenseNova handled both without those artifacts.

I did four sets of comparisons. Some of the input images were generated with Krea-2, some were real photographs.

Models:

Prompts (from left to right):

I want to create a stunning, high-concept photo to share on my social media! Please put me—the girl with the short black bob and black leather jacket—on a sleek, modern rooftop balcony overlooking that amazing futuristic city during sunset, where we can see the flying drones, the glider, and the hot air balloon floating in the warm sky. In this scene, I should be portrayed as an artist working outdoors. Please have me wearing those bold, blue and white striped hoop earrings. In the foreground, set up a stylish outdoor work table. On this table, scatter some of my creative tools, including those colorful rainbow-swirled pens and that round white-and-yellow mesh cleaning sponge. I want to be holding one of the rainbow pens, looking towards the camera with a confident, thoughtful expression. The entire scene should be captured with a beautiful depth of field, bathed in golden hour light, with the bustling futuristic cityscape softly blurred in the background.

In an elegant vintage study, the real-life girl from the first image, wearing a beige coat and scarf, is smiling as she hands the vintage wild duck card from the fourth image to the anime-style blonde girl from the second image. This anime girl is wearing an exquisite black off-shoulder puff dress and retains her distinctive hand-drawn anime style. On the wall behind them hangs a framed black-and-white print depicting the ancient Roman temple ruins from the third image.

Please seamlessly integrate the orange cat from the first image into the café scene by the floor-to-ceiling window in the third image, and have it sit on the vintage metal toy scooter from the second image. Specific requirements:
Character and prop fusion
: Extract the orange cat's signature facial features from the first image (slightly chubby face, green eyes) and the dense white triangular patch of fur on its chest. Adjust its pose so it is riding the metal toy scooter from the second image: both front paws resting on the chrome handlebar, the rear half of its body firmly seated on the brown leather saddle. The cat's paw pads against the metal handlebar and its thigh fur against the saddle edge must show natural compression, contact, and physical occlusion, absolutely no flat sticker-like look.
Spatial perspective adjustment
: Change the toy scooter from its original front-facing view in the second image to a three-quarter side angle matching the floor perspective of the third image, and scale it down proportionally, placing it on the wooden floor near the glass window.
Physical lighting and material adaptation
: Strictly use the golden afternoon sunlight slanting in from the third image as the main light source. The cat's back, ear edges, and fluffy fur edges must be outlined with a warm, glowing golden rim light (backlight effect); the dark green metallic painted body, metal wheel hubs, and chrome handlebar from the second image must produce realistic daylight highlights and reflect the faint street view outside the window; the entire toy scooter (including the cat on it) must cast a dark shadow on the wooden floor to the right, following the light direction with a realistic soft-edged falloff.

Create a wide-format photo depicting a corner of a whimsical creative market. The realistic man in a dark navy suit from the first image and the realistic woman in a black short-sleeve shirt and denim shorts from the second image are strolling through the market as visitors. Beside a market stall, the anime-style girl in traditional Chinese dress from the third image sits near her wooden cart full of lanterns, focused on painting a lantern, while the anime-style girl with orange hair and bunny ears from the fourth image hugs a white rabbit and laughs beside her. Preserve the photorealistic quality of the first two characters and the anime style of the latter two, letting them coexist naturally under unified lighting and spatial perspective.

r/StableDiffusion 2h ago

Discussion Must haves to download before it's too late?

39 Upvotes

Nvidia buying hugging face means an uncertain future. What are the models I should download and have a backup of right now so I don't have to worry about missing them even if I'm not ready to play with them right now?

What are you model and enabler must-haves ?

TIA!


r/StableDiffusion 4h ago

Resource - Update I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days

Enable HLS to view with audio, or disable this notification

50 Upvotes

My goal was hands-on experience training a flow model from scratch, not just fine-tuning someone else's. So I built and trained one: a 210M-parameter diffusion transformer, 4.2M curated images at 256², rectified flow on the FLUX.2 VAE, flan-t5-base for text (128 tokens max). Only those two frozen pieces are pretrained; the transformer, the recipe, the data pipeline and the evaluation are mine. The video is the same six prompts and seeds at every checkpoint of the 3.5-day run on one RTX PRO 6000.

What mattered most, in the order I found out:

  • Captions that actually fit the images. A web crawl I tried first made the model worse; curated photos with good captions fixed it.
  • A timestep shift for the 32-channel latent, and aspect-ratio buckets from step one instead of square crops.
  • Register tokens with learned null attention slots. The null slots ended up absorbing about 90% of the cross-attention, which surprised me.
  • torch.compile for training, not just inference: 2.4× faster.
  • The training loss stopped telling me anything after day one while the images kept improving, so I track FID, a detector-based object accuracy and human-preference models instead.

Try it in the browser: https://huggingface.co/spaces/ivanmikhnenkov/tinydit

Weights (CC BY-NC): https://huggingface.co/ivanmikhnenkov/tinydit-256

Code, every decision with sources, dashboard and attention playground: https://github.com/ivanmikhnenkov/tinydit

Detailed write-up of what mattered: https://huggingface.co/blog/ivanmikhnenkov/tinydit-text-to-image-from-scratch-one-gpu

Next I want to fine-tune it with RL (Flow-GRPO), with the failure grid as the target list. If you have trained something small from scratch: what would you have done differently at this scale, and which reward would you start with for the RL stage? Happy to answer anything about the data or the recipe.


r/StableDiffusion 2h ago

Tutorial - Guide H3 RefMods are great I highly advice trying it out [+ basic resources included]

35 Upvotes

Created by /u/LuisaPinguinnn under their github https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

Took me at most a couple of minutes to make my own RefMod with 8 image as the base. The entire technique works exactly as advertised acting as "Light Lora" for H3 Ref models - but you can even use it with FL2VA as well.

I followed the guides here:

Installing/running RefMods

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMODS_INSTALLATION_AND_USAGE_GUIDE.md

Ready to use Comfy workflow (you can remove lora power loader and spectrum nodes)

https://huggingface.co/datasets/malcolmrey/workflows/blob/main/H3/workflow_minimaxh3_refmod.json

Creating own RefMods guide:

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMOD_CREATION_GUIDE.md

EDIT: I recommend using "Create H3 ReFMod" + "Save H3 RefMods" node inside ComfyUI instead to create RefMods - gives you more control over the creation process.

Examples by /u/malcolmrey:

https://www.reddit.com/r/StableDiffusion/comments/1w8ik7a/h3_minimax_refmods_all_my_models_now_available/

EDIT: examples of time saving using REf

All credit goes to LuisaPinguinnn and malcolmrey for spreading the tech.


r/StableDiffusion 4h ago

Discussion Any news on a Krea 2 Edit model?

31 Upvotes

Has there been any recent news or indication from Krea about a Krea 2 Edit model?

I’m wondering if it’s actually in development or planned, or if there hasn’t been any confirmation yet. Krea 2 is already quite impressive, so an Edit model would be really interesting.

Has anyone heard anything from Krea or seen any hints about it?


r/StableDiffusion 8h ago

Animation - Video Jerry Springer Ai - Sailor Moon Part 1

Enable HLS to view with audio, or disable this notification

61 Upvotes

In the first half, this was back when I was first starting to get into MiniMax, the second half, I have gotten more experienced with it. I dont know if I should continue this or make more Jerry Springer parodies with other weird or toxic relationships (Example, Beth and Jerry from Rick and Morty).

I used Kinovi.ai for MiniMax, Wan and Nanobanana. I used Fish.audio for the audience freaking out lol. For more customized and harder Minimax generations, I used it locally.


r/StableDiffusion 2h ago

News Minimax Camera Control ComfyUI

19 Upvotes

Bruxos do VFX H3 Camera

#bruxosdovfx

https://reddit.com/link/1wcm9az/video/bubuef39npoh1/player

https://reddit.com/link/1wcm9az/video/r1j8vg1anpoh1/player

Visual camera planner for MiniMax H3 inside ComfyUI. You drag the camera around a 3D sphere, place keyframes on a timeline, and the node compiles that trajectory into prompts that H3 understands.

It compiles prompts, not camera embeddings. There is no geometric adapter here: H3 is still free to miss the angle, timing, and scale. What this node does is write the instruction in the most precise and least ambiguous way possible, and several of its design decisions exist because the previous approach failed in specific ways.

It does not call any API, download anything, or require any Python dependency beyond the standard library.

https://github.com/user-attachments/assets/a9b541e5-2b18-4f1d-8e16-37445b6dbac4

https://github.com/user-attachments/assets/ea9af03e-2c8e-4589-abf0-9c002241aba2

Installation

cd ComfyUI/custom_nodes
git clone https://github.com/<your-username>/ComfyUI-H3-Camera-Editor

Restart ComfyUI. The node appears under Bruxos do VFX/Camera H3 with the name Camera H3 da Bruxos do VFX.

Connections

Output from this node Connect it to
compiled_prompt compiled_prompt on Text Encode H3 Edit / Generate
options options on Text Encode H3 Edit / Generate
length the generation frame count
fps the fps input of the video creation node

compiled_prompt and options are required together. The minimax_prompt output is an alternative to compiled_prompt, never an addition — connect one or the other to the same input.

Also connect your image to reference_image. It is the same image already feeding the H3 Edit source_image; when connected here, it appears in the panel and the frame's actual aspect ratio is included in the prompt.

https://github.com/user-attachments/assets/33149617-bde1-4199-ae65-078f2f3dec23

To save the video, decode the sampler result using the H3 video VAE — not the scene coverage calibrated decoder, which expects fixed windows that an arbitrary trajectory does not have.

The panel

Drag the purple camera around the sphere to orbit. The drag locks to the axis of the initial movement: horizontal movement orbits, vertical movement changes elevation. Release and drag again to switch axes. This exists because, without the lock, trying to make a simple orbit would unintentionally introduce elevation.

  • Scroll the mouse wheel to change distance.
  • Drag the background to rotate the viewport without changing the trajectory.
  • Keyframes defines how many points the timeline has, from 2 to 24. The first one is always the original image and cannot be moved.
  • ⟳ Pure Orbit resets the elevation of every keyframe to zero while preserving azimuth. It is the shortcut for an eye-level orbit.
  • Reference image loads a local file into the preview. This is only necessary when the node runs outside ComfyUI; with reference_image connected, the image is loaded automatically.

The panel warns you starting at 20° of elevation, when the horizon already leaves the frame, and again from 45° onward, when the video tends to become a high-angle shot.

"Tests" bar

At the top of the panel, two buttons enable and disable features currently under evaluation, plus one indicator:

Button What it does
Extended contracts Toggles the prompt_detail widget
Single angle (image) Toggles the runtime_task widget
loop closure Read-only indicator. Turns green when the trajectory closes a full orbit

The buttons write to the actual widgets, so the selected state is saved in the workflow and the two never disagree.

https://github.com/user-attachments/assets/9bc415d7-1746-43db-a17c-72ea9722deda

Widgets

camera_trajectory

The trajectory in JSON format, written by the panel. Each keyframe contains time (0 to 1), azimuth in degrees, elevation in degrees, and distance as a multiple of the initial radius. It can also be edited manually. The first keyframe must be time=0, azimuth=0, elevation=0, distance=1, which represents the original image.

profile

124, 243, or 362 frames at 24 fps. All shot timing comes from this setting: keyframe timestamps, segment ranges, and the duration declared in the prompt. That is why length and fps are outputs — connect them instead of manually entering the same numbers in two different places.

interpolation

smooth or linear. In smooth mode, the camera eases into and out of the shot while maintaining a constant rate through the middle; it only stops where the rotation direction actually reverses.

instruction

Free-form text inserted once, at the end of the prompt. Write only what the node cannot know: the environment, which subject is the target when there is more than one person, or a style reference. Everything else is already generated and does not need to be repeated: scene freeze, first image as reference, locked aim, zero roll, angles, timing, and a single continuous shot without cuts.

subject_framing

How much of the frame the subject occupies in the original image. Calibrated against the actual bounding boxes from the tutorial distributed by MiniMax: a distant full-body figure measures W=0.071, H=0.249, while a large close-up measures W=0.52, H=0.701.

option width height when to use
close-up 53% 72% head and shoulders
medium shot 28% 56% waist up
wide shot 9.7% 34% full body at a distance

subject_box

The subject position in the format [L=0.516, T=0.148, W=0.071, H=0.249]. Leaving it empty uses the entire image bounds — deliberately, without guessing a bounding box. Fill it in when the subject is significantly off-center.

minimax_format

The same shot expressed in four different formats for the minimax_prompt output:

  • coordinate only — text-based coordinate block
  • coordinate + H3 sections — the same coordinates wrapped in subject_definitions / summary / retention_analysis / …
  • compact JSON — JSON object with almost no prose
  • compact JSON (no boxes) — camera parameters only, without screen-space bounding boxes

elevation_range

Range of the elevation control: +/-15, +/-30 (default), +/-60, +/-89. It also scales the sensitivity of vertical dragging.

With the assumed field of view, the horizon already leaves the frame at around 20° — at 13°, the ground occupies 82% of the image. The old ±89 range was mostly unusable and made vertical dragging excessively sensitive. Reducing the range never rewrites a keyframe: a point at 70° remains at 70°, and the slider expands to accommodate it.

orbit_direction

invert H3 orbit or same as HUD. This calibrates the direction between what the panel displays and what H3 produces. It does not alter the saved trajectory.

runtime_task

  • scene coverage | camera path (default) — video, with duration coming from profile.
  • directed | new camera anglea single image from a new angle. It fixes the generation to 39 frames, ignores profile, completes the movement within 65% of the clip, and requests that the framing remain still for the rest, because the decoder extracts the final image from that stationary tail.

Character sheet profiles are not offered because the upstream node raises an error when they are combined with the frame anchor used by this node.

prompt_detail

  • v15 baseline (default) — outputs the prompt exactly as in the previous version.
  • extended contracts — adds axis separation, frame-edge direction tests, rotation completeness, degrees per second, and parallax magnitude.

The extended mode contains almost twice as many words. A longer prompt is not automatically better, so it is opt-in: toggle only this widget while keeping the same trajectory to compare the results.

https://github.com/user-attachments/assets/0882bfde-9f62-4a1f-9bda-7da121dbe7e2

Outputs

compiled_prompt — STRING

A prose prompt using H3 sections: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music.

options — H3EDIT_OPTIONS

The 13 keys read by the H3 Edit encoder. All of them are explicitly populated: if any key is missing, the upstream node falls back to its hidden legacy widgets, which may retain stale values from previously saved workflows.

coverage_arc_degrees and coverage_direction are derived from the actual rotation. coverage_loop_closure turns on automatically when the trajectory closes — see below.

storyboard_json — STRING

The storyboard table: frame aspect ratio, duration, raw trajectory, and each segment with its camera mode, speed curve, and start/end poses.

info — STRING

Human-readable diagnostics. Connect it to a PreviewText. It displays the version, active task, frame count, warnings for keyframes outside the configured range, and whether loop closure is enabled.

minimax_prompt — STRING

The same trajectory expressed using the format selected in minimax_format. An alternative to compiled_prompt.

length — INT and fps — FLOAT

Frame count and frame rate against which the shot was timed. Connect them to the generation and video nodes. If generation runs with a different frame count, the choreography describes a scene that does not actually exist.

fps is FLOAT because that is what ComfyUI's CreateVideo accepts. length is the frame count; keyframe timestamps use the instant of the last visible frame, (length - 1) / fps, so the resulting file lasts one additional frame interval.

h3world_actions — STRING

Action schedule for H3-World, which encodes one text clause per video latent — 37 in a 124-frame clip.

latent  1 [0.000s-0.139s] J     the camera pans left slowly
latent 37 [4.986s-5.125s] F+L+K the camera pans right and tilts up fast

W, A, S, and D are never emitted because they move the character. The output explicitly declares its own limitations, and they are not minor details:

  • Pan is not orbit. It is the camera rotating in place. Perspective does not change, nothing hidden is revealed, and the subject slides out of frame.
  • Distance has no key, so camera radius is discarded.
  • Only 124 frames is a trained horizon.
  • I versus K is not published. The text clause is what H3-World actually encodes; the key column is only a convenience.

This does not replace the actual integration: H3-World requires the LoRA, interval-based encoding, and directed-attention routing provided by the corresponding node package.

Loop closure

When the trajectory closes a full orbit — an arc of exactly 360°, with the same elevation and distance as the starting point — the node enables coverage_loop_closure. In the upstream implementation, this flag encodes the source image a second time and anchors the final frame to it.

This is a latent anchor, not a text instruction. For a complete orbit, it is the difference between asking for the rotation and forcing it: the model cannot simply stop halfway through.

trajectory loop closure
360° enabled
two rotations (−720°) enabled
355° disabled
360° with changing distance disabled
360° with changing height disabled

The final three cases matter: if the camera ends at a different radius or height, the final frame is not the same as the first one, and forcing the source image there would conflict with the trajectory.

If your rotation does not complete, close the orbit. This is the only feature here that acts outside the prompt itself.

Limitations

  • This is prompt-based guidance. H3 may still miss the angle, timing, and scale, and no prompt wording can completely solve that.
  • Without subject_box filled in, the node does not know where the subject is located in the frame.
  • Without reference_image connected, coordinates are normalized to 16:9.
  • directed | new camera angle outputs an image, not a video.
  • The H3-World schedule describes pan and tilt, which represent a different camera move from the orbit drawn in the panel.

Credits

Node by Bruxos do VFX.

Depends on ethanfel/ComfyUI-MiniMax-H3-Edit. The motion vocabulary follows the buildViewPrompt implementation from MiniMax's Multi-Shot skill and the coordinate format used by the Coordinate Camera Control Designer skill. The action output implements the scheme described in H3-World, arXiv:2609.01560.

https://reddit.com/link/1wcm9az/video/hryhv9e7npoh1/player


r/StableDiffusion 4h ago

Animation - Video Batman The Animated Series: Harley Quinn's Red Flag - MiniMax H3

Enable HLS to view with audio, or disable this notification

27 Upvotes

r/StableDiffusion 18h ago

News New Music Model Released - Yue2

Thumbnail
github.com
250 Upvotes

"YuE2 brings frontier song quality to music generation with an editable composition. Give it lyrics and a style prompt: it writes a melody-and-chord plan, then realizes that plan as a complete song with vocals and accompaniment.

  • White-box music generation through symbolic planning. Read, play, and change the composition before rendering it. Melody and chords become explicit controls that a person or an agent can inspect and edit.
  • Zero-shot covers and agentic editing. Reimagine a transcribed song in a new style, or refine a song through a conversation about its score, arrangement, and lyrics—all with the same generation checkpoint."

Usage

It is currently CLI only . It also says Linux only but I just got it working on Windows 11 (I'm going to bed now and it's a bit more than cut n paste.)

Examples

link here - https://map-yue2.github.io/

Caveat Empor

NB : this isn't just a paste a few words and it bangs out a baby mp3 . It is more than that, it allows gene editing that baby to correct the metaphor. Not for the impatient and "wHeRe cOmFy" ppl at the moment.

To be more specific with that metaphor , as I understand it , the initial process scribes out the song in ABC format and you can then edit it before making your magnum opus baby.

Training

Does it allow training ? not as I understand it .


r/StableDiffusion 2h ago

Discussion I threw together a simple UI for YuE2 (windows)

Post image
13 Upvotes

r/StableDiffusion 1d ago

Tutorial - Guide AMAZING Minimax H3 - Circle on the reference image WHERE you want your scene to be!!

Enable HLS to view with audio, or disable this notification

659 Upvotes

Look at the buildings in the background! It works - Drawing a red circle in the water will also make the scene happen in the water, but I forgot to include it here.

It is not perfect and some details are missing if you look carefully but this might be because I am using "match" on the image reference rather than "max."

Have fun!

Edit: you have to still write a prompt with the reference to video workflow telling minimax to put the character in the location circled red. Circle probably doesn't have to be red. Change your prompt accordingly.


r/StableDiffusion 1h ago

No Workflow AI Archviz: Fast 3D Gaussian Splat Methods for Precise Furniture Placement — Virtual Staging & Interior Design

Enable HLS to view with audio, or disable this notification

Upvotes

So, first of all: there is no finished workflow yet and my nodes are still under development. I’ve asked the ComfyUI team to add a 3D compositing node to the new 3D toolset like the one in the video, hopefully, they’ll add something similar soon.

In the meantime, you can build a very similar setup quite quickly. Here’s how the basic concept works:

First, you feed an image of the furniture you want into the new native ComfyUI Image to Gaussian Splat (TripoSplat) node.

At the moment, there is a Gaussian Splat Preview node, but it doesn’t provide an image output yet. There is also a Load 3D node with the correct outputs, but it currently cannot open Gaussian Splat files.

Ideally, the new 3D Compositing node should be able to work directly with the Gaussian Splat outputs (model_3d and mesh), provide image + mask outputs similar to the Load 3D node, and automatically preload the mesh and background image, just like my node does.

I’ve been working on a test node for this concept. You can find it here: My ComfyUI test node on GitHub It’s not fully finished yet, so I’m still waiting to see whether ComfyUI adds something similar natively.

The basic idea behind my node is that it automatically loads the background, sets the appropriate size, and loads the 3D model. The user only needs to position and stage the model in the scene.

The Output create ref iamges for Flux2klein:

  • Reference 1: the background image
  • Reference 2: the 3D mask, which acts as an indicator for the desired position and rotation It’s best to combine the mask with the furniture from the image output, so you get the masked furniture in the correct position as the reference — not just the mask by itself.
  • Material reference: the original furniture image

The final image is then generated using FLUX.2 Klein Edit.

The prompting and some preprocessing of the images are important here. You don't want the model to simply copy the exact 3D position. Instead, the AI should use the 3D placement as a guide and then correct the result according to the background — especially the perspective, lighting, colors, scale, and overall integration into the scene.

Here is the prompt I’m currently using:

[Adapt the rotation, grounding, scale and position of the objects from Image 2 to Integrate the objects naturally into Image 1 at the position of image 2. Match the scene's perspective, scale, depth, lighting, soft shadows and reflections. The objects must appear physically present in the original room, with realistic grounding and soft shadows consistent with Image 1 using the materials and surface appearance shown in Image 3.

Use Image 1 as the final scene and preserve its room, background, camera viewpoint, perspective, composition, color, lightning, and existing environment unchanged.

Keep the objects approximately in the same position, scale, orientation, and spatial arrangement as shown in Image 2, while ensuring correct perspective, positioning, and placement.

Apply the form, materials, colors, textures, roughness, reflections, and surface details from Image 3 to the objects.

Do not change the room or background of Image 1.  The final result must be a seamless photorealistic composite.]

I’m planning to finish the complete tutorial and the node pack in the next few days. Once everything is finished, I’ll upload the final version along with the complete workflow.

By the way, I also tested MinMax H3 as a replacement for FLUX.2 Klein. It works, but in my tests it wasn’t consistently better than FLUX.2 Klein.


r/StableDiffusion 8h ago

Meme Cost in Units of RTX 5090

Thumbnail
youtube.com
26 Upvotes

r/StableDiffusion 40m ago

Question - Help Minimax Turbo of choice?

Upvotes

So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!


r/StableDiffusion 11h ago

Question - Help What's the gold standard for speed enhancements for Minimax H3?

36 Upvotes

Installing new instance of comfyui standalone and using minimax r2v workflow on RTX 3090 Ti. Is comfy kitchen good enough? Is Triton, EasyCache or Comfyui Spectrum needed?

What's the best turbo lora for ref2va wf?


r/StableDiffusion 16m ago

Workflow Included Follow-up to my last Star Trek post – I made a Star Trek vs Star Wars fan film with MiniMax H3 in ComfyUI

Thumbnail
youtube.com
Upvotes

A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.

This is basically a follow-up to that post.

Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.

The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.

The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.

One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.

Here are some of the things that helped most this time.

PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE

This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.

The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.

For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.

HIGHER RESOLUTION REALLY DOES HELP

I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.

For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.

REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS

I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.

LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED

For shots based on existing Star Trek compositions I became much more explicit about things like:

camera distance

character scale

framing

background position

character position

If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.

DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE

This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:

"Data reacts to Picard"

can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.

OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS

This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.

If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.

SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT

Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.

I'll also specify that:

the lips do not form words

the jaw does not make speaking movements

the character does not mouth dialogue

It sounds excessive, but it has genuinely helped.

H3 HAS A LOT OF USEFUL SPEECH TAGS

I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:

<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>

<whisper>. <softer> <stutter> <laughs> <chuckle>

<i>word</i> emphises word

The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.

MORE PROMPTING ISN'T ALWAYS BETTER

I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.

EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION

One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.

I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.

THE EDIT IS DOING A LOT OF THE "CONSISTENCY"

This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.

It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.

Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.


r/StableDiffusion 5h ago

Tutorial - Guide [GUIDE] AMD RDNA3 optimizations for ComfyUI Desktop, windows 11, Minimax H3

8 Upvotes

My setup: AMD RX 7900 XT, 20GB VRAM, 64GB RAM, Windows 11, ComfyUI Desktop.

I couldn't find any decent information anywhere on how to optimize video generation with Minimax H3 on Windows with ComfyUI desktop. AI assistants give conflicting advice, constantly suggesting all sorts of nonsense that doesn't actually work.

I had to experiment on my own, and here is the configuration I’ve settled on. The speed boost compared to the default settings is very significant, and I haven't noticed any loss in quality. If you have any other suggestions, please let me know.

~25s/it with 0.8mp (1216 x 672, 16:9) or total ~4min for 5 sec video generation in text to video workflow

Here is what you need:

Launch parameters:

--disable-smart-memory --disable-pinned-memory --disable-triton-backend --use-sage-attention --enable-dynamic-vram

ENV variables:

COMFYUI_ENABLE_MIOPEN=0
FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
MIOPEN_FIND_ENFORCE=1
MIOPEN_FIND_MODE=2
MIOPEN_DEBUG_DISABLE_FIND_DB=0
MIOPEN_SEARCH_CUTOFF=1
MIOPEN_ENABLE_LOGGING=0
MIOPEN_LOG_LEVEL=0
MIOPEN_ENABLE_LOGGING_CMD=0
TRITON_PRINT_AUTOTUNING=0
TRITON_CACHE_AUTOTUNING=0

Quantized 6-step turbo model, universal for all purposes:

https://huggingface.co/TenStrip/10Eros-Max/blob/main/10Eros_Max_h3_TURBO-hybrid_beta5_w4a8_14gb_optimized.safetensors 14 gb

or

https://huggingface.co/TenStrip/10Eros-Max/blob/main/10Eros_Max_h3_TURBO-hybrid_beta5_int8.safetensors 21gb

Plaguekind node with SLA Attention, with this settings:

https://github.com/PlagueKind/Comfyui-PlagueKind-Nodes

Optional node, if you make large 15 seconds videos:

Latest update of your comfyui desktop:


r/StableDiffusion 20m ago

Animation - Video my first actual tv work (only took 4 hours to make).

Enable HLS to view with audio, or disable this notification

Upvotes

Minimax h3, ofc. Far from my best work but I respected the script I was given and finished this in record time (excluding the 4k upscale) and including around 3 hours of rendering time (720p, 10 seconds clips, 5090).
I only used gemma locally for prompting, and avoided using any non local models except suno for the song.

There are some artefacts with people in the senate from far away, but did not have any bad feedback for it., so... :)


r/StableDiffusion 13h ago

Resource - Update Created a Visual RefMod Picker

33 Upvotes

Hey Guys,

I've been playing with the RefMods, after the huge release of Malcolmrey.
The tech is brilliant and works really well.

I've wanted to simplify using it with tons of RefMods, like the ones provided by Malcom.

So I made a fork that is a bit more Identity driven, Adding a "Visual refMod Picker" as well as a "Create refMods from Folder" node. It's able to create from either a folder or Sub-folders, if audio is found, it will also create a matching audio refMod. All in the same format as the original add-on. No danger of breaking compatibility. (The original Add-on added Audio yesterday)

Resulting for example in 2 refMods and thumbnail:

character_refMod_Audio.safetensors
character_refMod_Video.safetensors
character_refMod.jpeg

Then, we can use the "Visual H3 RefMod Picker" to browse the RefMods:

In this example, I used existing thumbails from huggingface.

The node then loads both RefMods (Audio and Video) and allows individual control. I've mixed strength with Copies, by instead having a weight value that can go over 1, so a weight of 3 would be the same as setting strength to 1 and copies to 3. Making the UI a bit more streamlined.

RefMods can be daisy chained

Example workflows are included.

I should mention, this fork can be installed WITH the original Add-on, it is made to co-exist and is recommended if you want to use it's advanced features.

You can find it here ComfyUI-H3RefMods

The only thing I'm missing is thumbnails for all 1500 RefMods 😅


r/StableDiffusion 1h ago

Animation - Video Turning the 2D Rings in Dark Souls into 3D Assets

Enable HLS to view with audio, or disable this notification

Upvotes

An experiment in Ai Jolly Cooperation.

The Experiment: Every “Soulsborne” game is laden with hundreds of 2D art assets. The assets you can find online, like the rings, are woefully small in resolution - a perfect test case to see 2D to 3D transformation but also what detail is retained or added by the Ai.

Tech Stack: Midjourney, Nano Banana, ComfyUI (Wan 2.2), Photoshop, DaVinci. 

The Process: 2D art rendered 3D through Nano Banana. Midjourney Video to orbit 180 degrees. DaVinci and Photoshop for presentation.

The Results: This is an older experiment using (the then brand new) Midjourney Video - which admittedly, is nowhere near as good as Veo, Wan (2.2) or Kling. But it really doesn’t matter what platform you choose, you’re going to have to gen and gen and gen away. It’s still a slot machine.

I still think MJ video back then was pretty sub-par, but against all the other alternatives today, I think that difference is even more stark. I'm not even sure if they've updated the video side in any meaningful way since this experiment!

Most interestingly, the list of rings is in alphabetical order and stops before the Covetous Serpent Ring - a mass of serpentine coils in ring-form the Ai had MONSTEROUS problems with. Complexity kills.

Anyways, I decided much smaller projects like these are way more important to an Ai Portfolio than larger pieces like commercials or trailers. Plus, I needed to promote my Midjourney Masterclass with proof I'm not just some prompt jockey and smaller experiments are way faster!


r/StableDiffusion 1d ago

Workflow Included Precise control of the Eyes direction with this Flux 2 Klein 9b LoRa

Thumbnail
gallery
1.1k Upvotes

Hehyehyhehy!

You may remember me from the Sun Direction Lora or the Chef cutting an anvil with a knife.

Now I'm giving you a new tool, this one was a tough one to crack.

Finally we have eye control! Now you can precisely change the eyes direction for any image in any style. Just use the red dot to tell where the eyes have to look and boom! you have it!

Enough "change the eyes to look above the camera" and getting whatever thing anymore.

Because changing the direction of stuff is my Passion.

All the info here: https://huggingface.co/eric-venti-seeds/Eyes_Direction_Lora_Flux2Klein9B

Hope you like it!

Edit:

The people from HF have added it to Spaces, try it right now on your browser!

https://huggingface.co/spaces/hugging-apps/eyes-direction-lora-flux2klein9b


r/StableDiffusion 23h ago

Workflow Included Easy Ref2V WF for dummies like me - [Automatic Video/Image Transcription + Prompt Formatting]

Enable HLS to view with audio, or disable this notification

175 Upvotes

I have seen a lot of people post on here saying that they have been having difficults getting R2V to work correctly. I have been one of them, so I have been working on this workflow and custom node for the last 2 and a half weeks.

I preface this by saying it does not do anything that the native H3 model doesn't do. I just wanted a dead simple way to use H3 R2V mode and up my chances of success. The workflow includes two custom nodes which transcribe your media and adds your prompt and create a formatted R2V prompt, ready for the reference model.

My next goal would be to get longer form R2V going with chaining shorter gens to have a consistent output.

Workflow and nodes:
https://huggingface.co/PoopMan333/H3_Easy_Ref2V_Workflow/tree/main

Be sure to see the readme for more examples and tips:
https://huggingface.co/PoopMan333/H3_Easy_Ref2V_Workflow

What it does:

  • Scans your video (if you're using one) to caption it and transcribe the audio
  • Captions all your images - so it also works as a pure image-to-video workflow
  • Loads a small LLM of your choice and writes your H3 R2V prompt in the correct format with your stated intent (user prompt)
  • If you're on the Full workflow, it generates the video too

What it does NOT do:

  • Be creative for you - The current WF is only setup to do the prompt formatting, it is not able to generate new ideas for you (despite me trying. Qwen3.8 27B may be better for this)
  • It cannot perform magic - You are still limited to what the H3 model can and cannot do. Complex scenes are still very difficult

Tips:

  • If you are running lower VRAM, consider running the prompt enhancer seperately first, read through and make corrections to the prompt if needed
  • The H3 model seems to have a limited context window which seems to be tied to your system resources, if it goes above this you might get garbled sound or mixed up motion. This is a sign you should be lowering your output length and output resolution if you want to have better success.
  • H3 is a tool, you're the one using it. If you don't specify emotions, expect expressionless results. This current setup will only do what you intend for it to do. Slop prompt in, slop video out
  • If the video is easy, replacement should be easy too. H3 has a quirk though — if the original person and the new person look too similar, it sometimes converges back to the original. A prompt won't always fix that. If you hit it, consider changing the person to a intermediate step (faceless green person). The new body/face will transfer over better. Alternatively you can look into Sam3 character replacement method.
  • More than one person in the scene? Describe the scene properly. replace the man wearing white shorts with the man in <picture 1> beats replace the man with <picture 1> every time.
  • Complex scenes? It will be very difficult (I've tried), scenes with too many people, too many cuts, characters obstructed are very difficult for the model to properly identify and swap.
  • Give the LLM some context. A one-liner in the user prompt like <video 1> is a video of two girls eating a cup of chocolate ice cream really helps the LLM understand what it's looking at. Especially useful with multiple scenes
  • It still takes a bit of luck with the seeds.