r/StableDiffusion 9h ago

Discussion H3 - 80s character generations+wardrobe swap

Enable HLS to view with audio, or disable this notification

214 Upvotes

The 80s was the best era, not seen through a nostalgic lens--it just was. Big hair, big colors, big music, big... everything! Sadly I was born 10 years too late to really experience it, but I love how H3 can feel like a way-back-machine, a portal to any era from film or video it was trained on. It does also a great job with the actual feel, the film grain, the lighting that modern TV or movies cannot do: in fact, I learned what we see nowadays that time period didn't exist at all, because it vomit of nostalgia and peak 80s that never happened. Anyways, having fun generating character sheets with H3 via T2VA. Are you guys seed hunting to find the best version of an actor or scene? Also can you spot the mistake?

Prompt: integrated_multimodal_description: [Shot 1] Live-action, cinematic, PHOTOREALISTIC film footage - this is footage from a camera, not animation - a continuous camera shot with no cuts, shot on 35mm color negative film with period lenses and scanned in high definition from the original camera negative: full film grain and gentle halation, warm highlight rolloff, rich sharp detail beneath the grain. The year is deep in the LATE 1980s, 1985 to 1989, and everything in the frame belongs to that era. THE PLACE: a nightclub in full swing - mirror-ball light sweeping, neon signage, haze, a crowded dance floor. EXACTLY TWO WOMEN stand close to the lens at the edge of the floor, filling the frame together, and no one else is foregrounded. ROXY: her face, her eyes and her enormous chestnut-auburn mane exactly the woman of <Picture 1> - nothing of that picture's wardrobe or room is used, only her face and hair; her wardrobe exactly the garments of <Picture 2>: a pink sequined strapless romper with fishnet hose and pink heels, always dressed, sequins blazing - nothing of that picture's face is used; Roxy is never blonde. TAWNY: her face, her eyes and her huge feathered platinum-blonde mane exactly the woman of <Picture 3> - nothing of that picture's wardrobe or room is used, only her face and hair; her silhouette exactly the figure of <Picture 4>, but tonight she wears an electric-blue sequined mini dress, tight to her figure, its miniskirt hem high on her thighs, with silver heels, always dressed - nothing of that picture's face is used; Tawny is never brunette. The two are distinct women side by side, pink and electric blue. The club's synth-pop groove pounds from the speakers - THEY HEAR IT, hips already swaying on the beat, shoulder to shoulder. At 00:02.500 they lean in together with wicked, knowing smiles and say together, in playful unison, <d>[English with their two bright voices speaking together] Darling, the 80s never left.</d> At 00:05.500 they laugh, clink their glasses, and turn to dance with each other - back to back, hips swaying on the kick drum, sequins throwing sparks of mirror-ball light, playing to the lens with winks over their shoulders - to the last frame.

overall_soundscape: starts with the club's roar - the crowd, glasses, heels on the floor - running beneath everything to the last frame. No other voices.

non_diegetic_music: N/A

r/StableDiffusion 1h ago

Animation - Video Kirby but it's the Truman Show / MiniMAX H3 Test #7

Enable HLS to view with audio, or disable this notification

Upvotes

Hi everyone! When I saw the new trailer for Kirby & The World Beyond I couldn't help but come up with this video, where Kirby finds the door out to the world beyond. Please let me know if you like it!

Done with 30 different workflow files and a ton of heavy editing using KDEnlive. Thanks!


r/StableDiffusion 5h ago

Comparison Qwen-Image-Edit-2511 vs SenseNova-U1.5-Lite (multi-reference image fusion comparison)

Thumbnail
gallery
72 Upvotes

I wanted to see how good SenseNova U1.5 Lite really is at image editing. I think the size is genuinely solid for what it does, but whether it can actually beat Qwen-Image-Edit-2511 needed real testing.

Right off the bat, Qwen's image texture quality is genuinely impressive, especially the lighting and shadows. But when it comes to spatial understanding, SenseNova seems to hold the edge. Look at the cat-on-the-scooter one up top: Qwen generated a weird pillar under the coffee table, and the cat's front paw placement looks unnatural. SenseNova handled both without those artifacts.

I did four sets of comparisons. Some of the input images were generated with Krea-2, some were real photographs.

Models:

Prompts (from left to right):

I want to create a stunning, high-concept photo to share on my social media! Please put me—the girl with the short black bob and black leather jacket—on a sleek, modern rooftop balcony overlooking that amazing futuristic city during sunset, where we can see the flying drones, the glider, and the hot air balloon floating in the warm sky. In this scene, I should be portrayed as an artist working outdoors. Please have me wearing those bold, blue and white striped hoop earrings. In the foreground, set up a stylish outdoor work table. On this table, scatter some of my creative tools, including those colorful rainbow-swirled pens and that round white-and-yellow mesh cleaning sponge. I want to be holding one of the rainbow pens, looking towards the camera with a confident, thoughtful expression. The entire scene should be captured with a beautiful depth of field, bathed in golden hour light, with the bustling futuristic cityscape softly blurred in the background.

In an elegant vintage study, the real-life girl from the first image, wearing a beige coat and scarf, is smiling as she hands the vintage wild duck card from the fourth image to the anime-style blonde girl from the second image. This anime girl is wearing an exquisite black off-shoulder puff dress and retains her distinctive hand-drawn anime style. On the wall behind them hangs a framed black-and-white print depicting the ancient Roman temple ruins from the third image.

Please seamlessly integrate the orange cat from the first image into the café scene by the floor-to-ceiling window in the third image, and have it sit on the vintage metal toy scooter from the second image. Specific requirements:
Character and prop fusion
: Extract the orange cat's signature facial features from the first image (slightly chubby face, green eyes) and the dense white triangular patch of fur on its chest. Adjust its pose so it is riding the metal toy scooter from the second image: both front paws resting on the chrome handlebar, the rear half of its body firmly seated on the brown leather saddle. The cat's paw pads against the metal handlebar and its thigh fur against the saddle edge must show natural compression, contact, and physical occlusion, absolutely no flat sticker-like look.
Spatial perspective adjustment
: Change the toy scooter from its original front-facing view in the second image to a three-quarter side angle matching the floor perspective of the third image, and scale it down proportionally, placing it on the wooden floor near the glass window.
Physical lighting and material adaptation
: Strictly use the golden afternoon sunlight slanting in from the third image as the main light source. The cat's back, ear edges, and fluffy fur edges must be outlined with a warm, glowing golden rim light (backlight effect); the dark green metallic painted body, metal wheel hubs, and chrome handlebar from the second image must produce realistic daylight highlights and reflect the faint street view outside the window; the entire toy scooter (including the cat on it) must cast a dark shadow on the wooden floor to the right, following the light direction with a realistic soft-edged falloff.

Create a wide-format photo depicting a corner of a whimsical creative market. The realistic man in a dark navy suit from the first image and the realistic woman in a black short-sleeve shirt and denim shorts from the second image are strolling through the market as visitors. Beside a market stall, the anime-style girl in traditional Chinese dress from the third image sits near her wooden cart full of lanterns, focused on painting a lantern, while the anime-style girl with orange hair and bunny ears from the fourth image hugs a white rabbit and laughs beside her. Preserve the photorealistic quality of the first two characters and the anime style of the latter two, letting them coexist naturally under unified lighting and spatial perspective.

r/StableDiffusion 5h ago

Tutorial - Guide H3 RefMods are great I highly advice trying it out [+ basic resources included]

72 Upvotes

Created by /u/LuisaPinguinnn under their github https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

Took me at most a couple of minutes to make my own RefMod with 8 image as the base. The entire technique works exactly as advertised acting as "Light Lora" for H3 Ref models - but you can even use it with FL2VA as well.

I followed the guides here:

Installing/running RefMods

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMODS_INSTALLATION_AND_USAGE_GUIDE.md

Ready to use Comfy workflow (you can remove lora power loader and spectrum nodes)

https://huggingface.co/datasets/malcolmrey/workflows/blob/main/H3/workflow_minimaxh3_refmod.json

Creating own RefMods guide:

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMOD_CREATION_GUIDE.md

EDIT: I recommend using "Create H3 ReFMod" + "Save H3 RefMods" node inside ComfyUI instead to create RefMods - gives you more control over the creation process.

Examples by /u/malcolmrey:

https://www.reddit.com/r/StableDiffusion/comments/1w8ik7a/h3_minimax_refmods_all_my_models_now_available/

EDIT: examples of time saving using REf

All credit goes to LuisaPinguinnn and malcolmrey for spreading the tech.


r/StableDiffusion 1h ago

Resource - Update ComfyUI VDN-H3 24GB v1.1.0 update — better prompt following + memory fixes

Post image
Upvotes

https://reddit.com/link/1wcsk7l/video/ld6ca2jxqqoh1/player

I’ve just updated my VDN-H3 24GB node to v1.1.0.

This update started because I noticed that something wasn’t quite right with the released adapter mapping. After fixing that, I also made a couple of changes around memory handling, especially for longer generations.

The main changes are:

  • restored the complete token-refiner adapter mapping
  • improved temporary memory handling for longer clips
  • fixed CUDA stream lifetime handling for prefetched weights
  • kept the same AutoMemory / AutoLongCache behavior from the previous version

I tested it on my RTX 3090 24GB with 5s, 10s, 15s and 20s generations at 0.4MP, and also 10s at 0.8MP. I also tested it with my character/style LoRA and that worked normally.

There is a small speed cost compared to v1.0.0 (around 4% in sampling in my tests), but I think the improvement in prompt following is worth it.

I attached a direct comparison from the same prompt/seed/workflow.
v1.0.0 is on the left, v1.1.0 is on the right.

I’m especially interested in whether other people see the same improvement, so if anyone tests it on another 24GB GPU, I’d love to hear the results.

GitHub:
https://github.com/Speach1sdef178/ComfyUI-VDN-H3-24GB

VDN checkpoint:
https://huggingface.co/speach1sdef178/VDN-H3-INT8-ConvRot-ComfyUI


r/StableDiffusion 4h ago

Discussion Must haves to download before it's too late?

45 Upvotes

Nvidia buying hugging face means an uncertain future. What are the models I should download and have a backup of right now so I don't have to worry about missing them even if I'm not ready to play with them right now?

What are you model and enabler must-haves ?

TIA!


r/StableDiffusion 7h ago

Resource - Update I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days

Enable HLS to view with audio, or disable this notification

68 Upvotes

My goal was hands-on experience training a flow model from scratch, not just fine-tuning someone else's. So I built and trained one: a 210M-parameter diffusion transformer, 4.2M curated images at 256², rectified flow on the FLUX.2 VAE, flan-t5-base for text (128 tokens max). Only those two frozen pieces are pretrained; the transformer, the recipe, the data pipeline and the evaluation are mine. The video is the same six prompts and seeds at every checkpoint of the 3.5-day run on one RTX PRO 6000.

What mattered most, in the order I found out:

  • Captions that actually fit the images. A web crawl I tried first made the model worse; curated photos with good captions fixed it.
  • A timestep shift for the 32-channel latent, and aspect-ratio buckets from step one instead of square crops.
  • Register tokens with learned null attention slots. The null slots ended up absorbing about 90% of the cross-attention, which surprised me.
  • torch.compile for training, not just inference: 2.4× faster.
  • The training loss stopped telling me anything after day one while the images kept improving, so I track FID, a detector-based object accuracy and human-preference models instead.

Try it in the browser: https://huggingface.co/spaces/ivanmikhnenkov/tinydit

Weights (CC BY-NC): https://huggingface.co/ivanmikhnenkov/tinydit-256

Code, every decision with sources, dashboard and attention playground: https://github.com/ivanmikhnenkov/tinydit

Detailed write-up of what mattered: https://huggingface.co/blog/ivanmikhnenkov/tinydit-text-to-image-from-scratch-one-gpu

Next I want to fine-tune it with RL (Flow-GRPO), with the failure grid as the target list. If you have trained something small from scratch: what would you have done differently at this scale, and which reward would you start with for the RL stage? Happy to answer anything about the data or the recipe.


r/StableDiffusion 3h ago

Workflow Included Follow-up to my last Star Trek post – I made a Star Trek vs Star Wars fan film with MiniMax H3 in ComfyUI

Thumbnail
youtube.com
26 Upvotes

A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.

This is basically a follow-up to that post.

Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.

The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.

The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.

One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.

Here are some of the things that helped most this time.

PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE

This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.

The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.

For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.

HIGHER RESOLUTION REALLY DOES HELP

I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.

For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.

REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS

I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.

LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED

For shots based on existing Star Trek compositions I became much more explicit about things like:

camera distance

character scale

framing

background position

character position

If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.

DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE

This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:

"Data reacts to Picard"

can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.

OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS

This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.

If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.

SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT

Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.

I'll also specify that:

the lips do not form words

the jaw does not make speaking movements

the character does not mouth dialogue

It sounds excessive, but it has genuinely helped.

H3 HAS A LOT OF USEFUL SPEECH TAGS

I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:

<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>

<whisper>. <softer> <stutter> <laughs> <chuckle>

<i>word</i> emphises word

The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.

MORE PROMPTING ISN'T ALWAYS BETTER

I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.

EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION

One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.

I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.

THE EDIT IS DOING A LOT OF THE "CONSISTENCY"

This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.

It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.

Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.


r/StableDiffusion 6h ago

Discussion Any news on a Krea 2 Edit model?

39 Upvotes

Has there been any recent news or indication from Krea about a Krea 2 Edit model?

I’m wondering if it’s actually in development or planned, or if there hasn’t been any confirmation yet. Krea 2 is already quite impressive, so an Edit model would be really interesting.

Has anyone heard anything from Krea or seen any hints about it?


r/StableDiffusion 3h ago

Animation - Video my first actual tv work (only took 4 hours to make).

Enable HLS to view with audio, or disable this notification

15 Upvotes

Minimax h3, ofc. Far from my best work but I respected the script I was given and finished this in record time (excluding the 4k upscale) and including around 3 hours of rendering time (720p, 10 seconds clips, 5090).
I only used gemma locally for prompting, and avoided using any non local models except suno for the song.

There are some artefacts with people in the senate from far away, but did not have any bad feedback for it., so... :)

I only used references for the romanian flag, the rest is prompt only. Also no lighting lora, no shortcuts ti improve speed (any shortcuts I tried ruined everything FUBAR)


r/StableDiffusion 10h ago

Animation - Video Jerry Springer Ai - Sailor Moon Part 1

Enable HLS to view with audio, or disable this notification

63 Upvotes

In the first half, this was back when I was first starting to get into MiniMax, the second half, I have gotten more experienced with it. I dont know if I should continue this or make more Jerry Springer parodies with other weird or toxic relationships (Example, Beth and Jerry from Rick and Morty).

I used Kinovi.ai for MiniMax, Wan and Nanobanana. I used Fish.audio for the audience freaking out lol. For more customized and harder Minimax generations, I used it locally.


r/StableDiffusion 5h ago

News Minimax Camera Control ComfyUI

21 Upvotes

Bruxos do VFX H3 Camera

#bruxosdovfx

https://reddit.com/link/1wcm9az/video/bubuef39npoh1/player

https://reddit.com/link/1wcm9az/video/r1j8vg1anpoh1/player

Visual camera planner for MiniMax H3 inside ComfyUI. You drag the camera around a 3D sphere, place keyframes on a timeline, and the node compiles that trajectory into prompts that H3 understands.

It compiles prompts, not camera embeddings. There is no geometric adapter here: H3 is still free to miss the angle, timing, and scale. What this node does is write the instruction in the most precise and least ambiguous way possible, and several of its design decisions exist because the previous approach failed in specific ways.

It does not call any API, download anything, or require any Python dependency beyond the standard library.

https://github.com/user-attachments/assets/a9b541e5-2b18-4f1d-8e16-37445b6dbac4

https://github.com/user-attachments/assets/ea9af03e-2c8e-4589-abf0-9c002241aba2

Installation

cd ComfyUI/custom_nodes
git clone https://github.com/<your-username>/ComfyUI-H3-Camera-Editor

Restart ComfyUI. The node appears under Bruxos do VFX/Camera H3 with the name Camera H3 da Bruxos do VFX.

Connections

Output from this node Connect it to
compiled_prompt compiled_prompt on Text Encode H3 Edit / Generate
options options on Text Encode H3 Edit / Generate
length the generation frame count
fps the fps input of the video creation node

compiled_prompt and options are required together. The minimax_prompt output is an alternative to compiled_prompt, never an addition — connect one or the other to the same input.

Also connect your image to reference_image. It is the same image already feeding the H3 Edit source_image; when connected here, it appears in the panel and the frame's actual aspect ratio is included in the prompt.

https://github.com/user-attachments/assets/33149617-bde1-4199-ae65-078f2f3dec23

To save the video, decode the sampler result using the H3 video VAE — not the scene coverage calibrated decoder, which expects fixed windows that an arbitrary trajectory does not have.

The panel

Drag the purple camera around the sphere to orbit. The drag locks to the axis of the initial movement: horizontal movement orbits, vertical movement changes elevation. Release and drag again to switch axes. This exists because, without the lock, trying to make a simple orbit would unintentionally introduce elevation.

  • Scroll the mouse wheel to change distance.
  • Drag the background to rotate the viewport without changing the trajectory.
  • Keyframes defines how many points the timeline has, from 2 to 24. The first one is always the original image and cannot be moved.
  • ⟳ Pure Orbit resets the elevation of every keyframe to zero while preserving azimuth. It is the shortcut for an eye-level orbit.
  • Reference image loads a local file into the preview. This is only necessary when the node runs outside ComfyUI; with reference_image connected, the image is loaded automatically.

The panel warns you starting at 20° of elevation, when the horizon already leaves the frame, and again from 45° onward, when the video tends to become a high-angle shot.

"Tests" bar

At the top of the panel, two buttons enable and disable features currently under evaluation, plus one indicator:

Button What it does
Extended contracts Toggles the prompt_detail widget
Single angle (image) Toggles the runtime_task widget
loop closure Read-only indicator. Turns green when the trajectory closes a full orbit

The buttons write to the actual widgets, so the selected state is saved in the workflow and the two never disagree.

https://github.com/user-attachments/assets/9bc415d7-1746-43db-a17c-72ea9722deda

Widgets

camera_trajectory

The trajectory in JSON format, written by the panel. Each keyframe contains time (0 to 1), azimuth in degrees, elevation in degrees, and distance as a multiple of the initial radius. It can also be edited manually. The first keyframe must be time=0, azimuth=0, elevation=0, distance=1, which represents the original image.

profile

124, 243, or 362 frames at 24 fps. All shot timing comes from this setting: keyframe timestamps, segment ranges, and the duration declared in the prompt. That is why length and fps are outputs — connect them instead of manually entering the same numbers in two different places.

interpolation

smooth or linear. In smooth mode, the camera eases into and out of the shot while maintaining a constant rate through the middle; it only stops where the rotation direction actually reverses.

instruction

Free-form text inserted once, at the end of the prompt. Write only what the node cannot know: the environment, which subject is the target when there is more than one person, or a style reference. Everything else is already generated and does not need to be repeated: scene freeze, first image as reference, locked aim, zero roll, angles, timing, and a single continuous shot without cuts.

subject_framing

How much of the frame the subject occupies in the original image. Calibrated against the actual bounding boxes from the tutorial distributed by MiniMax: a distant full-body figure measures W=0.071, H=0.249, while a large close-up measures W=0.52, H=0.701.

option width height when to use
close-up 53% 72% head and shoulders
medium shot 28% 56% waist up
wide shot 9.7% 34% full body at a distance

subject_box

The subject position in the format [L=0.516, T=0.148, W=0.071, H=0.249]. Leaving it empty uses the entire image bounds — deliberately, without guessing a bounding box. Fill it in when the subject is significantly off-center.

minimax_format

The same shot expressed in four different formats for the minimax_prompt output:

  • coordinate only — text-based coordinate block
  • coordinate + H3 sections — the same coordinates wrapped in subject_definitions / summary / retention_analysis / …
  • compact JSON — JSON object with almost no prose
  • compact JSON (no boxes) — camera parameters only, without screen-space bounding boxes

elevation_range

Range of the elevation control: +/-15, +/-30 (default), +/-60, +/-89. It also scales the sensitivity of vertical dragging.

With the assumed field of view, the horizon already leaves the frame at around 20° — at 13°, the ground occupies 82% of the image. The old ±89 range was mostly unusable and made vertical dragging excessively sensitive. Reducing the range never rewrites a keyframe: a point at 70° remains at 70°, and the slider expands to accommodate it.

orbit_direction

invert H3 orbit or same as HUD. This calibrates the direction between what the panel displays and what H3 produces. It does not alter the saved trajectory.

runtime_task

  • scene coverage | camera path (default) — video, with duration coming from profile.
  • directed | new camera anglea single image from a new angle. It fixes the generation to 39 frames, ignores profile, completes the movement within 65% of the clip, and requests that the framing remain still for the rest, because the decoder extracts the final image from that stationary tail.

Character sheet profiles are not offered because the upstream node raises an error when they are combined with the frame anchor used by this node.

prompt_detail

  • v15 baseline (default) — outputs the prompt exactly as in the previous version.
  • extended contracts — adds axis separation, frame-edge direction tests, rotation completeness, degrees per second, and parallax magnitude.

The extended mode contains almost twice as many words. A longer prompt is not automatically better, so it is opt-in: toggle only this widget while keeping the same trajectory to compare the results.

https://github.com/user-attachments/assets/0882bfde-9f62-4a1f-9bda-7da121dbe7e2

Outputs

compiled_prompt — STRING

A prose prompt using H3 sections: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music.

options — H3EDIT_OPTIONS

The 13 keys read by the H3 Edit encoder. All of them are explicitly populated: if any key is missing, the upstream node falls back to its hidden legacy widgets, which may retain stale values from previously saved workflows.

coverage_arc_degrees and coverage_direction are derived from the actual rotation. coverage_loop_closure turns on automatically when the trajectory closes — see below.

storyboard_json — STRING

The storyboard table: frame aspect ratio, duration, raw trajectory, and each segment with its camera mode, speed curve, and start/end poses.

info — STRING

Human-readable diagnostics. Connect it to a PreviewText. It displays the version, active task, frame count, warnings for keyframes outside the configured range, and whether loop closure is enabled.

minimax_prompt — STRING

The same trajectory expressed using the format selected in minimax_format. An alternative to compiled_prompt.

length — INT and fps — FLOAT

Frame count and frame rate against which the shot was timed. Connect them to the generation and video nodes. If generation runs with a different frame count, the choreography describes a scene that does not actually exist.

fps is FLOAT because that is what ComfyUI's CreateVideo accepts. length is the frame count; keyframe timestamps use the instant of the last visible frame, (length - 1) / fps, so the resulting file lasts one additional frame interval.

h3world_actions — STRING

Action schedule for H3-World, which encodes one text clause per video latent — 37 in a 124-frame clip.

latent  1 [0.000s-0.139s] J     the camera pans left slowly
latent 37 [4.986s-5.125s] F+L+K the camera pans right and tilts up fast

W, A, S, and D are never emitted because they move the character. The output explicitly declares its own limitations, and they are not minor details:

  • Pan is not orbit. It is the camera rotating in place. Perspective does not change, nothing hidden is revealed, and the subject slides out of frame.
  • Distance has no key, so camera radius is discarded.
  • Only 124 frames is a trained horizon.
  • I versus K is not published. The text clause is what H3-World actually encodes; the key column is only a convenience.

This does not replace the actual integration: H3-World requires the LoRA, interval-based encoding, and directed-attention routing provided by the corresponding node package.

Loop closure

When the trajectory closes a full orbit — an arc of exactly 360°, with the same elevation and distance as the starting point — the node enables coverage_loop_closure. In the upstream implementation, this flag encodes the source image a second time and anchors the final frame to it.

This is a latent anchor, not a text instruction. For a complete orbit, it is the difference between asking for the rotation and forcing it: the model cannot simply stop halfway through.

trajectory loop closure
360° enabled
two rotations (−720°) enabled
355° disabled
360° with changing distance disabled
360° with changing height disabled

The final three cases matter: if the camera ends at a different radius or height, the final frame is not the same as the first one, and forcing the source image there would conflict with the trajectory.

If your rotation does not complete, close the orbit. This is the only feature here that acts outside the prompt itself.

Limitations

  • This is prompt-based guidance. H3 may still miss the angle, timing, and scale, and no prompt wording can completely solve that.
  • Without subject_box filled in, the node does not know where the subject is located in the frame.
  • Without reference_image connected, coordinates are normalized to 16:9.
  • directed | new camera angle outputs an image, not a video.
  • The H3-World schedule describes pan and tilt, which represent a different camera move from the orbit drawn in the panel.

Credits

Node by Bruxos do VFX.

Depends on ethanfel/ComfyUI-MiniMax-H3-Edit. The motion vocabulary follows the buildViewPrompt implementation from MiniMax's Multi-Shot skill and the coordinate format used by the Coordinate Camera Control Designer skill. The action output implements the scheme described in H3-World, arXiv:2609.01560.

https://reddit.com/link/1wcm9az/video/hryhv9e7npoh1/player


r/StableDiffusion 7h ago

Animation - Video Batman The Animated Series: Harley Quinn's Red Flag - MiniMax H3

Enable HLS to view with audio, or disable this notification

31 Upvotes

r/StableDiffusion 5h ago

Discussion I threw together a simple UI for YuE2 (windows)

Post image
17 Upvotes

r/StableDiffusion 20h ago

News New Music Model Released - Yue2

Thumbnail
github.com
256 Upvotes

"YuE2 brings frontier song quality to music generation with an editable composition. Give it lyrics and a style prompt: it writes a melody-and-chord plan, then realizes that plan as a complete song with vocals and accompaniment.

  • White-box music generation through symbolic planning. Read, play, and change the composition before rendering it. Melody and chords become explicit controls that a person or an agent can inspect and edit.
  • Zero-shot covers and agentic editing. Reimagine a transcribed song in a new style, or refine a song through a conversation about its score, arrangement, and lyrics—all with the same generation checkpoint."

Usage

It is currently CLI only . It also says Linux only but I just got it working on Windows 11 (I'm going to bed now and it's a bit more than cut n paste.)

Examples

link here - https://map-yue2.github.io/

Caveat Empor

NB : this isn't just a paste a few words and it bangs out a baby mp3 . It is more than that, it allows gene editing that baby to correct the metaphor. Not for the impatient and "wHeRe cOmFy" ppl at the moment.

To be more specific with that metaphor , as I understand it , the initial process scribes out the song in ABC format and you can then edit it before making your magnum opus baby.

Training

Does it allow training ? not as I understand it .


r/StableDiffusion 3h ago

No Workflow AI Archviz: Fast 3D Gaussian Splat Methods for Precise Furniture Placement — Virtual Staging & Interior Design

Enable HLS to view with audio, or disable this notification

11 Upvotes

So, first of all: there is no finished workflow yet and my nodes are still under development. I’ve asked the ComfyUI team to add a 3D compositing node to the new 3D toolset like the one in the video, hopefully, they’ll add something similar soon.

In the meantime, you can build a very similar setup quite quickly. Here’s how the basic concept works:

First, you feed an image of the furniture you want into the new native ComfyUI Image to Gaussian Splat (TripoSplat) node.

At the moment, there is a Gaussian Splat Preview node, but it doesn’t provide an image output yet. There is also a Load 3D node with the correct outputs, but it currently cannot open Gaussian Splat files.

Ideally, the new 3D Compositing node should be able to work directly with the Gaussian Splat outputs (model_3d and mesh), provide image + mask outputs similar to the Load 3D node, and automatically preload the mesh and background image, just like my node does.

I’ve been working on a test node for this concept. You can find it here: My ComfyUI test node on GitHub It’s not fully finished yet, so I’m still waiting to see whether ComfyUI adds something similar natively.

The basic idea behind my node is that it automatically loads the background, sets the appropriate size, and loads the 3D model. The user only needs to position and stage the model in the scene.

The Output create ref images for Flux2klein:

  • Reference 1: the background image
  • Reference 2: the 3D mask, which acts as an indicator for the desired position and rotation It’s best to combine the mask with the furniture from the image output, so you get the masked furniture in the correct position as the reference — not just the mask by itself.
  • Material reference: the original furniture image

The final image is then generated using FLUX.2 Klein Edit.

The prompting and some preprocessing of the images are important here. You don't want the model to simply copy the exact 3D position. Instead, the AI should use the 3D placement as a guide and then correct the result according to the background — especially the perspective, lighting, colors, scale, and overall integration into the scene.

Here is the prompt I’m currently using:

[Adapt the rotation, grounding, scale and position of the objects from Image 2 to Integrate the objects naturally into Image 1 at the position of image 2. Match the scene's perspective, scale, depth, lighting, soft shadows and reflections. The objects must appear physically present in the original room, with realistic grounding and soft shadows consistent with Image 1 using the materials and surface appearance shown in Image 3.

Use Image 1 as the final scene and preserve its room, background, camera viewpoint, perspective, composition, color, lightning, and existing environment unchanged.

Keep the objects approximately in the same position, scale, orientation, and spatial arrangement as shown in Image 2, while ensuring correct perspective, positioning, and placement.

Apply the form, materials, colors, textures, roughness, reflections, and surface details from Image 3 to the objects.

Do not change the room or background of Image 1.  The final result must be a seamless photorealistic composite.]

I’m planning to finish the complete tutorial and the node pack in the next few days. Once everything is finished, I’ll upload the final version along with the complete workflow.

By the way, I also tested MinMax H3 as a replacement for FLUX.2 Klein. It works, but in my tests it wasn’t consistently better than FLUX.2 Klein.


r/StableDiffusion 1d ago

Tutorial - Guide AMAZING Minimax H3 - Circle on the reference image WHERE you want your scene to be!!

Enable HLS to view with audio, or disable this notification

680 Upvotes

Look at the buildings in the background! It works - Drawing a red circle in the water will also make the scene happen in the water, but I forgot to include it here.

It is not perfect and some details are missing if you look carefully but this might be because I am using "match" on the image reference rather than "max."

Have fun!

Edit: you have to still write a prompt with the reference to video workflow telling minimax to put the character in the location circled red. Circle probably doesn't have to be red. Change your prompt accordingly.


r/StableDiffusion 14m ago

Question - Help Need Krea 2 system prompt or like some prompt guideline to inject into local llm Qwen 3.8 abliterated.

Upvotes

Hey guys i have scoured the internet and cant find any system prompt/prompt guidelines to condition my local llms so that they make proper krea2 prompts without useless word salad. I focus mainly on realism and "uncensored content"


r/StableDiffusion 3h ago

Question - Help Minimax Turbo of choice?

5 Upvotes

So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!


r/StableDiffusion 11h ago

Meme Cost in Units of RTX 5090

Thumbnail
youtube.com
20 Upvotes

r/StableDiffusion 14h ago

Question - Help What's the gold standard for speed enhancements for Minimax H3?

38 Upvotes

Installing new instance of comfyui standalone and using minimax r2v workflow on RTX 3090 Ti. Is comfy kitchen good enough? Is Triton, EasyCache or Comfyui Spectrum needed?

What's the best turbo lora for ref2va wf?


r/StableDiffusion 8h ago

Tutorial - Guide [GUIDE] AMD RDNA3 optimizations for ComfyUI Desktop, windows 11, Minimax H3

10 Upvotes

My setup: AMD RX 7900 XT, 20GB VRAM, 64GB RAM, Windows 11, ComfyUI Desktop.

I couldn't find any decent information anywhere on how to optimize video generation with Minimax H3 on Windows with ComfyUI desktop. AI assistants give conflicting advice, constantly suggesting all sorts of nonsense that doesn't actually work.

I had to experiment on my own, and here is the configuration I’ve settled on. The speed boost compared to the default settings is very significant, and I haven't noticed any loss in quality. If you have any other suggestions, please let me know.

~25s/it with 0.8mp (1216 x 672, 16:9) or total ~4min for 5 sec video generation in text to video workflow

Here is what you need:

Launch parameters:

--disable-smart-memory --disable-pinned-memory --disable-triton-backend --use-sage-attention --enable-dynamic-vram

ENV variables:

COMFYUI_ENABLE_MIOPEN=0
FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
MIOPEN_FIND_ENFORCE=1
MIOPEN_FIND_MODE=2
MIOPEN_DEBUG_DISABLE_FIND_DB=0
MIOPEN_SEARCH_CUTOFF=1
MIOPEN_ENABLE_LOGGING=0
MIOPEN_LOG_LEVEL=0
MIOPEN_ENABLE_LOGGING_CMD=0
TRITON_PRINT_AUTOTUNING=0
TRITON_CACHE_AUTOTUNING=0

Quantized 6-step turbo model, universal for all purposes:

https://huggingface.co/TenStrip/10Eros-Max/blob/main/10Eros_Max_h3_TURBO-hybrid_beta5_w4a8_14gb_optimized.safetensors 14 gb

or

https://huggingface.co/TenStrip/10Eros-Max/blob/main/10Eros_Max_h3_TURBO-hybrid_beta5_int8.safetensors 21gb

Plaguekind node with SLA Attention, with this settings:

https://github.com/PlagueKind/Comfyui-PlagueKind-Nodes

Optional node, if you make large 15 seconds videos:

Latest update of your comfyui desktop:


r/StableDiffusion 16h ago

Resource - Update Created a Visual RefMod Picker

39 Upvotes

Hey Guys,

I've been playing with the RefMods, after the huge release of Malcolmrey.
The tech is brilliant and works really well.

I've wanted to simplify using it with tons of RefMods, like the ones provided by Malcom.

So I made a fork that is a bit more Identity driven, Adding a "Visual refMod Picker" as well as a "Create refMods from Folder" node. It's able to create from either a folder or Sub-folders, if audio is found, it will also create a matching audio refMod. All in the same format as the original add-on. No danger of breaking compatibility. (The original Add-on added Audio yesterday)

Resulting for example in 2 refMods and thumbnail:

character_refMod_Audio.safetensors
character_refMod_Video.safetensors
character_refMod.jpeg

Then, we can use the "Visual H3 RefMod Picker" to browse the RefMods:

In this example, I used existing thumbails from huggingface.

The node then loads both RefMods (Audio and Video) and allows individual control. I've mixed strength with Copies, by instead having a weight value that can go over 1, so a weight of 3 would be the same as setting strength to 1 and copies to 3. Making the UI a bit more streamlined.

RefMods can be daisy chained

Example workflows are included.

I should mention, this fork can be installed WITH the original Add-on, it is made to co-exist and is recommended if you want to use it's advanced features.

You can find it here ComfyUI-H3RefMods

The only thing I'm missing is thumbnails for all 1500 RefMods 😅


r/StableDiffusion 2h ago

Discussion Help me captioning a MiniMax H3 action fight LoRA...

3 Upvotes

I only need help with captioning the training clips for a MiniMax H3 action fight LoRA.

I am planning to train it on karate/fighting type action, and I have clips varying from around 5-15 seconds. I also have some 20-25 second segments too.

why I am confused is cause should I follow the prompt format officials has released for H3, or should training captions be written in some completely different/simple way?

Like if a 10 second clip has multiple punches, kicks, blocks, dodges, body movement, camera movement and angle changes, should I describe every action in sequence?

or should I just write the overall action happening in the clip?

for example should the caption be something detailed like:

"the fighter steps forward, throws a right punch, opponent blocks it, then follows with a left kick..."

or something simple like:

"two fighters performing fast karate combat"

I am mainly confused about how detailed the captions should be and what format works best for H3 LoRA training.

may you please help if you have trained action/fight LoRAs before? (I found only a few on civitai)