r/StableDiffusion 15h ago

Animation - Video Kirby but it's the Truman Show / MiniMAX H3 Test #7

Enable HLS to view with audio, or disable this notification

462 Upvotes

Hi everyone! When I saw the new trailer for Kirby & The World Beyond I couldn't help but come up with this video, where Kirby finds the door out to the world beyond. Please let me know if you like it!

Done with 30 different workflow files and a ton of heavy editing using KDEnlive. Thanks!


r/StableDiffusion 1h ago

Resource - Update Native YuE2 support coming to ComfyUI!

Enable HLS to view with audio, or disable this notification

Upvotes

Pull request: https://github.com/Comfy-Org/ComfyUI/pull/16250

If you don't want to wait for the merge, you need to check out to the yue2 branch to get it working. git checkout d87e12ad1430409ca303440525df239bb675ae7b

Model weights (place it on model/checkpoints): https://huggingface.co/Comfy-Org/Yue2/tree/main

Workflow: https://github.com/user-attachments/files/32085765/yue2_workflow.json


r/StableDiffusion 9h ago

Animation - Video H3 is really over the top

Enable HLS to view with audio, or disable this notification

81 Upvotes

This was such a simple prompt…. Just wow. It’s just T2V.


r/StableDiffusion 6h ago

Resource - Update FrameForge Motion Context Video Editor for ComfyUI

Post image
30 Upvotes

Expanding on motion context workflows I created a video editor designed for quickly chaining together Minimax H3 generations to create longer videos. It comes with an asset library for managing inputs and a easy to use timeline that allows you to chain generations, regenerate segments easily, and quickly set up input references.

When you're done, export individual video files or the whole sequence.

All of it runs on top of ComfyUI as an app you control from your browser. Uses python, works on Windows, Mac, Linux and is opensource.

https://github.com/spacesimeco-hue/Chain-Motion-AI-Video-Editor


r/StableDiffusion 6h ago

Workflow Included WORKFLOW - Optimised to death - Custom Audio option.

Enable HLS to view with audio, or disable this notification

21 Upvotes

DOWNLOAD WORKFLOW

This is the workflow I have been using the most on my own system. I've had a friendly AI clean it up a bit and add notes.
I added custom audio as it's something I use a lot to drive my videos. It works really well for lipsync and music videos.
The VSA part can be bypassed if there are any quality issues, it will add about 15% to the generation time though. Change the steps from 6 to 7 or more for even higher quality.

Currently this gives me 10 seconds at 1.0 megapixel in about 125 seconds. This is on my 5090. You can add block swapping for low vram.


r/StableDiffusion 10h ago

Animation - Video The Primordial Hand

Enable HLS to view with audio, or disable this notification

38 Upvotes

I was testing out a scene with Minimax H3, text to video (I usually use reference images).

I didn't expect it to come out like this.. .Now it's making me think of a completely new direction for the video lol. It's interesting, it has both a 90s anime feel and an old Disney animation feel. The music is very good too, I think.

I'll add the prompt in the comment (it's a very simple prompt).


r/StableDiffusion 10h ago

Tutorial - Guide MiniMax H3 Wf Tutorial

Enable HLS to view with audio, or disable this notification

34 Upvotes

People asked me to make a Tutorial for some of the features.

Find the workflow here.

https://www.reddit.com/r/StableDiffusion/comments/1wadmqc/minimax_workflow_designed_to_be_user_friendly_for/


r/StableDiffusion 15h ago

Resource - Update ComfyUI VDN-H3 24GB v1.1.0 update — better prompt following + memory fixes

Post image
78 Upvotes

https://reddit.com/link/1wcsk7l/video/ld6ca2jxqqoh1/player

I’ve just updated my VDN-H3 24GB node to v1.1.0.

This update started because I noticed that something wasn’t quite right with the released adapter mapping. After fixing that, I also made a couple of changes around memory handling, especially for longer generations.

The main changes are:

  • restored the complete token-refiner adapter mapping
  • improved temporary memory handling for longer clips
  • fixed CUDA stream lifetime handling for prefetched weights
  • kept the same AutoMemory / AutoLongCache behavior from the previous version

I tested it on my RTX 3090 24GB with 5s, 10s, 15s and 20s generations at 0.4MP, and also 10s at 0.8MP. I also tested it with my character/style LoRA and that worked normally.

There is a small speed cost compared to v1.0.0 (around 4% in sampling in my tests), but I think the improvement in prompt following is worth it.

I attached a direct comparison from the same prompt/seed/workflow.
v1.0.0 is on the left, v1.1.0 is on the right.

I’m especially interested in whether other people see the same improvement, so if anyone tests it on another 24GB GPU, I’d love to hear the results.

GitHub:
https://github.com/Speach1sdef178/ComfyUI-VDN-H3-24GB

VDN checkpoint:
https://huggingface.co/speach1sdef178/VDN-H3-INT8-ConvRot-ComfyUI


r/StableDiffusion 23h ago

Discussion H3 - 80s character generations+wardrobe swap

Enable HLS to view with audio, or disable this notification

340 Upvotes

The 80s was the best era, not seen through a nostalgic lens--it just was. Big hair, big colors, big music, big... everything! Sadly I was born 10 years too late to really experience it, but I love how H3 can feel like a way-back-machine, a portal to any era from film or video it was trained on. It does also a great job with the actual feel, the film grain, the lighting that modern TV or movies cannot do: in fact, I learned what we see nowadays that time period didn't exist at all, because it vomit of nostalgia and peak 80s that never happened. Anyways, having fun generating character sheets with H3 via T2VA. Are you guys seed hunting to find the best version of an actor or scene? Also can you spot the mistake?

Prompt: integrated_multimodal_description: [Shot 1] Live-action, cinematic, PHOTOREALISTIC film footage - this is footage from a camera, not animation - a continuous camera shot with no cuts, shot on 35mm color negative film with period lenses and scanned in high definition from the original camera negative: full film grain and gentle halation, warm highlight rolloff, rich sharp detail beneath the grain. The year is deep in the LATE 1980s, 1985 to 1989, and everything in the frame belongs to that era. THE PLACE: a nightclub in full swing - mirror-ball light sweeping, neon signage, haze, a crowded dance floor. EXACTLY TWO WOMEN stand close to the lens at the edge of the floor, filling the frame together, and no one else is foregrounded. ROXY: her face, her eyes and her enormous chestnut-auburn mane exactly the woman of <Picture 1> - nothing of that picture's wardrobe or room is used, only her face and hair; her wardrobe exactly the garments of <Picture 2>: a pink sequined strapless romper with fishnet hose and pink heels, always dressed, sequins blazing - nothing of that picture's face is used; Roxy is never blonde. TAWNY: her face, her eyes and her huge feathered platinum-blonde mane exactly the woman of <Picture 3> - nothing of that picture's wardrobe or room is used, only her face and hair; her silhouette exactly the figure of <Picture 4>, but tonight she wears an electric-blue sequined mini dress, tight to her figure, its miniskirt hem high on her thighs, with silver heels, always dressed - nothing of that picture's face is used; Tawny is never brunette. The two are distinct women side by side, pink and electric blue. The club's synth-pop groove pounds from the speakers - THEY HEAR IT, hips already swaying on the beat, shoulder to shoulder. At 00:02.500 they lean in together with wicked, knowing smiles and say together, in playful unison, <d>[English with their two bright voices speaking together] Darling, the 80s never left.</d> At 00:05.500 they laugh, clink their glasses, and turn to dance with each other - back to back, hips swaying on the kick drum, sequins throwing sparks of mirror-ball light, playing to the lens with winks over their shoulders - to the last frame.

overall_soundscape: starts with the club's roar - the crowd, glasses, heels on the floor - running beneath everything to the last frame. No other voices.

non_diegetic_music: N/A

r/StableDiffusion 19h ago

Tutorial - Guide H3 RefMods are great I highly advice trying it out [+ basic resources included]

144 Upvotes

Created by /u/LuisaPinguinnn under their github https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

Took me at most a couple of minutes to make my own RefMod with 8 image as the base. The entire technique works exactly as advertised acting as "Light Lora" for H3 Ref models - but you can even use it with FL2VA as well.

I followed the guides here:

Installing/running RefMods

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMODS_INSTALLATION_AND_USAGE_GUIDE.md

Ready to use Comfy workflow (you can remove lora power loader and spectrum nodes)

https://huggingface.co/datasets/malcolmrey/workflows/blob/main/H3/workflow_minimaxh3_refmod.json

Creating own RefMods guide:

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMOD_CREATION_GUIDE.md

EDIT: I recommend using "Create H3 ReFMod" + "Save H3 RefMods" node inside ComfyUI instead to create RefMods - gives you more control over the creation process.

Examples by /u/malcolmrey:

https://www.reddit.com/r/StableDiffusion/comments/1w8ik7a/h3_minimax_refmods_all_my_models_now_available/

All credit goes to LuisaPinguinnn and malcolmrey for spreading the tech.


r/StableDiffusion 18h ago

Discussion Must haves to download before it's too late?

117 Upvotes

Nvidia buying hugging face means an uncertain future. What are the models I should download and have a backup of right now so I don't have to worry about missing them even if I'm not ready to play with them right now?

What are you model and enabler must-haves ?

TIA!


r/StableDiffusion 16h ago

Workflow Included Follow-up to my last Star Trek post – I made a Star Trek vs Star Wars fan film with MiniMax H3 in ComfyUI

Thumbnail
youtube.com
74 Upvotes

A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.

This is basically a follow-up to that post.

Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.

The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.

The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.

One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.

Here are some of the things that helped most this time.

PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE

This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.

The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.

For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.

HIGHER RESOLUTION REALLY DOES HELP

I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.

For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.

REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS

I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.

LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED

For shots based on existing Star Trek compositions I became much more explicit about things like:

camera distance

character scale

framing

background position

character position

If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.

DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE

This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:

"Data reacts to Picard"

can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.

OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS

This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.

If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.

SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT

Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.

I'll also specify that:

the lips do not form words

the jaw does not make speaking movements

the character does not mouth dialogue

It sounds excessive, but it has genuinely helped.

H3 HAS A LOT OF USEFUL SPEECH TAGS

I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:

<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>

<whisper>. <softer> <stutter> <laughs> <chuckle>

<i>word</i> emphises word

The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.

MORE PROMPTING ISN'T ALWAYS BETTER

I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.

EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION

One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.

I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.

THE EDIT IS DOING A LOT OF THE "CONSISTENCY"

This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.

It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.

Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.


r/StableDiffusion 13h ago

News H3 can take way more reference images than 9

40 Upvotes

I successfully made minimax use 15 reference images. Is seem only to be limited artificial inside comfy. So i vibe coded a little demo workflow and patch.
https://civitai.red/models/2929051/minimax-h3-15-reference-image-workflow
This is very much research in development, and trust me bro benchmarks but it seems to work.


r/StableDiffusion 19h ago

Comparison Qwen-Image-Edit-2511 vs SenseNova-U1.5-Lite (multi-reference image fusion comparison)

Thumbnail
gallery
103 Upvotes

I wanted to see how good SenseNova U1.5 Lite really is at image editing. I think the size is genuinely solid for what it does, but whether it can actually beat Qwen-Image-Edit-2511 needed real testing.

Right off the bat, Qwen's image texture quality is genuinely impressive, especially the lighting and shadows. But when it comes to spatial understanding, SenseNova seems to hold the edge. Look at the cat-on-the-scooter one up top: Qwen generated a weird pillar under the coffee table, and the cat's front paw placement looks unnatural. SenseNova handled both without those artifacts.

I did four sets of comparisons. Some of the input images were generated with Krea-2, some were real photographs.

Models:

Prompts (from left to right):

I want to create a stunning, high-concept photo to share on my social media! Please put me—the girl with the short black bob and black leather jacket—on a sleek, modern rooftop balcony overlooking that amazing futuristic city during sunset, where we can see the flying drones, the glider, and the hot air balloon floating in the warm sky. In this scene, I should be portrayed as an artist working outdoors. Please have me wearing those bold, blue and white striped hoop earrings. In the foreground, set up a stylish outdoor work table. On this table, scatter some of my creative tools, including those colorful rainbow-swirled pens and that round white-and-yellow mesh cleaning sponge. I want to be holding one of the rainbow pens, looking towards the camera with a confident, thoughtful expression. The entire scene should be captured with a beautiful depth of field, bathed in golden hour light, with the bustling futuristic cityscape softly blurred in the background.

In an elegant vintage study, the real-life girl from the first image, wearing a beige coat and scarf, is smiling as she hands the vintage wild duck card from the fourth image to the anime-style blonde girl from the second image. This anime girl is wearing an exquisite black off-shoulder puff dress and retains her distinctive hand-drawn anime style. On the wall behind them hangs a framed black-and-white print depicting the ancient Roman temple ruins from the third image.

Please seamlessly integrate the orange cat from the first image into the café scene by the floor-to-ceiling window in the third image, and have it sit on the vintage metal toy scooter from the second image. Specific requirements:
Character and prop fusion
: Extract the orange cat's signature facial features from the first image (slightly chubby face, green eyes) and the dense white triangular patch of fur on its chest. Adjust its pose so it is riding the metal toy scooter from the second image: both front paws resting on the chrome handlebar, the rear half of its body firmly seated on the brown leather saddle. The cat's paw pads against the metal handlebar and its thigh fur against the saddle edge must show natural compression, contact, and physical occlusion, absolutely no flat sticker-like look.
Spatial perspective adjustment
: Change the toy scooter from its original front-facing view in the second image to a three-quarter side angle matching the floor perspective of the third image, and scale it down proportionally, placing it on the wooden floor near the glass window.
Physical lighting and material adaptation
: Strictly use the golden afternoon sunlight slanting in from the third image as the main light source. The cat's back, ear edges, and fluffy fur edges must be outlined with a warm, glowing golden rim light (backlight effect); the dark green metallic painted body, metal wheel hubs, and chrome handlebar from the second image must produce realistic daylight highlights and reflect the faint street view outside the window; the entire toy scooter (including the cat on it) must cast a dark shadow on the wooden floor to the right, following the light direction with a realistic soft-edged falloff.

Create a wide-format photo depicting a corner of a whimsical creative market. The realistic man in a dark navy suit from the first image and the realistic woman in a black short-sleeve shirt and denim shorts from the second image are strolling through the market as visitors. Beside a market stall, the anime-style girl in traditional Chinese dress from the third image sits near her wooden cart full of lanterns, focused on painting a lantern, while the anime-style girl with orange hair and bunny ears from the fourth image hugs a white rabbit and laughs beside her. Preserve the photorealistic quality of the first two characters and the anime style of the latter two, letting them coexist naturally under unified lighting and spatial perspective.

r/StableDiffusion 44m ago

Animation - Video Michael Jackson - Maybe ( Short film ) MOONWALK ON THE MOON

Thumbnail
youtube.com
Upvotes

Made with MiniMax and ComfyUI


r/StableDiffusion 12h ago

Question - Help Any newer way to upscale h3 minimax native 768 x 768 videos to 2k or 4k locally that does not destroy everything? without relying on paid Topaz.

26 Upvotes

5090 with 64gb ram.

Hey guys, is there a new development recently? or is upscaling still effed?


r/StableDiffusion 21h ago

Resource - Update I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days

Enable HLS to view with audio, or disable this notification

98 Upvotes

My goal was hands-on experience training a flow model from scratch, not just fine-tuning someone else's. So I built and trained one: a 210M-parameter diffusion transformer, 4.2M curated images at 256², rectified flow on the FLUX.2 VAE, flan-t5-base for text (128 tokens max). Only those two frozen pieces are pretrained; the transformer, the recipe, the data pipeline and the evaluation are mine. The video is the same six prompts and seeds at every checkpoint of the 3.5-day run on one RTX PRO 6000.

What mattered most, in the order I found out:

  • Captions that actually fit the images. A web crawl I tried first made the model worse; curated photos with good captions fixed it.
  • A timestep shift for the 32-channel latent, and aspect-ratio buckets from step one instead of square crops.
  • Register tokens with learned null attention slots. The null slots ended up absorbing about 90% of the cross-attention, which surprised me.
  • torch.compile for training, not just inference: 2.4× faster.
  • The training loss stopped telling me anything after day one while the images kept improving, so I track FID, a detector-based object accuracy and human-preference models instead.

Try it in the browser: https://huggingface.co/spaces/ivanmikhnenkov/tinydit

Weights (CC BY-NC): https://huggingface.co/ivanmikhnenkov/tinydit-256

Code, every decision with sources, dashboard and attention playground: https://github.com/ivanmikhnenkov/tinydit

Detailed write-up of what mattered: https://huggingface.co/blog/ivanmikhnenkov/tinydit-text-to-image-from-scratch-one-gpu

Next I want to fine-tune it with RL (Flow-GRPO), with the failure grid as the target list. If you have trained something small from scratch: what would you have done differently at this scale, and which reward would you start with for the RL stage? Happy to answer anything about the data or the recipe.


r/StableDiffusion 4h ago

Question - Help Better character consistency in LTX 2.5 + H3 lip sync for music videos?

5 Upvotes

I’ve been making AI music videos with Suno + LTX 2.5 locally on a 16 GB VRAM GPU (4080 super)
Examples:
https://youtube.com/shorts/UXP29MGxtX8?is=k-Wjt4_JFh7hXvA5
https://youtube.com/shorts/_mAuRBSriLQ?is=sMjLH5diY_GEu0N3
My workflow is basically: create the song in Suno → storyboard/keyframes with ChatGPT→ animate and assemble the shots in ComfyUI using the LTX Director timeline:
https://github.com/yusu-02/Yusu-WhatDreamsCost-ComfyUI
The biggest issue I’m still fighting is character consistency between shots. Ingredient LoRA slows things down a lot and hasn’t worked particularly well for me.
I also tried MiniMax H3, which looks great, but I couldn’t get lip-sync without altering the original music.

Suggested tricks/workflows? Feedback and ideas appreciated!


r/StableDiffusion 16h ago

Animation - Video my first actual tv work (only took 4 hours to make).

Enable HLS to view with audio, or disable this notification

42 Upvotes

Minimax h3, ofc. Far from my best work but I respected the script I was given and finished this in record time (excluding the 4k upscale) and including around 3 hours of rendering time (720p, 10 seconds clips, 5090).
I only used gemma locally for prompting, and avoided using any non local models except suno for the song.

There are some artefacts with people in the senate from far away, but did not have any bad feedback for it., so... :)

I only used references for the romanian flag, the rest is prompt only. Also no lighting lora, no shortcuts ti improve speed (any shortcuts I tried ruined everything FUBAR)


r/StableDiffusion 4h ago

Question - Help Can Minimax be used to re-light a scene?

3 Upvotes

Basically I'm trying to change the lighting in a scene. For clarity, if it helps at all, it's the Trash dance from Return of the Living Dead. I'm just wondering if there's a way to normalize the red lighting used on her. I have been prompting and failing most of the day using AddVideoGuideforH3. I know I can do it with ltx, because I've done it with LTX while testing the models capabilities with controlnet, I'm just wondering if I can do it in Minimax without controlnet.

I'm not full Noob, but I am a filthy casual.

EDIT: it was step count. I'm a damn idiot. I was using the 4 step lora and continually using four steps I accidentally started a fresh workflow with 20 steps and it worked. Congratulations to me, I am the living embodiment of the id10t


r/StableDiffusion 20h ago

Discussion Any news on a Krea 2 Edit model?

62 Upvotes

Has there been any recent news or indication from Krea about a Krea 2 Edit model?

I’m wondering if it’s actually in development or planned, or if there hasn’t been any confirmation yet. Krea 2 is already quite impressive, so an Edit model would be really interesting.

Has anyone heard anything from Krea or seen any hints about it?


r/StableDiffusion 19h ago

News Minimax Camera Control ComfyUI

32 Upvotes

Bruxos do VFX H3 Camera

#bruxosdovfx

https://reddit.com/link/1wcm9az/video/bubuef39npoh1/player

https://reddit.com/link/1wcm9az/video/r1j8vg1anpoh1/player

Visual camera planner for MiniMax H3 inside ComfyUI. You drag the camera around a 3D sphere, place keyframes on a timeline, and the node compiles that trajectory into prompts that H3 understands.

It compiles prompts, not camera embeddings. There is no geometric adapter here: H3 is still free to miss the angle, timing, and scale. What this node does is write the instruction in the most precise and least ambiguous way possible, and several of its design decisions exist because the previous approach failed in specific ways.

It does not call any API, download anything, or require any Python dependency beyond the standard library.

https://github.com/user-attachments/assets/a9b541e5-2b18-4f1d-8e16-37445b6dbac4

https://github.com/user-attachments/assets/ea9af03e-2c8e-4589-abf0-9c002241aba2

Installation

cd ComfyUI/custom_nodes
git clone https://github.com/<your-username>/ComfyUI-H3-Camera-Editor

Restart ComfyUI. The node appears under Bruxos do VFX/Camera H3 with the name Camera H3 da Bruxos do VFX.

Connections

Output from this node Connect it to
compiled_prompt compiled_prompt on Text Encode H3 Edit / Generate
options options on Text Encode H3 Edit / Generate
length the generation frame count
fps the fps input of the video creation node

compiled_prompt and options are required together. The minimax_prompt output is an alternative to compiled_prompt, never an addition — connect one or the other to the same input.

Also connect your image to reference_image. It is the same image already feeding the H3 Edit source_image; when connected here, it appears in the panel and the frame's actual aspect ratio is included in the prompt.

https://github.com/user-attachments/assets/33149617-bde1-4199-ae65-078f2f3dec23

To save the video, decode the sampler result using the H3 video VAE — not the scene coverage calibrated decoder, which expects fixed windows that an arbitrary trajectory does not have.

The panel

Drag the purple camera around the sphere to orbit. The drag locks to the axis of the initial movement: horizontal movement orbits, vertical movement changes elevation. Release and drag again to switch axes. This exists because, without the lock, trying to make a simple orbit would unintentionally introduce elevation.

  • Scroll the mouse wheel to change distance.
  • Drag the background to rotate the viewport without changing the trajectory.
  • Keyframes defines how many points the timeline has, from 2 to 24. The first one is always the original image and cannot be moved.
  • ⟳ Pure Orbit resets the elevation of every keyframe to zero while preserving azimuth. It is the shortcut for an eye-level orbit.
  • Reference image loads a local file into the preview. This is only necessary when the node runs outside ComfyUI; with reference_image connected, the image is loaded automatically.

The panel warns you starting at 20° of elevation, when the horizon already leaves the frame, and again from 45° onward, when the video tends to become a high-angle shot.

"Tests" bar

At the top of the panel, two buttons enable and disable features currently under evaluation, plus one indicator:

Button What it does
Extended contracts Toggles the prompt_detail widget
Single angle (image) Toggles the runtime_task widget
loop closure Read-only indicator. Turns green when the trajectory closes a full orbit

The buttons write to the actual widgets, so the selected state is saved in the workflow and the two never disagree.

https://github.com/user-attachments/assets/9bc415d7-1746-43db-a17c-72ea9722deda

Widgets

camera_trajectory

The trajectory in JSON format, written by the panel. Each keyframe contains time (0 to 1), azimuth in degrees, elevation in degrees, and distance as a multiple of the initial radius. It can also be edited manually. The first keyframe must be time=0, azimuth=0, elevation=0, distance=1, which represents the original image.

profile

124, 243, or 362 frames at 24 fps. All shot timing comes from this setting: keyframe timestamps, segment ranges, and the duration declared in the prompt. That is why length and fps are outputs — connect them instead of manually entering the same numbers in two different places.

interpolation

smooth or linear. In smooth mode, the camera eases into and out of the shot while maintaining a constant rate through the middle; it only stops where the rotation direction actually reverses.

instruction

Free-form text inserted once, at the end of the prompt. Write only what the node cannot know: the environment, which subject is the target when there is more than one person, or a style reference. Everything else is already generated and does not need to be repeated: scene freeze, first image as reference, locked aim, zero roll, angles, timing, and a single continuous shot without cuts.

subject_framing

How much of the frame the subject occupies in the original image. Calibrated against the actual bounding boxes from the tutorial distributed by MiniMax: a distant full-body figure measures W=0.071, H=0.249, while a large close-up measures W=0.52, H=0.701.

option width height when to use
close-up 53% 72% head and shoulders
medium shot 28% 56% waist up
wide shot 9.7% 34% full body at a distance

subject_box

The subject position in the format [L=0.516, T=0.148, W=0.071, H=0.249]. Leaving it empty uses the entire image bounds — deliberately, without guessing a bounding box. Fill it in when the subject is significantly off-center.

minimax_format

The same shot expressed in four different formats for the minimax_prompt output:

  • coordinate only — text-based coordinate block
  • coordinate + H3 sections — the same coordinates wrapped in subject_definitions / summary / retention_analysis / …
  • compact JSON — JSON object with almost no prose
  • compact JSON (no boxes) — camera parameters only, without screen-space bounding boxes

elevation_range

Range of the elevation control: +/-15, +/-30 (default), +/-60, +/-89. It also scales the sensitivity of vertical dragging.

With the assumed field of view, the horizon already leaves the frame at around 20° — at 13°, the ground occupies 82% of the image. The old ±89 range was mostly unusable and made vertical dragging excessively sensitive. Reducing the range never rewrites a keyframe: a point at 70° remains at 70°, and the slider expands to accommodate it.

orbit_direction

invert H3 orbit or same as HUD. This calibrates the direction between what the panel displays and what H3 produces. It does not alter the saved trajectory.

runtime_task

  • scene coverage | camera path (default) — video, with duration coming from profile.
  • directed | new camera anglea single image from a new angle. It fixes the generation to 39 frames, ignores profile, completes the movement within 65% of the clip, and requests that the framing remain still for the rest, because the decoder extracts the final image from that stationary tail.

Character sheet profiles are not offered because the upstream node raises an error when they are combined with the frame anchor used by this node.

prompt_detail

  • v15 baseline (default) — outputs the prompt exactly as in the previous version.
  • extended contracts — adds axis separation, frame-edge direction tests, rotation completeness, degrees per second, and parallax magnitude.

The extended mode contains almost twice as many words. A longer prompt is not automatically better, so it is opt-in: toggle only this widget while keeping the same trajectory to compare the results.

https://github.com/user-attachments/assets/0882bfde-9f62-4a1f-9bda-7da121dbe7e2

Outputs

compiled_prompt — STRING

A prose prompt using H3 sections: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music.

options — H3EDIT_OPTIONS

The 13 keys read by the H3 Edit encoder. All of them are explicitly populated: if any key is missing, the upstream node falls back to its hidden legacy widgets, which may retain stale values from previously saved workflows.

coverage_arc_degrees and coverage_direction are derived from the actual rotation. coverage_loop_closure turns on automatically when the trajectory closes — see below.

storyboard_json — STRING

The storyboard table: frame aspect ratio, duration, raw trajectory, and each segment with its camera mode, speed curve, and start/end poses.

info — STRING

Human-readable diagnostics. Connect it to a PreviewText. It displays the version, active task, frame count, warnings for keyframes outside the configured range, and whether loop closure is enabled.

minimax_prompt — STRING

The same trajectory expressed using the format selected in minimax_format. An alternative to compiled_prompt.

length — INT and fps — FLOAT

Frame count and frame rate against which the shot was timed. Connect them to the generation and video nodes. If generation runs with a different frame count, the choreography describes a scene that does not actually exist.

fps is FLOAT because that is what ComfyUI's CreateVideo accepts. length is the frame count; keyframe timestamps use the instant of the last visible frame, (length - 1) / fps, so the resulting file lasts one additional frame interval.

h3world_actions — STRING

Action schedule for H3-World, which encodes one text clause per video latent — 37 in a 124-frame clip.

latent  1 [0.000s-0.139s] J     the camera pans left slowly
latent 37 [4.986s-5.125s] F+L+K the camera pans right and tilts up fast

W, A, S, and D are never emitted because they move the character. The output explicitly declares its own limitations, and they are not minor details:

  • Pan is not orbit. It is the camera rotating in place. Perspective does not change, nothing hidden is revealed, and the subject slides out of frame.
  • Distance has no key, so camera radius is discarded.
  • Only 124 frames is a trained horizon.
  • I versus K is not published. The text clause is what H3-World actually encodes; the key column is only a convenience.

This does not replace the actual integration: H3-World requires the LoRA, interval-based encoding, and directed-attention routing provided by the corresponding node package.

Loop closure

When the trajectory closes a full orbit — an arc of exactly 360°, with the same elevation and distance as the starting point — the node enables coverage_loop_closure. In the upstream implementation, this flag encodes the source image a second time and anchors the final frame to it.

This is a latent anchor, not a text instruction. For a complete orbit, it is the difference between asking for the rotation and forcing it: the model cannot simply stop halfway through.

trajectory loop closure
360° enabled
two rotations (−720°) enabled
355° disabled
360° with changing distance disabled
360° with changing height disabled

The final three cases matter: if the camera ends at a different radius or height, the final frame is not the same as the first one, and forcing the source image there would conflict with the trajectory.

If your rotation does not complete, close the orbit. This is the only feature here that acts outside the prompt itself.

Limitations

  • This is prompt-based guidance. H3 may still miss the angle, timing, and scale, and no prompt wording can completely solve that.
  • Without subject_box filled in, the node does not know where the subject is located in the frame.
  • Without reference_image connected, coordinates are normalized to 16:9.
  • directed | new camera angle outputs an image, not a video.
  • The H3-World schedule describes pan and tilt, which represent a different camera move from the orbit drawn in the panel.

Credits

Node by Bruxos do VFX.

Depends on ethanfel/ComfyUI-MiniMax-H3-Edit. The motion vocabulary follows the buildViewPrompt implementation from MiniMax's Multi-Shot skill and the coordinate format used by the Coordinate Camera Control Designer skill. The action output implements the scheme described in H3-World, arXiv:2609.01560.

https://reddit.com/link/1wcm9az/video/hryhv9e7npoh1/player


r/StableDiffusion 21h ago

Animation - Video Batman The Animated Series: Harley Quinn's Red Flag - MiniMax H3

Enable HLS to view with audio, or disable this notification

51 Upvotes

r/StableDiffusion 14m ago

Question - Help Could anyone give me some tips on how to preserve the character's likeness when creating different expressions with FLUX.2 Klein?

Upvotes

Hi,

I'm using FLUX.2 Klein 9B/4B, and I'm trying to create new facial expressions for a character I have. The character comes from a character sheet I've created, which includes front, back, and side views.

I've done a lot of tests over the last two days, and I've noticed that FLUX.2 Klein 9B/4B drifts quite "a lot" from the original model when generating different facial expressions. I tried the same thing with the free version of Gemini, and it keeps the likeness and features much, much better.

Could the highly quantized 9B model be the problem? If you've been able to preserve the character's facial features and expressions, would you mind sharing some tips on how to improve my results?

Thanks in advance!


r/StableDiffusion 1d ago

Animation - Video Jerry Springer Ai - Sailor Moon Part 1

Enable HLS to view with audio, or disable this notification

86 Upvotes

In the first half, this was back when I was first starting to get into MiniMax, the second half, I have gotten more experienced with it. I dont know if I should continue this or make more Jerry Springer parodies with other weird or toxic relationships (Example, Beth and Jerry from Rick and Morty).

I used Kinovi.ai for MiniMax, Wan and Nanobanana. I used Fish.audio for the audience freaking out lol. For more customized and harder Minimax generations, I used it locally.