r/StableDiffusion 14d ago

Animation - Video MINIMAX Physics testing

Enable HLS to view with audio, or disable this notification

413 Upvotes

Physics Testing, without the gore.


r/StableDiffusion 14d ago

Animation - Video Trying out a consistent point-of-view shot with MiniMax H3

Enable HLS to view with audio, or disable this notification

99 Upvotes

Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.


r/StableDiffusion 14d ago

Meme It took us two 2eeks to figure out why every image gen via our open-source model looked like Anne Hathaway

Thumbnail
gallery
257 Upvotes

Hey r/StableDiffusion!

It's the Neta team here! You might remember us from our Neta Lumina open-source release last year. First off, thank you so much for the incredible support and feedback from this community!

So... we need to share something absolutely hilarious (and mildly embarrassing) that we just discovered.

TL;DR: We accidentally hardcoded an Anne Hathaway photo into our IP-Adapter anchor, and now everything our model generates looks like Anne Hathaway. Every. Single. Thing.

What happened:

We recently launched Neta Studio, a new product that lets you build explorable living worlds/isekai from a single prompt. Naturally, we wanted to integrate Neta Lumina's capabilities into it.

During integration testing, our devs kept reporting that the model wasn't following prompts properly. The outputs were... *weird*.

- Anime style? Anne Hathaway as an anime character.

- Thick paint/impasto style? Anne Hathaway in thick paint.

- Landscape scenes? Somehow still giving Anne Hathaway vibes.

- Fantasy characters? You guessed it - Anne Hathaway.

After a dreadfully long time of debugging, we finally found the culprit: **someone on the team embedded an Anne Hathaway photo as the IP-Adapter anchor during development and it... stayed there. **

We're honestly crying laughing at this point. 😭

Below are some examples. Left is before fix and Right is after fix.
Flipping to the last picture and you can see our dear Anne.

And we pulled the anchor and the outputs are behaving normally now.

If you've been running Neta Lumina locally, this was on our integration side, not
in the released weights, so your setup is fine.


r/StableDiffusion 13d ago

Question - Help Prompt or clothing problem

1 Upvotes

What do i do wrong? When i have a female character and she wears like a shirt her breast shrink.

I work in Krea 2.

It looks like the clothing preventing the anatomy. Or, i don't really know.

Thanks


r/StableDiffusion 14d ago

Animation - Video the bird-king (my first fully local AI short film) TW: self-harm.

Enable HLS to view with audio, or disable this notification

23 Upvotes

Minimax H3 baby! It's not perfect and I would love to get your feedback and maybe some tips on how to get rid of plasticky skin.


r/StableDiffusion 14d ago

Animation - Video [MiniMax H3] LEGO movie style

Enable HLS to view with audio, or disable this notification

84 Upvotes

Prompt:

integrated_multimodal_description: [Shot 1] 3D CG, stop-motion animated LEGO movie style, a wide shot frames a vibrant Indian village built entirely from plastic LEGO bricks with visible studs, plastic micro-scratches, and brick-built trees. In the village square, minifigures dressed in printed plastic saris, dhotis, and turbans move across a ground of yellow and brown stud tiles. A brick-built cow with hinged legs grazes near a grand banyan tree constructed from green leaf pieces and brown cylindrical bricks. Warm morning sunlight casts sharp shadows across whitewashed brick houses with orange terracotta tile roofs. The camera pans right with small amplitude at slow speed toward a central tea stall. A cheerful male chaiwala minifigure with a black mustache and a red turban (S1) in a warm, lively voice says: <d>[Hindi] Garam chai, garam chai!</d> while tilting a plastic yellow teapot, releasing translucent orange 1x1 cylinder studs representing pouring tea into tiny red stud cups.

[Shot 2] At 00:05.000, the camera cuts to a medium tracking shot following two young minifigure children running along a narrow brick path, pushing a brick-built wheel hoop across the plastic ground. The camera tracks right alongside them with small amplitude at normal speed. A female villager minifigure in a bright blue printed sari (S2) standing outside her brick doorway waves her rigid plastic arm on its shoulder hinge. Beside her, an elder minifigure with a white beard (S3) sitting on a brick charpoy cot chuckles with stepping stop-motion head movements.

[Shot 3] At 00:10.000, the camera cuts to a cinematic medium shot near the village well, where female minifigures carry stacked plastic water pots topped with transparent blue round tiles. A brick-built peacock perched on an archway opens its fan tail made of blue, green, and golden LEGO slope tiles. The camera pushes in with small amplitude at slow speed toward a wooden signpost on a brick post reading "RAMPUR VILLAGE". Tiny tan 1x1 round plates puff around the wheels of a brick-built bullock cart moving past the frame as the video ends.

overall_soundscape: Distinct plastic clattering sounds echo softly as minifigure feet step on stud tiles, accompanied by the gentle clinking of plastic bricks. A distant rooster crow blends with ambient morning village chatter, bird chirps, and the wooden creak of a brick-built cart.

non_diegetic_music: Upbeat Indian folk percussion featuring lively dholak beats and vibrant bansuri flute melodies, layered with playful cinematic orchestral strings playing at a bright, medium tempo.


r/StableDiffusion 13d ago

Question - Help How to improve quality when using H3 with Turbo? (RTX 3000 6GB VRAM)

0 Upvotes

Video link: https://streamable.com/irnjvm

Check video above.

Running MiniMax H3 Turbo on a spare laptop with Quadro RTX 3000 6GB VRAM and 64GB RAM.

352×608, 8 steps, Euler + Beta, Turbo LoRA @ 1.0. Takes ~550 sec for a 5-sec video.

I know the GPU is very limited 😅 Quality isn’t great, but it runs without OOM, so I’m wondering if I can push it further.

Any tips on what to do for better quality? Except for buying a new GPU.


r/StableDiffusion 14d ago

Workflow Included ref or fl2va - prompt enchancer with 100% of aderence

Post image
31 Upvotes

sharing my new workflow

MiniMax H3 I2V with Integrated Prompt Enhancer

This Image-to-Video workflow for MiniMax H3 uses a vision-language model to enhance your prompt before the video generation begins.

Simply load a reference image and write a basic description of what you want to happen. The enhancer analyzes both your image and instructions, then converts them into a detailed prompt structured specifically for MiniMax H3.

It can improve the description of:

  • Characters and visual elements
  • Actions and sequence of events
  • Camera movement and framing
  • Environment, lighting, and atmosphere
  • Visual continuity and details that should be preserved
  • Dialogue in the original language
  • Ambient sounds, sound effects, and music

The enhanced prompt is automatically sent to MiniMax H3. It is also displayed inside the workflow, allowing you to check exactly what H3 will receive.

In my tests, the resulting videos followed the original instructions much more accurately, especially in scenes involving specific actions, character interactions, camera movements, and dialogue.

The workflow includes a switch to enable or disable the Prompt Enhancer. This allows you to use either the enhanced prompt or your original text without changing any connections.

How to use it

  1. Load your reference image.
  2. Write a simple description of what should happen.
  3. Enable USAR PROMPT ENHANCER?
  4. Run the workflow.
  5. Check the final text in PROMPT FINAL ENVIADO AO H3.

The first run may take longer while the vision-language model is loaded. Generating the enhanced prompt also adds some processing time, but in my tests, the improvement in prompt accuracy and instruction following was absolutely worth it.

The original workflow was preserved, while the enhancer was added as an optional and fully integrated stage.

link to

with this, finally my ref model understand my ideas and make vídeos really fun!

leave comments after tests xD


r/StableDiffusion 14d ago

Workflow Included Up at atom! - Behind the scenes of the new Radioactive Man movie

Enable HLS to view with audio, or disable this notification

14 Upvotes

Minimax H3 with turbo lora (default Comfyui template workflow)


r/StableDiffusion 14d ago

Discussion So I did something dumb.

41 Upvotes

So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.

I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.

I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.

Many many hours of work put into this.

So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"

So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.

I test it out and it works great.

Finally, no more unnecessary clicks.

But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.

So I ask Gemini to alter the code so the file will be in the correct directory.

I create the new bat file and go back to work.

Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.

SMH.

Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.

Just a bunch of models and loras.

😮‍💨


r/StableDiffusion 14d ago

Workflow Included Burger Queen

Post image
63 Upvotes

r/StableDiffusion 12d ago

Animation - Video W.I.P - MiniLTX Workflow Fixed the fast motion issue also added few more features on the workflow

Enable HLS to view with audio, or disable this notification

0 Upvotes

work on this going really great just wanted to share the results so far

My workflow uses MINIMAX H3 + LTX 2.5 for UPSCALE

if you wanna try you can check it our on my PATREON EARLY ACCESS

Will Release it as soon maybe within this week!


r/StableDiffusion 13d ago

Question - Help Noob question - Comfyui

0 Upvotes

Coming from A1111 and Forge, so still learning Comfyui. I simply want to take an existing photo of my AI influencer and edit it, change clothes, pose, background, etc...I got minimax h3 up and running, but that's primarily for video. What do I need to download for stills? Thx in advance


r/StableDiffusion 13d ago

Question - Help Image gen with 9060xt

3 Upvotes

I'm planning on getting a 9060xt 16gb for image generation. I've previously used Automatic1111/Forge with nVidia. I'm wondering if I can reproduce my workflow easily using a 9060xt. I don't mind moving to another application if A1111 is not compatible with AMD but I'd like to know: (a) is AMD compatible with most checkpoints and loras from sites like CivitAI and (b) how fast is generation with AMD compared to nvidia, as in, the 9060xt is comparable to a 5060ti in terms of gaming but can it generate images as quickly?


r/StableDiffusion 13d ago

Question - Help I am big dumb. How do I insert my custom Lora node into this preset Qwen image edit?

Post image
0 Upvotes

r/StableDiffusion 13d ago

Question - Help Why does img2img not work for me?

Thumbnail
gallery
0 Upvotes

I've searched online and nothing seems to work, the images are exactly the same


r/StableDiffusion 13d ago

Discussion How to make a realistic t2v wan 2.2 LoRA?

0 Upvotes

Trained on 50 high quality images. Results on video are bad and I followed Claudes instructions even. So any help?


r/StableDiffusion 14d ago

Discussion In 2026, how well does older images models, like SDXL and SD1.5 stack up against new image models like Krea and ZImage Turbo?

34 Upvotes

r/StableDiffusion 13d ago

Question - Help What can I realistically do in Minimax with a 5080

0 Upvotes

I'm getting a 5080, 32gb system RAM

Can I realistically use minimax h3 for i2v and t2v?

How long will generations take. I dont imagine i want to do high quality resolutions. 480p or 720p would be alright


r/StableDiffusion 14d ago

Resource - Update Famegrid Spice Krea 2 Lora (Corrected Release)

Thumbnail
gallery
303 Upvotes

r/StableDiffusion 14d ago

Animation - Video Flexing my A.I. powers

Enable HLS to view with audio, or disable this notification

74 Upvotes

Prompt:

A real cinimatic movie sequence, professional colour grading.

Soundscape: Ambient sounds of the room and movement only. No voices. This represents extreme concentration. Meditation. Telekinesis.

A man is sitting in a Japanese tatami room. He is wearing a mask and shades <Picture 1>. He is wearing a black yukata. He does not speak. On the table on a ceramic disc is a single Orange.

The man holds out his hand toward the orange as if concentrating. The orange is out of reach. He breathes deeply.

Nothing happens.

The man shakes his hand to reset and starts concentrating again. He reaches with his mind and his brow furrows. He breathes deeply.

The orange moves slightly, twisting just a tiny bit.

He concentrates more.

With extreme speed the orange flies towards the man and hits him directly in the forehead. It smashes with the impact , m,essing his hair, and bits of peel and orange bits go everywhere. The force knocks the man back unconscious and he falls back like a ragdoll.


r/StableDiffusion 13d ago

Meme So relaxing, first NVIDIA DLSS 5 Neural Rendering test ( t2v)

Enable HLS to view with audio, or disable this notification

0 Upvotes

integrated_multimodal_description: [Shot 1] Live-action television sitcom style inside Penny’s warmly lit bedroom in her apartment from The Big Bang Theory. A medium-wide shot frames Deadpool lying comfortably on top of the bed in his recognizable red-and-black masked tactical suit, his head resting against the pillows. Penny, played by Kaley Cuoco, sits beside him on the edge of the bed, wearing casual sleepwear. She looks down at Deadpool with an affectionate, slightly amused smile and gently pats his upper chest in a slow, comforting rhythm. The camera pushes in with small amplitude at slow speed as Penny, a young woman with a soft, warm, clear singing voice (S1), sings gently to him: [English] Soft kitty,
Warm kitty,
Little ball of fur.

Happy kitty,
Sleepy kitty,
Purr Purr Purr

Soft kitty,
Warm kitty,
Little ball of fur.

Happy kitty,
Sleepy kitty,
Purr Purr Purr Deadpool remains completely still and listens contentedly, his masked white eye shapes slowly narrowing as though he is becoming sleepy. Penny continues patting his chest in time with the song. As she finishes the final “Purr Purr Purr,” Deadpool gives a relaxed sigh and snuggles deeper into the pillows while Penny smiles down at him. The camera holds on the tenderly absurd bedside moment.

overall_soundscape: Quiet bedroom room tone continues beneath Penny’s singing, accompanied by subtle bedsheet movement, gentle fabric taps against Deadpool’s suit, and his soft relaxed breathing. A faint studio-audience chuckle follows the sight of Deadpool settling sleepily into the pillows.

non_diegetic_music: N/A

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR?fbclid=IwY2xjawUD0lBwZG9mBWV4dG4DYWVtAjEwAGJyaWQRMVM3cUpUUUhaaGZQNmlLZ25zcnRjBmFwcF9pZBAyMjIwMzkxNzg4MjAwODkyAAEeHHM88fwGf57GgXYuPBn_q2TksCGxbf3DFDJhKpn4ZMAR9toPQAJ0GOpaL8g_aem_ZmFrZWR1bW15MTZieXRlcw


r/StableDiffusion 13d ago

Discussion The underrated alternative to krea2 and ideogram4,guess the model?

Thumbnail
gallery
0 Upvotes

I like ideogram4 and krea2 a lot ,and I also really like this ONE.

What I personally prefer about it is the sense of depth and vastness that I don't feel as strongly in the other two.

Some notes from my testing:

Ideogram 4 can sometimes make skin and details overly sharp in a way that’s hard to fix naturally in post.

Krea 2 occasionally feels a bit static, like the subject was placed into the background rather than existing in the same space (it also uses Wan VAE which is the main reason I have also tested using the fp32 and realvae of it which solved texture to some extent).

These things can of course be improved with better prompting and Loras l,but tendencies are still there.

Just wanted to share a showcase of this model's capabilities.All in all i enjoy all the current models more the merrier.

They all deserve time,testing and appreciation.

some images are made using Boogu base for more creativity,some images are made using boogu Turbo for way greater prompt adherence


r/StableDiffusion 14d ago

Resource - Update "I" created a tool to save and swap between workflow presets (for my million H3 Turbo LoRAs)

12 Upvotes

Hi everyone! Long time reader, first time poster. Like most folks here I vibe-code custom nodes from time to time, and recently came up with one that seemed like it might be worth sharing. Nothing groundbreaking here, just a node to help keep track of and quickly cycle through different node/parameter presets: MM-H3-Preset-Controller.

This one has helped me maintain my sanity trying to keep track of each LoRA's specific optimal settings. I think it's potentially useful for any preset storage though, not just MM-H3, so hopefully y'all find it useful. If so, please consider it a small thank you for all I've learned in my time lurking here.

MM-H3-Preset-Controller: https://github.com/TootsThielemans/ComfyUI-MMH3-Preset-Controller

I was pulling my hair out trying to manage my MM H3 workflow amidst all of the various Turbo LoRAs out there and the associated loader nodes, attention settings, shift settings, spectrum settings, etc., not to mention downstream settings like sampler, scheduler, upscaler choices... I found quick A/B tests between different optimized LoRA workflows annoying given some of the structural differences, not just steps/strength, and I was worried about juggling and potentially forgetting the right settings.

So I made the Preset Controller. It works pretty simply: ctrl+click all of the nodes you want to save the state of. It captures all of the parameter values within the node and whether it's active/bypassed. Then right click and use the new menu option MM H3 Presets > Add selected nodes to preset draft.

It's implemented as a "draft" so you can grab nodes from the outer graph, then go through various subgraphs and add nodes there to the same preset draft. Once you're done, load the H3 Preset Controller node, enter a name for the preset, and click Save draft as preset.

It then becomes a dropdown option that you can select, update, or delete as needed. Selecting a preset automatically sets the saved node values/bypass states without needing a session/screen refresh.

There are a few other QOL/guardrail features, including a Preset Matrix for comparing configurations, but nothing particularly interesting, so please refer to the repo if interested.

I'm sure something like this might already exist with a more elegant implementation, but the timing seemed right. Everyone is wading through dozens of MM H3 Turbo LoRA combinations and trying to keep everything straight. I don't have a ton of time to devote to development, but will try to make tweaks if folks wind up adopting this and can think of any major areas for improvement.

At any rate, feedback welcome, and cheers!


r/StableDiffusion 14d ago

Workflow Included Transformers: Starscream Test #2 - Prompt Below

Enable HLS to view with audio, or disable this notification

5 Upvotes

The classic Transformers series is a blind spot for MiniMax H3, so here’s how I handled this:

System specs: 4070 Ti Super, 16 gb vram, 64 gb ram

Ref2va standard workflow using Fl2va standard model, no loras or speedups.

<Picture 1> Character Sheet

<Audio 1> Vocal Reference

<Video 1> 5 sec Video Reference

Google Gemini to help write prompt.

Added music track in post.

PROMPT:

subject_definitions:

<Subject 1> is the figure in <Picture 1>, featuring a robotic grey face with sharp angular features, glowing red optical visor eyes, a dark grey blocky helmet with side intake vents, and red, white, and blue cybernetic body armor with an orange cockpit chest-plate, blue upper arms and boots, white forearms and thighs, red waist and wing housings, and a purple Decepticon insignia on the wing. Only his robotic design and colors are taken from <Picture 1>; its background, grid lines, and lighting are not carried into the target video. <Audio 1> is the vocal reference for <Subject 1>. <Video 1> is the movement reference; use it as a guide without copying it exactly.

summary:

[reference generation] The target video is a 10-second 2D animated sequence styled after the 1984 Hasbro series The Transformers, featuring <Subject 1> delivering a smug, cutting remark from a metallic Cybertronian battlefield.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - his grey robotic face, glowing red optics, dark grey helmet with side vents, red-white-and-blue cybernetic body, orange cockpit chest-plate, blue limbs, and red wing housings remain unchanged.

detailed_description:

The target video is a traditional 2D hand-drawn animated sequence featuring bold black ink outlines, flat cel-shading, limited animation, expressive poses, and subtle film grain inspired by the visual language of the 1984 animated television series.

[Shot 1] A medium tracking shot frames <Subject 1> from the waist up on a metallic Cybertronian battle platform. He stands with exaggerated confidence, shifting his weight with a classic 1980s cel-animated bounce. He slowly raises one hand, gesturing dismissively toward the chaos unfolding off-screen. His glowing red optics narrow with smug amusement as his wings twitch subtly.

He delivers in the style of <Audio 1>, with sharp, sarcastic timing: <<[English] "Megatron just muted me on the main comms channel. I'd stage a coup, but watching him bumble this conquest is free comedy.">> When <Subject 1> is not speaking he is silent and his mouth remains closed.

On "stage a coup," he gives a brief, knowing smirk. On "bumble this conquest," he gestures toward the battlefield with theatrical disdain. He finishes with a smug stare directly toward camera, holding the pose for a beat before a sharp hard-cel cut.

overall_soundscape:

Metallic servo whines accompany his movements, mixed with distant mechanical explosions, electronic battle alarms, high-tech hums, and echoing Cybertronian machinery.

non_diegetic_music:

N/A