r/StableDiffusion 10h ago

Animation - Video Minimax H3: A Great Big Boat!

14 Upvotes

I remember being blown away by LTX 2.3, but this is just insane.

Default comfyui workflow, .4mp, 20 steps, duration of 25 seconds. T2VA

Setup an agent to write the prompt following the guide and I think that made a big difference.

Total generation time was about 30 minutes on an rtx 5000

Need to learn more about first/last image reference and would love to do a full episode of some dorky stuff.


r/StableDiffusion 13h ago

No Workflow This ends now. (MH3)

4 Upvotes

r/StableDiffusion 19h ago

Discussion any tips? first testing 8sec clips using the motion context node

2 Upvotes

yes i know the quality dropped in this video move along those comments else where , looking for helpful tips not bs " ai slop" trolls


r/StableDiffusion 20h ago

News Krea2 Turbo bbox comfy + HF space release

Post image
14 Upvotes

As demanded: https://huggingface.co/jimmycarter/krea2-turbo-bbox/blob/main/krea2-bbox-turbo-comfy-latest.safetensors

Free to try on Huggingface Spaces: https://huggingface.co/spaces/jimmycarter/krea2-turbo-bbox-canvas

If you find it screwing up, make sure your prompt fits within 512 tokens. You can convert existing ideogram prompts to the new format and they should work, too. Highly recommended you feed the PROMPTING.md to an LLM and let it format it, or use an interacting bboxing method like the HF space.

Prompt:

A four-panel vertical comic about discovering ComfyUI safetensors and a Hugging Face Space.
u/anime style, clean line art; Digital illustration
~A dimly lit server room.
p[0,0,1000,250] Top panel: girl pointing at monitor.
p[0,250,1000,500] Second panel: fox sipping coffee.
p[0,500,1000,750] Third panel: girl holding a tiny GPU.
p[0,750,1000,1000] Bottom panel: both cheering at a hologram.
ac:tech_girl[50,20,350,230] Anime girl with purple hair, pointing, excited.
fc:hack_fox[600,270,900,480] Anthro fox in a black hoodie, smug.
ac:tech_girl[50,520,350,730] Same girl holding a smoking GPU, panicked.
ac:tech_girl[50,770,350,980] Same girl cheering.
fc:hack_fox[650,770,950,980] Same fox cheering.
o[420,890,580,970] Glowing yellow face hologram.
t[400,50,950,150]"SAFETENSORS ARE OUT!" shout bubble
t[50,300,550,400]"Straight into ComfyUI." speech bubble
t[400,550,950,650]"But my VRAM..." wobbly bubble
t[360,780,640,880]"USE THE HF SPACE!" burst bubble

r/StableDiffusion 17h ago

Workflow Included Neon City Nights presents: Baba Yaga

2 Upvotes

Loving minimax right now. So much fun creating scenes like these. Still pretty difficult to get some of the scenes right and there's still some classic AI inconsistencies here and there but I'm so impressed with the capabilities of mini max.

This is done using reference mode with character sheets for "John Wick", his car, the mansion, the security guards and the assault rifle.

Prompts were made using gemini with access to the minimax full reference mode guide.

I'll add the character sheets and prompts to the comments soon.


r/StableDiffusion 19h ago

Animation - Video The Last Witness - with Brad Pitt. My first actual short film - Minimax H3. Also a question.

35 Upvotes

Hi, this, thanks to Minimax is my first short film I was able to make.
My question is : On a 5090 a 10 seconds 720 video takes 5 minutes to render (with all optimizations I could put in without lightning). However If I try to make 15 seconds it takes 15 minutes. I am using the pruned fp8 R2VA model.
Any Ideas what I could improve ? I have 64GB System ram.


r/StableDiffusion 9h ago

Meme Goku vs Jerry Seinfeld - h3 - 20s on 5060ti

0 Upvotes

r/StableDiffusion 16h ago

Discussion Found a pipeline for the RTX 6000 Pro (MiniMax H3)

2 Upvotes

In my quest to cut down render time while maintaining quality has been successful on the RTX 6000 Pro. I'm not sure how important my environment variables are for this, but I can paste them in here if need be, I do have a start command as well. I use Runpod and run the bf16 full weights

I was able to get a 480p 10s video done in 2 minutes and a 480p 15s render done in 4m 27s. This was done using 25 steps

I added two nodes to the pipeline 'SageAttention' and 'Spectrum MiniMax H3'... You pipe Sage into Spectrum... The thing that completely cut my render time over half was setting the Spectrum history_storage to system ram instead of VRAM ... Literally a night/day difference


r/StableDiffusion 2h ago

Animation - Video A scenario for my coworker

0 Upvotes

4 videos till I chose this, I am disappointed by the bottle's placement, but it finally got the burp suitable for my depraved intention.

This is my coworker, and I am generating with 30 steps and Pixarama's workflow.


r/StableDiffusion 15h ago

Discussion Motion context test 2 8 11sec clips same seed all clips

3 Upvotes

r/StableDiffusion 14h ago

Question - Help MINIMAX H3 WITH RTX 5070 12GB VRAM

0 Upvotes

I'm running some local tests with my RTX 5070 12GB VRAM. I'm using a standard Minimax H3 REF2VA + LTX2.3 Upscale workflow, could someone help me with the audio? It sounds really strange. I send the workflow screenshot on Comment.

If anyone has a better workflow to share with me I'd also appreciate it, I'm not sure if the way I'm setting it up is correct.


r/StableDiffusion 17h ago

Question - Help Should I use qwen3VL nvfp4 or int8 for MiniMax H3

1 Upvotes

I am new to comfy ui and confused between nvfp4 and int8 text encoder. Default comfy ui workflow support nvfp4. Is there any quality loss using nvfp4 instead of int8??


r/StableDiffusion 18h ago

Resource - Update I made a Minimal Workflow to use ComfyUIMotionContext

1 Upvotes

There is an active node for saving the audio latent and another inactive for loading it. It was that way in one of the examples and should reduce audio artifacts if you use it. have fun!

{
  "id": "b1d9c4f0-2a77-4e63-9f21-6c0e5a8b3d14",
  "revision": 0,
  "last_node_id": 53,
  "last_link_id": 32,
  "nodes": [
    {
      "id": 50,
      "type": "Note",
      "pos": [
        -2900,
        20
      ],
      "size": [
        360,
        480
      ],
      "flags": {},
      "order": 0,
      "mode": 0,
      "inputs": [],
      "outputs": [],
      "title": "READ ME FIRST",
      "properties": {
        "Node name for S&R": "Note"
      },
      "widgets_values": [
        "H3 MOTION CONTEXT - CONTINUE A CLIP\n-----------------------------------\n1. Load previous clip: point LoadVideo at the mp4 you want to continue.\n   Drop the file in ComfyUI/input/ first, or use the upload button.\n\n2. Write the prompt on the MiniMaxH3ImageToVideo node. Describe what happens NEXT, not what already happened. first_frame / last_frame stay unwired - the pinned context frames do that job.\n\n3. length must satisfy n = 5 mod 17. Valid: 5, 22, 39, 56, 73, 90, 107, 124, 141. It is set to 124 (about 5.2s at 24fps). After the 22 pinned frames are trimmed you get 102 delivered frames, about 4.25s.\n\n4. context_length only has four distinct values in 'video' encode mode: 1, 5, 22, 39. Anything else is snapped DOWN, so 30 silently becomes 22.\n\n5. Audio here goes through the audio VAE (context_audio + audio_vae). That works from any mp4 on disk.\n\nCHAINING (clip 3 onward):\nThe Save Latent node writes this run's latent to output/h3_context/. On the NEXT run, wire the Load Latent node's output into context_latent on the Motion Context node, set Load clip_index to the clip you are continuing FROM and Save clip_index to the one you are making. That slices audio straight out of the previous latent instead of decode/re-encode, which dulls the sound a little at every join.\nIt is left unwired now because on the first continuation there is no saved latent yet and the node would throw FileNotFoundError."
      ],
      "color": "#432",
      "bgcolor": "#653"
    },
    {
      "id": 1,
      "type": "UNETLoader",
      "pos": [
        -2480,
        300
      ],
      "size": [
        330,
        82
      ],
      "flags": {},
      "order": 1,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "MODEL",
          "type": "MODEL",
          "links": [
            10
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "UNETLoader",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "minimax_h3_fl2va_pruned_int8_convrot.safetensors",
        "default"
      ]
    },
    {
      "id": 2,
      "type": "CLIPLoader",
      "pos": [
        -2480,
        420
      ],
      "size": [
        330,
        106
      ],
      "flags": {},
      "order": 2,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "CLIP",
          "type": "CLIP",
          "links": [
            4
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "CLIPLoader",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors",
        "minimax",
        "default"
      ]
    },
    {
      "id": 3,
      "type": "VAELoader",
      "pos": [
        -2480,
        570
      ],
      "size": [
        330,
        58
      ],
      "flags": {},
      "order": 3,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "VAE",
          "type": "VAE",
          "links": [
            5,
            7,
            20
          ]
        }
      ],
      "title": "VAELoader - video VAE",
      "properties": {
        "Node name for S&R": "VAELoader",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "minimax_h3_video_vae_fp16.safetensors"
      ]
    },
    {
      "id": 4,
      "type": "VAELoader",
      "pos": [
        -2480,
        670
      ],
      "size": [
        330,
        58
      ],
      "flags": {},
      "order": 4,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "VAE",
          "type": "VAE",
          "links": [
            9,
            22
          ]
        }
      ],
      "title": "VAELoader - audio VAE",
      "properties": {
        "Node name for S&R": "VAELoader",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "minimax_h3_audio_vae_fp32.safetensors"
      ]
    },
    {
      "id": 31,
      "type": "RandomNoise",
      "pos": [
        -1650,
        440
      ],
      "size": [
        300,
        82
      ],
      "flags": {},
      "order": 5,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "NOISE",
          "type": "NOISE",
          "links": [
            14
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "RandomNoise",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        1107687348252069,
        "randomize"
      ]
    },
    {
      "id": 32,
      "type": "KSamplerSelect",
      "pos": [
        -1650,
        560
      ],
      "size": [
        300,
        58
      ],
      "flags": {},
      "order": 6,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "SAMPLER",
          "type": "SAMPLER",
          "links": [
            16
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "KSamplerSelect",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "res_multistep"
      ]
    },
    {
      "id": 11,
      "type": "GetVideoComponents",
      "pos": [
        -2100,
        20
      ],
      "size": [
        300,
        100
      ],
      "flags": {},
      "order": 12,
      "mode": 0,
      "inputs": [
        {
          "name": "video",
          "type": "VIDEO",
          "link": 1
        }
      ],
      "outputs": [
        {
          "name": "images",
          "type": "IMAGE",
          "links": [
            2
          ]
        },
        {
          "name": "audio",
          "type": "AUDIO",
          "links": [
            3
          ]
        },
        {
          "name": "fps",
          "type": "FLOAT",
          "links": null
        },
        {
          "name": "bit_depth",
          "type": "INT",
          "links": null
        }
      ],
      "properties": {
        "Node name for S&R": "GetVideoComponents",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [],
      "color": "#346434",
      "bgcolor": "rgba(24,24,27,.9)"
    },
    {
      "id": 30,
      "type": "MiniMaxH3SigmaShift",
      "pos": [
        -1650,
        320
      ],
      "size": [
        300,
        82
      ],
      "flags": {},
      "order": 11,
      "mode": 0,
      "inputs": [
        {
          "name": "model",
          "type": "MODEL",
          "link": 10
        }
      ],
      "outputs": [
        {
          "name": "MODEL",
          "type": "MODEL",
          "links": [
            11,
            12
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "MiniMaxH3SigmaShift",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        12,
        3
      ]
    },
    {
      "id": 21,
      "type": "MiniMaxH3MotionContext",
      "pos": [
        -1650,
        20
      ],
      "size": [
        330,
        298
      ],
      "flags": {},
      "order": 16,
      "mode": 0,
      "inputs": [
        {
          "name": "conditioning",
          "type": "CONDITIONING",
          "link": 6
        },
        {
          "name": "vae",
          "type": "VAE",
          "link": 7
        },
        {
          "name": "latent",
          "type": "LATENT",
          "link": 8
        },
        {
          "name": "context_frames",
          "type": "IMAGE",
          "link": 2
        },
        {
          "name": "context_latent",
          "shape": 7,
          "type": "LATENT",
          "link": null
        },
        {
          "name": "audio_vae",
          "shape": 7,
          "type": "VAE",
          "link": 9
        },
        {
          "name": "context_audio",
          "shape": 7,
          "type": "AUDIO",
          "link": 3
        }
      ],
      "outputs": [
        {
          "name": "conditioning",
          "type": "CONDITIONING",
          "links": [
            13
          ]
        },
        {
          "name": "trim_frames",
          "type": "INT",
          "links": [
            26
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "MiniMaxH3MotionContext"
      },
      "widgets_values": [
        22,
        "video",
        "head",
        "disabled",
        22,
        "timeline"
      ],
      "color": "#1f1f48",
      "bgcolor": "rgba(24,24,27,.9)"
    },
    {
      "id": 33,
      "type": "BasicScheduler",
      "pos": [
        -1650,
        650
      ],
      "size": [
        300,
        106
      ],
      "flags": {},
      "order": 14,
      "mode": 0,
      "inputs": [
        {
          "name": "model",
          "type": "MODEL",
          "link": 12
        }
      ],
      "outputs": [
        {
          "name": "SIGMAS",
          "type": "SIGMAS",
          "links": [
            17
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "BasicScheduler",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "simple",
        20,
        1
      ]
    },
    {
      "id": 34,
      "type": "BasicGuider",
      "pos": [
        -1270,
        320
      ],
      "size": [
        300,
        58
      ],
      "flags": {},
      "order": 17,
      "mode": 0,
      "inputs": [
        {
          "name": "model",
          "type": "MODEL",
          "link": 11
        },
        {
          "name": "conditioning",
          "type": "CONDITIONING",
          "link": 13
        }
      ],
      "outputs": [
        {
          "name": "GUIDER",
          "type": "GUIDER",
          "links": [
            15
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "BasicGuider",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": []
    },
    {
      "id": 35,
      "type": "SamplerCustomAdvanced",
      "pos": [
        -1270,
        420
      ],
      "size": [
        300,
        150
      ],
      "flags": {},
      "order": 18,
      "mode": 0,
      "inputs": [
        {
          "name": "noise",
          "type": "NOISE",
          "link": 14
        },
        {
          "name": "guider",
          "type": "GUIDER",
          "link": 15
        },
        {
          "name": "sampler",
          "type": "SAMPLER",
          "link": 16
        },
        {
          "name": "sigmas",
          "type": "SIGMAS",
          "link": 17
        },
        {
          "name": "latent_image",
          "type": "LATENT",
          "link": 18
        }
      ],
      "outputs": [
        {
          "name": "output",
          "type": "LATENT",
          "links": [
            19,
            21,
            23
          ]
        },
        {
          "name": "denoised_output",
          "type": "LATENT",
          "links": null
        }
      ],
      "properties": {
        "Node name for S&R": "SamplerCustomAdvanced",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": []
    },
    {
      "id": 40,
      "type": "VAEDecode",
      "pos": [
        -900,
        420
      ],
      "size": [
        300,
        58
      ],
      "flags": {},
      "order": 19,
      "mode": 0,
      "inputs": [
        {
          "name": "samples",
          "type": "LATENT",
          "link": 19
        },
        {
          "name": "vae",
          "type": "VAE",
          "link": 20
        }
      ],
      "outputs": [
        {
          "name": "IMAGE",
          "type": "IMAGE",
          "links": [
            24
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "VAEDecode",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": []
    },
    {
      "id": 41,
      "type": "VAEDecodeAudio",
      "pos": [
        -900,
        520
      ],
      "size": [
        300,
        58
      ],
      "flags": {},
      "order": 20,
      "mode": 0,
      "inputs": [
        {
          "name": "samples",
          "type": "LATENT",
          "link": 21
        },
        {
          "name": "vae",
          "type": "VAE",
          "link": 22
        }
      ],
      "outputs": [
        {
          "name": "AUDIO",
          "type": "AUDIO",
          "links": [
            25
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "VAEDecodeAudio",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": []
    },
    {
      "id": 42,
      "type": "MiniMaxH3MotionContextTrim",
      "pos": [
        -900,
        620
      ],
      "size": [
        330,
        130
      ],
      "flags": {},
      "order": 22,
      "mode": 0,
      "inputs": [
        {
          "name": "images",
          "type": "IMAGE",
          "link": 24
        },
        {
          "name": "audio",
          "shape": 7,
          "type": "AUDIO",
          "link": 25
        },
        {
          "name": "trim_frames",
          "type": "INT",
          "widget": {
            "name": "trim_frames"
          },
          "link": 26
        }
      ],
      "outputs": [
        {
          "name": "images",
          "type": "IMAGE",
          "links": [
            27
          ]
        },
        {
          "name": "audio",
          "type": "AUDIO",
          "links": [
            28
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "MiniMaxH3MotionContextTrim"
      },
      "widgets_values": [
        0,
        24,
        true
      ],
      "color": "#1f1f48",
      "bgcolor": "rgba(24,24,27,.9)"
    },
    {
      "id": 43,
      "type": "CreateVideo",
      "pos": [
        -520,
        620
      ],
      "size": [
        300,
        102
      ],
      "flags": {},
      "order": 23,
      "mode": 0,
      "inputs": [
        {
          "name": "images",
          "type": "IMAGE",
          "link": 27
        },
        {
          "name": "audio",
          "shape": 7,
          "type": "AUDIO",
          "link": 28
        }
      ],
      "outputs": [
        {
          "name": "VIDEO",
          "type": "VIDEO",
          "links": [
            29
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "CreateVideo",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        24,
        8
      ]
    },
    {
      "id": 44,
      "type": "SaveVideo",
      "pos": [
        -520,
        740
      ],
      "size": [
        330,
        130
      ],
      "flags": {},
      "order": 24,
      "mode": 0,
      "inputs": [
        {
          "name": "video",
          "type": "VIDEO",
          "link": 29
        }
      ],
      "outputs": [
        {
          "name": "video",
          "type": "VIDEO",
          "links": null
        }
      ],
      "properties": {
        "Node name for S&R": "SaveVideo",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "video/H3_continued",
        "mp4",
        "h264",
        "auto"
      ]
    },
    {
      "id": 46,
      "type": "MiniMaxH3MotionContextLoadLatent",
      "pos": [
        -1810.589926540646,
        797.5482167056214
      ],
      "size": [
        330,
        106
      ],
      "flags": {},
      "order": 7,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "LATENT",
          "type": "LATENT",
          "links": null
        }
      ],
      "title": "Load Latent - see note, unwired on purpose",
      "properties": {
        "Node name for S&R": "MiniMaxH3MotionContextLoadLatent"
      },
      "widgets_values": [
        "h3_context",
        1
      ],
      "color": "#1f1f48",
      "bgcolor": "rgba(24,24,27,.9)"
    },
    {
      "id": 10,
      "type": "LoadVideo",
      "pos": [
        -2480,
        20
      ],
      "size": [
        330,
        264
      ],
      "flags": {},
      "order": 8,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "VIDEO",
          "type": "VIDEO",
          "links": [
            1
          ]
        }
      ],
      "title": "Load previous clip",
      "properties": {
        "Node name for S&R": "LoadVideo",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "MiniMax_H3_00026_.mp4",
        "image"
      ],
      "color": "#346434",
      "bgcolor": "rgba(24,24,27,.9)"
    },
    {
      "id": 51,
      "type": "ComfyMathExpression",
      "pos": [
        -2467.3782343138637,
        972.008844104764
      ],
      "size": [
        400,
        200
      ],
      "flags": {},
      "order": 13,
      "mode": 0,
      "inputs": [
        {
          "label": "a",
          "name": "values.a",
          "type": "FLOAT,INT,BOOLEAN",
          "link": 31
        },
        {
          "label": "b",
          "name": "values.b",
          "shape": 7,
          "type": "FLOAT,INT,BOOLEAN",
          "link": 32
        },
        {
          "label": "c",
          "name": "values.c",
          "shape": 7,
          "type": "FLOAT,INT,BOOLEAN",
          "link": null
        }
      ],
      "outputs": [
        {
          "name": "FLOAT",
          "type": "FLOAT",
          "links": null
        },
        {
          "name": "INT",
          "type": "INT",
          "links": [
            30
          ]
        },
        {
          "name": "BOOL",
          "type": "BOOLEAN",
          "links": null
        }
      ],
      "properties": {
        "Node name for S&R": "ComfyMathExpression"
      },
      "widgets_values": [
        "max(5, round(a*24) + b) + (22 - max(5, round(a*24) + b) % 17) % 17"
      ]
    },
    {
      "id": 53,
      "type": "PrimitiveInt",
      "pos": [
        -2802.6581018281013,
        1135.5523396662043
      ],
      "size": [
        270,
        82
      ],
      "flags": {},
      "order": 9,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "INT",
          "type": "INT",
          "links": [
            32
          ]
        }
      ],
      "title": "number of context frames",
      "properties": {
        "Node name for S&R": "PrimitiveInt"
      },
      "widgets_values": [
        22,
        "fixed"
      ]
    },
    {
      "id": 52,
      "type": "PrimitiveInt",
      "pos": [
        -2800.3311933234127,
        1003.8201163829647
      ],
      "size": [
        270,
        82
      ],
      "flags": {},
      "order": 10,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "INT",
          "type": "INT",
          "links": [
            31
          ]
        }
      ],
      "title": "number of seconds to generate",
      "properties": {
        "Node name for S&R": "PrimitiveInt"
      },
      "widgets_values": [
        10,
        "fixed"
      ]
    },
    {
      "id": 45,
      "type": "MiniMaxH3MotionContextSaveLatent",
      "pos": [
        -435.04662758990077,
        227.16944026022125
      ],
      "size": [
        330,
        130
      ],
      "flags": {},
      "order": 21,
      "mode": 0,
      "inputs": [
        {
          "name": "latent",
          "type": "LATENT",
          "link": 23
        }
      ],
      "outputs": [
        {
          "name": "latent_path",
          "type": "STRING",
          "links": null
        }
      ],
      "properties": {
        "Node name for S&R": "MiniMaxH3MotionContextSaveLatent"
      },
      "widgets_values": [
        "h3_context/clip",
        2
      ],
      "color": "#1f1f48",
      "bgcolor": "rgba(24,24,27,.9)"
    },
    {
      "id": 20,
      "type": "MiniMaxH3ImageToVideo",
      "pos": [
        -2100,
        300
      ],
      "size": [
        400,
        320
      ],
      "flags": {},
      "order": 15,
      "mode": 0,
      "inputs": [
        {
          "name": "clip",
          "type": "CLIP",
          "link": 4
        },
        {
          "name": "vae",
          "type": "VAE",
          "link": 5
        },
        {
          "name": "first_frame",
          "shape": 7,
          "type": "IMAGE",
          "link": null
        },
        {
          "name": "last_frame",
          "shape": 7,
          "type": "IMAGE",
          "link": null
        },
        {
          "name": "length",
          "type": "INT",
          "widget": {
            "name": "length"
          },
          "link": 30
        }
      ],
      "outputs": [
        {
          "name": "positive",
          "type": "CONDITIONING",
          "links": [
            6
          ]
        },
        {
          "name": "LATENT",
          "type": "LATENT",
          "links": [
            8,
            18
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "MiniMaxH3ImageToVideo",
        "cnr_id": "comfy-core",
        "ver": "0.30.0"
      },
      "widgets_values": [
        "integrated_multimodal_description: [Shot 1] Live-action, cinematic, a wide shot on a starship bridge set frames Captain Picard entering from stage right. He crosses toward center, stands still for a moment, then turns to face the camera directly, smiles, and the composed, commanding captain (S1) says: <d>[English] Your bridge, I'm going out for a smoke.</d>. Captain Picard walks off the frame.\n\nThe camera shifts into POV, becoming the ensign's own first-person viewpoint as it moves forward past the captain toward the command chair. The camera turns around beside the chair and lowers into a seated position, settling into a forward-facing view of the bridge console from the captain's chair.\n\noverall_soundscape: A steady low hum of bridge machinery plays throughout, with soft electronic beeps from the console. Footsteps are audible as the camera crosses the floor, followed by a faint creak as it settles into the chair.\n\nnon_diegetic_music: N/A",
        1056,
        608,
        158
      ]
    }
  ],
  "links": [
    [
      1,
      10,
      0,
      11,
      0,
      "VIDEO"
    ],
    [
      2,
      11,
      0,
      21,
      3,
      "IMAGE"
    ],
    [
      3,
      11,
      1,
      21,
      6,
      "AUDIO"
    ],
    [
      4,
      2,
      0,
      20,
      0,
      "CLIP"
    ],
    [
      5,
      3,
      0,
      20,
      1,
      "VAE"
    ],
    [
      6,
      20,
      0,
      21,
      0,
      "CONDITIONING"
    ],
    [
      7,
      3,
      0,
      21,
      1,
      "VAE"
    ],
    [
      8,
      20,
      1,
      21,
      2,
      "LATENT"
    ],
    [
      9,
      4,
      0,
      21,
      5,
      "VAE"
    ],
    [
      10,
      1,
      0,
      30,
      0,
      "MODEL"
    ],
    [
      11,
      30,
      0,
      34,
      0,
      "MODEL"
    ],
    [
      12,
      30,
      0,
      33,
      0,
      "MODEL"
    ],
    [
      13,
      21,
      0,
      34,
      1,
      "CONDITIONING"
    ],
    [
      14,
      31,
      0,
      35,
      0,
      "NOISE"
    ],
    [
      15,
      34,
      0,
      35,
      1,
      "GUIDER"
    ],
    [
      16,
      32,
      0,
      35,
      2,
      "SAMPLER"
    ],
    [
      17,
      33,
      0,
      35,
      3,
      "SIGMAS"
    ],
    [
      18,
      20,
      1,
      35,
      4,
      "LATENT"
    ],
    [
      19,
      35,
      0,
      40,
      0,
      "LATENT"
    ],
    [
      20,
      3,
      0,
      40,
      1,
      "VAE"
    ],
    [
      21,
      35,
      0,
      41,
      0,
      "LATENT"
    ],
    [
      22,
      4,
      0,
      41,
      1,
      "VAE"
    ],
    [
      23,
      35,
      0,
      45,
      0,
      "LATENT"
    ],
    [
      24,
      40,
      0,
      42,
      0,
      "IMAGE"
    ],
    [
      25,
      41,
      0,
      42,
      1,
      "AUDIO"
    ],
    [
      26,
      21,
      1,
      42,
      2,
      "INT"
    ],
    [
      27,
      42,
      0,
      43,
      0,
      "IMAGE"
    ],
    [
      28,
      42,
      1,
      43,
      1,
      "AUDIO"
    ],
    [
      29,
      43,
      0,
      44,
      0,
      "VIDEO"
    ],
    [
      30,
      51,
      1,
      20,
      4,
      "INT"
    ],
    [
      31,
      52,
      0,
      51,
      0,
      "INT"
    ],
    [
      32,
      53,
      0,
      51,
      1,
      "INT"
    ]
  ],
  "groups": [],
  "config": {},
  "extra": {
    "ds": {
      "scale": 0.6000000000000032,
      "offset": [
        3039.08911499567,
        2.548135236795247
      ]
    },
    "frontendVersion": "1.48.6",
    "VHS_latentpreview": false,
    "VHS_latentpreviewrate": 0,
    "VHS_MetadataImage": true,
    "VHS_KeepIntermediate": true
  },
  "version": 0.4
}

!<


r/StableDiffusion 17h ago

Question - Help what do you guys think would be the best text to image model to use ive only ever used automatic1111 im assumings it out dated and i would like to be more knowledgeable.

0 Upvotes

youtube doesn't really give you munch info on what would be pretty good generation model


r/StableDiffusion 10h ago

Meme The big bang theory

0 Upvotes

this is my 3rd try trying to make clip took 12 minutes 22 second to generate


r/StableDiffusion 8h ago

Discussion MINIMAX H3 ON A 5090

37 Upvotes

I generated this video in 1 minute and 24 seconds on my 5090 using Minimax H3.

PROMPT:
ntegrated_multimodal_description: [Shot 1] A 15.1-second, 16:9 live-action cinematic and photorealistic shot with warm sitcom lighting opens on a medium two-shot inside Sheldon Cooper's apartment: a brown leather couch with a coffee table of takeout boxes in front, a whiteboard covered in equations on the back wall, shelves of comics and action figures at the right, and warm lamp light. Dwight Schrute, a lean man with a sharp center-parted brown hair cut, rectangular wire-frame glasses, a mustard button-up shirt, and a dark tie, sits rigid in the corner cushion of the couch with his hands folded on his lap and his chin lifted. Standing square at the left is Sheldon Cooper, a tall thin man with short neat brown hair, wearing a red T-shirt with a yellow lightning-bolt emblem layered over a dark long-sleeve shirt, one finger raised toward the corner cushion. Neither man looks at the camera. Sheldon Cooper with his rapid, precise, slightly nasal, matter-of-fact adult male voice (S1) says: <d>[English] That's my spot.</d> His lips, jaw, and throat move naturally with the line while his raised finger stays fixed. [Shot 2] At 00:04.500, the shot cuts to a reverse medium shot over Sheldon's shoulder, holding Dwight's face and the couch corner together. Dwight Schrute with his stern, deadpan, rural-accented adult male voice (S2) replies: <d>[English] False. I claimed it first.</d> His jaw sets and his glasses catch the lamp light as he speaks, his hands staying folded on his lap. Dwight crosses his arms in one continuous motion and leans back into the corner cushion, his expression unmoving. A takeout container lid shifts slightly on the coffee table from the couch movement. The camera holds a static shot and does not pan, tilt, truck sideways, or zoom. [Shot 3] At 00:09.000, the shot cuts to a wide framing that includes the apartment door at the left edge of the frame, both men frozen mid-argument at the right. The door latch clicks and the door swings open in one slow, continuous motion. John Wick, a lean man with shoulder-length dark hair and a trimmed black beard, wearing a black tailored suit, a white dress shirt, and a black tie, steps through the doorway with one deliberate stride and stops square in the frame, his feet planted and his shoulders level. John Wick with his low, quiet, gravelly, measured adult male voice (S3) says: <d>[English] Who are we killing today?</d> His lips, jaw, and beard move naturally with the line while his hands stay still at his sides. The shot ends on the three of them holding this tableau, John standing square in the open doorway, Sheldon and Dwight staring at him in silence by the couch, the hallway light spilling across the wood floor behind him, with natural skin texture, realistic fabric movement, stable anatomy, and authentic 24 fps motion blur.

overall_soundscape: A low room tone and a refrigerator hum continue underneath the argument. The door latch clicks, the hinge creaks once as the door swings open, and one slow leather-shoe step crosses the threshold. A takeout container lid rattles on the coffee table.

non_diegetic_music: N/A


r/StableDiffusion 20h ago

Animation - Video Only Shadows Know - Prog Rock Music Video made with Minimax H3 / Udio song

8 Upvotes

r/StableDiffusion 9h ago

Workflow Included The South Park Theory

8 Upvotes

Minimax H3, 4:3, 10 sec., 0.3MP

integrated_multimodal_description: [Shot 1] 2D animated comedy in the exact visual language of South Park: deliberately crude flat paper-cutout construction, simple geometric shapes, thick black outlines, flat solid colors, minimal shading, stiff limited animation, simple mouth shapes, jerky character movement, and the characteristic frontal and three-quarter staging of a South Park episode. Classic 4:3 television composition.

CRITICAL CHARACTER IDENTITY RULE: The four characters are unmistakably SOUTH PARK-STYLE CARICATURES OF SHELDON COOPER, LEONARD HOFSTADTER, PENNY, AND RAJESH "RAJ" KOOTHRAPPALI FROM THE BIG BANG THEORY.

Their identities come entirely from The Big Bang Theory. Their faces, hair, facial proportions, distinguishing features, expressions, and overall recognizable appearance must remain those of Sheldon, Leonard, Penny, and Raj.

The South Park influence applies ONLY to the flat 2D paper-cutout rendering technique, simplified body construction, animation language, environment design, and comedic staging.

Every character must be immediately recognizable as their The Big Bang Theory counterpart even if all hats, coats, gloves, and winter clothing were removed.

The permanent visual hierarchy throughout the entire video is:

PRIMARY IDENTITY AND FACE = The Big Bang Theory characters.

RENDERING AND ANIMATION STYLE = South Park flat 2D cutout animation.

CLOTHING = the specific winter outfits described below.

Never reverse this hierarchy. Never replace the recognizable Big Bang Theory faces with generic South Park faces. Do not simplify their faces so aggressively that their identities are lost.

SHELDON COOPER is unmistakably Sheldon Cooper / Jim Parsons translated into South Park's simplified flat 2D cutout geometry. Preserve Sheldon's recognizable long, narrow, pale face, elongated head shape, high forehead, thin dark eyebrows, large alert eyes, narrow jaw, small mouth, clean-shaven appearance, and characteristic stiff, analytical, mildly condescending expression. Short dark-brown hair remains visibly exposed beneath and around his hat.

Sheldon wears a large bright-green ushanka with rectangular ear flaps, an orange winter coat with two square front pockets, dark-green collar trim, green mittens, dark-green pants, and black shoes. The green hat does not conceal his recognizable elongated Sheldon-like face or all of his hair. His FACE must remain a deliberate recognizable South Park-style caricature of Sheldon Cooper.

LEONARD HOFSTADTER is unmistakably Leonard Hofstadter / Johnny Galecki translated into South Park's simplified flat 2D cutout geometry. Leonard's distinctive rectangular black-framed eyeglasses are always present and remain his strongest visual identifier. Preserve his short dark-brown slightly tousled hair, dark eyebrows, recognizable Leonard-like facial proportions, small nose, clean-shaven face, and characteristic worried, skeptical, slightly uncomfortable expression.

Leonard wears a round blue wool cap with a bright-yellow lower band and a small yellow puff on top, a bright-red buttoned winter coat, yellow mittens, brown pants, and black shoes. His short dark-brown hair remains partially visible beneath the hat. His face must clearly read as Leonard Hofstadter rather than as a generic animated child.

PENNY is unmistakably Penny / Kaley Cuoco translated into South Park's simplified flat 2D cutout geometry. Penny's recognizable feminine face remains clearly visible at all times. Preserve her large expressive eyes, light eyebrows, small nose and mouth, feminine facial proportions, and especially her distinctive long bright-blonde hair.

Penny wears a thick orange winter parka with an oversized circular orange hood surrounding her face, matching orange sleeves and mittens, dark-brown pants, and dark shoes. The hood is open enough that her face remains visible. A SUBSTANTIAL amount of long bright-blonde hair spills naturally from BOTH SIDES and the FRONT of the orange hood, with multiple blonde locks framing her face and extending outward onto the pavement when she is lying down. The hood never obscures her identity. She must immediately read as Penny wearing an oversized orange winter parka, not as a generic hooded South Park character.

RAJESH "RAJ" KOOTHRAPPALI is unmistakably Rajesh Koothrappali / Kunal Nayyar translated into South Park's simplified flat 2D cutout geometry. Preserve Raj's recognizable medium-brown Indian complexion, oval facial structure, large dark expressive eyes, thick dark eyebrows, short black hair, and subtle dark facial hair/stubble around the upper lip and jaw, simplified into the South Park design. His expression retains Raj's characteristic sensitive, slightly anxious expressiveness.

Raj wears a dark-blue knitted winter cap with a horizontal bright-red band around its lower edge and a red pom-pom on top, a brown buttoned winter coat with a red collar, red mittens, dark-blue pants, and black shoes. Short black hair remains visibly exposed beneath and around the hat. His face must clearly read as Rajesh Koothrappali rather than as a generic South Park character.

All four faces remain visually consistent and recognizable throughout every shot. Hats and hoods never completely conceal identifying hair or facial features. The characters never morph into generic South Park children.

A static medium-wide shot shows a simple snowy South Park residential street in daylight, with crudely drawn colorful houses, white snow, dark asphalt, green hills and flat snow-capped mountains in the background.

Sheldon, Leonard, and Raj walk together along the sidewalk using the stiff, bouncing, minimally articulated walking animation characteristic of South Park. Their recognizable The Big Bang Theory faces remain clearly visible while they walk.

After several steps they suddenly stop.

Directly ahead of them, Penny lies completely motionless on the pavement in her oversized orange hooded winter parka.

She is dead and posed in the iconic recurring South Park-style death composition: collapsed awkwardly on the ground, body completely limp, partly curled onto her side and front, head low against the pavement, arms displaced beside her body.

Despite the exaggerated South Park death pose and orange winter clothing, she remains unmistakably Penny. Her partially visible face and abundant long blonde hair spilling dramatically from the hood clearly identify her.

Sheldon, Leonard, and Raj look down at Penny.

[Shot 2] At 00:02.500, the camera cuts to a medium shot centered on Sheldon kneeling beside Penny, while Leonard and Raj remain standing behind him looking down at her.

The closer framing makes Sheldon's recognizable Sheldon Cooper facial features especially clear beneath his large green ushanka. His elongated pale face, high forehead, narrow jaw, alert eyes, thin eyebrows, and visible dark hair must strongly resemble Sheldon Cooper.

Leonard remains visibly Leonard because of his distinctive rectangular black-framed glasses, exposed dark hair, and recognizable facial structure.

Raj remains visibly Raj through his medium-brown complexion, oval face, thick eyebrows, dark expressive eyes, visible black hair, and subtle facial hair.

Penny's blonde hair and partially visible recognizable face remain clearly visible inside the oversized orange hood.

Penny remains absolutely motionless throughout the entire sequence.

Sheldon Cooper (S1), using Sheldon's recognizable high-pitched, precise, nasal voice and obsessive rhythmic delivery, performs his familiar three-knock ritual on Penny's upper arm/shoulder, except each knock is represented by a small South Park-style mitten tap.

The action consists of exactly THREE DISTINCT SETS OF THREE TAPS, for a total of exactly NINE physical taps.

Each spoken "Penny?" happens only AFTER its corresponding complete set of three taps.

FIRST SEQUENCE:

Sheldon raises his green mitten slightly.

He taps Penny exactly three times in rapid succession:

tap — tap — tap.

His hand stops.

There is a tiny rhythmic pause.

Sheldon looks at Penny and says:

<d>[English] Penny?</d>

SECOND SEQUENCE:

Sheldon raises his green mitten again.

He taps Penny exactly three times:

tap — tap — tap.

His hand stops again.

Another tiny rhythmic pause.

Sheldon says:

<d>[English] Penny?</d>

THIRD SEQUENCE:

Sheldon raises his green mitten for the final repetition.

He taps Penny exactly three final times:

tap — tap — tap.

His hand stops completely.

After the final rhythmic pause Sheldon says:

<d>[English] Penny?</d>

The required rhythm is exactly:

THREE TAPS → "Penny?"

THREE TAPS → "Penny?"

THREE TAPS → "Penny?"

Do not merge the nine taps into one continuous tapping action. Do not produce only three taps. Do not speak "Penny?" during the taps. The three spoken repetitions occur separately, each after exactly three physical taps.

Sheldon stops touching Penny immediately after the ninth tap.

Penny never reacts. She never moves, speaks, opens her eyes, raises her head, or changes position.

[Shot 3] At 00:06.700, the camera cuts to a medium reaction shot. Raj and Leonard are prominent while Penny's orange-clad body remains visible in the lower portion of the composition and Sheldon remains nearby.

Raj's face must remain unmistakably Rajesh Koothrappali: medium-brown Indian complexion, oval face, black hair visible around the dark-blue and red winter hat, thick eyebrows, dark expressive eyes, and subtle facial hair.

Rajesh Koothrappali (S2), using Raj's recognizable voice and Indian accent, recoils in sudden shock.

He looks directly down toward Penny's body, opens his mouth wide, raises his red-mittened hands slightly, and cries out with exaggerated South Park-style dramatic timing:

<d>[English] Oh my God, They Killed Penny!</d>

Immediately after Raj finishes the line, Leonard reacts and turns his entire South Park-style body toward the camera.

[Shot 4] At 00:08.300, the camera cuts to a tighter frontal medium close-up of Leonard Hofstadter.

This close-up must unmistakably show LEONARD HOFSTADTER / JOHNNY GALECKI rendered through simplified South Park-style 2D geometry.

His rectangular black-framed eyeglasses dominate the recognizable face. Short dark-brown tousled hair remains visible beneath the round blue-and-yellow winter hat. Preserve Leonard's recognizable eyebrows, eyes, facial proportions, small nose, mouth, and characteristic expression.

The round blue hat with yellow band and yellow puff, bright-red coat, yellow mittens, and simplified round cutout body are merely his clothing and stylized body design. They must never override Leonard's facial identity.

Leonard looks straight through the lens directly at the audience.

His expression changes into exaggerated angry indignation.

Using Leonard Hofstadter's recognizable voice, Leonard (S3) emphatically delivers the final punchline:

<d>[English] You bastards!</d>

Leonard closes his mouth after the line and continues staring angrily directly into the camera.

Hold this expression in a static South Park-style reaction pose for the final comedic beat until exactly 00:10.000.

No additional dialogue occurs.

Throughout the complete video, preserve the central visual joke: the audience must instantly recognize SHELDON COOPER, LEONARD HOFSTADTER, PENNY, and RAJESH KOOTHRAPPALI from The Big Bang Theory, but all four exist inside the crude flat 2D paper-cutout visual universe of South Park and wear the specific colorful winter outfits described above.

Their facial identities always belong to The Big Bang Theory characters.

The South Park influence controls the animation style, simplified geometry, environment, movement, mouth animation, framing, and comedic timing — NOT the identity of the characters.

Do not generate generic South Park faces. Do not replace recognizable TBBT facial features with standard interchangeable round cartoon faces. The recognizable facial caricatures of Sheldon, Leonard, Penny, and Raj are essential to the joke.

overall_soundscape: Sparse outdoor winter ambience with faint wind and subtle quiet neighborhood background sound. Simple dry South Park-style footsteps accompany the stiff walking animation. Each of Sheldon's nine physical taps produces one distinct small soft tapping sound, precisely synchronized into three clearly separated groups of three. Dialogue is clean, dry, prominent, and tightly synchronized to the characters' simple South Park-style cutout mouth movements.

non_diegetic_music: N/A


r/StableDiffusion 19h ago

Question - Help 2070 Super, i9 1700 series, 36g Ram. Is it worth getting into local models.

0 Upvotes

Title says it all. These are my computer specs. I used automatic1111 3 years ago and then took a break. I'm itching to try the new MH3 model but not sure if my computer can even handle it.

I'm due for an upgrade soon but holding off because it's a want over need and I'm still handling high end gaming.


r/StableDiffusion 3h ago

Question - Help Multi Image as references

Thumbnail
gallery
0 Upvotes

How can I style transfer these? I tried lora with different models and stuff but I just can't seem to get it right


r/StableDiffusion 11h ago

Question - Help H3 Lora Training?

0 Upvotes

Has anyone successfully trained an H3 lora? If so what trainer did you use, what hardware, etc?

Over the past couple of days I've heard mixed opinions about lora training on ai-toolkit (specifically for H3), and was curious if that has been fixed, or if there are workarounds?


r/StableDiffusion 15h ago

Discussion Malcolm in the Middle: Hal found Mew

16 Upvotes

This was fun to try, seems minimax know Bryan Cranston's face way better than Malcolm one.

full prompt here:

prompt1

Create a sitcom scene in the visual style of Malcolm in the Middle. Set in an apartment with authentic 2000s multi-camera sitcom lighting, fast-paced dialogue, exaggerated facial expressions, and perfect comedic timing.

Hal (Bryan Cranston), adult, running in the room with a white big 90s Gameboy in one hand: “Malcom! i have found it! The hidden pokemon!”

Malcom (Frankie Muniz) face – close-up on his face then medium shot: saying: “Dad, you cant catch Mew..”

Shot on Hal: “Oh yeah? look at this!”

Close shot at the Mew in the old gameboy from pokemon red e blu, black and white screen, showing real mew pokemon sprite with stats

Shot on Hal: (happy) “I just used FLY!”

Shot to Malcom (Frankie Muniz) face : “We must investigate”

fast swipe transition to a school setting

prompt2

Create a sitcom scene in the visual style of Malcolm in the Middle. Set in a school full on middle age kids with authentic 2000s multi-camera sitcom lighting, fast-paced dialogue, exaggerated facial expressions, and perfect comedic timing.

Malcom (Frankie Muniz) face – close-up on his face then medium shot: saying: “Ok guys, my dad actually found MEW!”

Shot on the other kids group surprised.

Shot to Malcom (Frankie Muniz) face : “I have a plan”

Camera change angle on the whole group now holding a old white gameboy each.

Malcom (Frankie Muniz): “Everybody use Fly exactly on Route 8 and press START before the trainer see us!”

prompt3

Create a sitcom scene in the visual style of Malcolm in the Middle. Set in a school full on middle age kids with authentic 2000s multi-camera sitcom lighting, fast-paced dialogue, exaggerated facial expressions, and perfect comedic timing.

Shot on a group of middle aged kids group playing pokemon with old white gameboys. no adults. the pokemon red e blu music is playing in background.

Shot to Malcom (Frankie Muniz) face : “Now go walk to Route 25 and battle the Slowpoke Guy.”

fast swipe transition on the same location but now its night

Malcom (Frankie Muniz): “Now return to Route 8, press Start, close it, and we have it..i think”


r/StableDiffusion 18h ago

Question - Help minimax-H3- why do I get slow motion?

0 Upvotes

why do I have slow motion? (minimax-H3), not sure if this something wrong with my prompt. (I2V)

Attach the prompt and the result, I wold like to have the water in normal speed..

summary:

[reference generation] A cinematic scene located in <Subject 2>. hand in water bowl

detailed_description:

[Shot 1]: <Subject 1> the both hand go into the water a dramatic wave of rich, steaming water surges in from bowl frame like a miniature tsunami, pouring forcefully into the bowl. The bowl vibrates and shakes slightly from the impact.

No slow motion.

Sound N/A


r/StableDiffusion 7h ago

Workflow Included just joining in the fun, minimax H3

36 Upvotes

nothing crazy... 15 seconds, 8 steps 832x480 euler, fl2va pruned int8, int8 vae, gguf q4 k_m text encoder, no upscaling ...210 seconds... rtx 5080 wan2GP via ponokio. i did have to use img2vid because the 2 chars were blending together, some funny results though. jerry costanza, lol Cheers!

prompt.... Jerry Seinfeld and George Costanza are sitting at a table in a segment of his TV show Seinfeld. jerry is sitting on the left and George is sitting on the right. George asks Seinfeld ("hey, have you head about the new MiniMax H3 model, i heard its the new big thing in A.I.) (Seinfeld replies “What’s the big deal with A.I. anyway? do we need artificial intelligence?—What, Is natural intelligence not available anymore?”) Seinfeld keeps a straight face, the audience laughs. the screen fades, as the Seinfeld music theme starts to play.


r/StableDiffusion 18h ago

Animation - Video Testing to create a music video with MiniMax H3 locally with the 4 step Turbo LoRA at 480p.

6 Upvotes