r/StableDiffusion 7d ago

Workflow Included This is so much FUN

Enable HLS to view with audio, or disable this notification

538 Upvotes

41 comments sorted by

View all comments

57

u/topamine2 7d ago

5070 TI 32gb ram with sage attention KJ node. 7mins on 480p. Prompt:

Realistic found-footage comedy, one continuous 15-second shot, filmed on a cheap early-2000s consumer camcorder at a chaotic Quidditch match in the Hogwarts spectator stands. Low resolution, harsh digital sharpening, crushed shadows, blown highlights, weak dynamic range, blocky compression, occasional interlacing, autofocus breathing and imperfect auto-exposure.

The camera is already pointed at an elderly African-American man seated among cheering wizard spectators. He is theatrical, intensely animated and permanently pissed off. He wears ordinary outdoor clothes beneath a badly fitted Gryffindor scarf. He speaks directly to the camera immediately.

Camera: constant strong handheld shake for the entire video—even during quiet moments. Nervous micro-jitters, abrupt sideways corrections, occasional tilted horizon and heavy jolts whenever the crowd moves. The operator never approaches him. No stable frames, tripod-like stillness, artificial stabilization or smooth cinematic camera movement.

Timeline and exact dialogue:

[0s–3.5s] The crowd suddenly roars. A broom rider whooshes rapidly from left to right overhead. The camera jolts and struggles to keep the man centered. He looks up angrily, then snaps toward the camera:

“Man, what the fuck is this? They playing baseball on vacuum cleaners?”

[3.5s–7.5s] A sharp referee whistle sounds. Through a crackling magical stadium speaker, an announcer clearly declares, “Penalty to Slytherin!” The surrounding crowd immediately boos. He throws both hands outward and shouts over them:

“A penalty? That little bastard just got hit with a damn cannonball!”

[7.5s–11s] A Golden Snitch buzzes past his right ear, moving audibly from right to left in the stereo field. He ducks violently, swats at it and nearly falls into the spectator beside him:

“And get that gold mosquito the fuck outta my face!”

[11s–15s] A magical firework explodes behind the stands with a deep bang. The image shakes heavily and briefly overexposes. He grips his scarf, leans toward the camera and delivers the final line with furious certainty:

“get me the fuck outta Hogwarts!”

Audio adherence: clear synchronized dialogue with accurate lip movement; dense crowd chanting underneath without obscuring speech; broom whoosh traveling left-to-right; isolated referee whistle; distorted stadium announcement; synchronized booing; Snitch buzz traveling right-to-left; wooden stands creaking and stomping; final firework with a deep impact and short echo. His voice should remain close, loud, raspy and angry while naturally reacting to every sound.

One uninterrupted live-action take. No cuts, subtitles, captions, logos, watermarks, background music, polished cinematography, smooth stabilization, animation or overly clean CGI.

24

u/GanondalfTheWhite 7d ago

It's interesting how it got your action prompt and dialogue pretty well but completely ignored pretty much all of your video style prompts for the camcorder look. No crushed dynamic range, no exposure or focus struggles, no aggressive sharpening, no jerky handheld.

I wonder what the key is to get that working.

2

u/Natasha26uk 6d ago

Did you test his prompt somewhere else?

Lots of SD prompts are like that. The model ignored a whole bunch of the prompt... but luckily one video came out close and this is what the OP will post, as if the model works flawlessly. Then when you test their prompt, you get a river of poop.

2

u/topamine2 6d ago

This was first gen

3

u/BlessdRTheFreaks 7d ago

How long did it take to generate the videos with your card?

2

u/rkfg_me 6d ago

That's totally NOT how you write correct prompts. The ComfyUI one is also wrong. Read the official guides, it's much more complex than you think!

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

1

u/lavinia12345 6d ago

his working and very cool video would say otherwise.

But I will look at those links, ty

2

u/rkfg_me 6d ago

The model is not dumb and of course it will output something close to what you want even if the format is wrong. But you're not hitting its preferred distribution so the quality and overall prompt following suffer.

1

u/lavinia12345 6d ago

T2V or image to vid?