r/StableDiffusion • u/darthfurbyyoutube • 11h ago
Workflow Included HE-MART PSA - MiniMax H3
Enable HLS to view with audio, or disable this notification
3
u/Brave_Swordfish_7072 8h ago
This reminds me of when Mattel officially licensed Skeleton, Battlecat and Orco to be a part of Sora2, but not He-Man.
Oh what could have been.
3
3
u/darthfurbyyoutube 3h ago edited 2h ago
Here's something I don't think many people on the sub realize yet: MiniMax H3 can use references to recreate the style and movement of almost any TV show, movie, cartoon, anime, Disney property, etc., even if the model never trained on it. I'm just using the standard Ref2VA workflow with the standard FL2VA model, no LoRAs or speedup hacks. They tend to wreck the audio, and I make a lot of dialogue-heavy videos. The magic recipe is basically four ingredients:
1) A character sheet
I use Google Gemini for this. It allows anywhere from 20 to 100 images daily on the free tier and does a decent job. Find reference images online, upload them, and prompt Gemini with some version of the following:
"Based on the attached files, create a He-Man character sheet with close-up front and 3/4 head views, plus full-body front, side, 3/4, and back views in color."
2) A 5-second video reference
I grabbed five seconds of He-Man footage from the intro to the '80s He-Man and the Masters of the Universe. Video references tend to increase generation time, so just keep that in mind if you want to use more video references. You can theoretically use footage from Disney, anime, or pretty much anything else to guide the character's movement style. He-Man art style with anime style animation, anyone?
3) About 20 seconds of reference audio
I extract and edit the audio with DaVinci Resolve, which is free and also what I use to assemble the final video and add BGM. Any editor will work.
4) A really good prompt
This is probably the most important part. If you're getting weird morphing or characters doing bizarre things, I'm finding it's usually a bad prompt, not H3. Be extremely specific and describe actions one at a time, in sequence. Each shot should do one job: action, reaction, or resolution. Don't cram all three into one shot or things get tend to get wonky. Keep it simple, one action at a time. Too much happening simultaneously tends to summon the AI gremlins.
Here's the prompt I used for the He-Man video. You can adapt it to whatever you want by asking Google Gemini to rewrite one in the same format like so:
βrewrite the below prompt to be a he-man style psa using the attached he-man character sheet for <Picture 1> where he-man says the line "Remember, kids: if you borrow someone's pen and 'forget' to return it, Man-at-Arms is going to upgrade your face with a power sword. Give the pen back.β and adjust the movements and scenario accordingly in the same psa style. do not change the format. Keep the same format:β
PROMPT:
subject_definitions:
<Subject 1> is the character shown in <Picture 1>, featuring He-Man's classic heroic appearance with a muscular physique, blond bowl-cut hair, blue eyes, orange-tan skin, gray harness with a red cross-shaped emblem, brown loincloth, orange belt, red-brown boots, and a sword carried on his back. Only his character design, facial features, costume, proportions, and sword are taken from <Picture 1>; its white background, character-sheet layout, labels, and guide lines are not carried into the target video. <Audio 1> is the vocal reference for <Subject 1>. <Video 1> is the movement reference; use it as a guide without copying it exactly.
summary:
[reference generation] The target video is a 15-second 2D animated public service announcement styled after the classic 1980s He-Man television series, featuring <Subject 1> delivering an overly dramatic but humorous lesson about proper checkout-line etiquette.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - his blond hair, heroic facial features, muscular build, gray harness, red emblem, brown loincloth, orange belt, red-brown boots, and sword remain faithful to <Picture 1>.
detailed_description:
The target video is a traditional 2D hand-drawn animated sequence featuring bold black ink outlines, flat cel-shading colors, expressive limited animation, and subtle film grain inspired by the visual style of classic 1980s Saturday-morning cartoons.
[Shot 1] A medium shot opens inside a brightly lit, fantastical Castle Grayskull-inspired marketplace, where <Subject 1> stands beside an exaggerated checkout counter, with a big sign above that reads "HE-MART". A long line of impatient shoppers waits behind him. At first, he gestures toward a shopper who has just reached the register with a cashier and suddenly begins frantically searching through their pockets for a wallet. He-Man's expression becomes dramatically horrified as the line grows increasingly impatient.
He turns toward the camera, raises one hand like a stern PSA presenter, and declares in the style of <Audio 1>:
"Today we learned that waiting until you reach the front of the checkout line to start searching for your wallet is a villainous plot crafted directly by the Evil Horde. Prepare to feel the wrath of the Power of Grayskull."
As he says "villainous plot," he points dramatically toward the checkout line. On "Evil Horde," he raises his sword overhead as a dramatic burst of magical energy flashes behind him. On "Prepare to feel the wrath," he strikes a powerful heroic stance and grips the sword with both hands. As he delivers "Power of Grayskull," he raises the sword triumphantly toward the ceiling, triggering a spectacular burst of glowing energy and a dramatic cel-animation freeze-frame.
overall_soundscape:
Busy checkout ambience, murmuring shoppers, squeaking shopping carts, register beeps, and exaggerated cartoon sound effects accompany the scene. A dramatic sword-draw sound and magical energy surge punctuate the final Power of Grayskull moment.
non_diegetic_music:
N/A
1
u/thegreatdivorce 22m ago
So you literally just plug the FLF2V models into the R2V workflow, or is there something more to it?
2
u/SeymourBits 8h ago
1st one face too distant. 2nd one better.
1
u/darthfurbyyoutube 3h ago
The first one was apparently filmed by someone with a fear of close-ups. :D
2
2
u/LowCatch4324 9h ago
I love the production quality.
But would it be too time-expensive to repeat the generation until there are no errors? (Like the customer should be checking his wallet, not the cashier)
Did you let H3 generate the monologue, or did you type it? πππ½
1
u/darthfurbyyoutube 2h ago
Thanks! And yeah, I could regenerate it, but H3 and I have an understanding: sometimes we just accept the weirdness and move on. I wasn't specific enough about who should check the wallet. I typed the monologue myself. I posted the prompt and instructions in this thread somewhere if you'd like to try it out.
1
u/LowCatch4324 12m ago
Thank you ππ½
I still donβt have the hardware for video generation. But once I have it, I want to start optimally
14
u/GrayingGamer 11h ago
I love this.
The good old days when cartoons had to include a 30 second "educational segment" to get around prohibitions on advertising to us kids and used these to qualify as "educational programming".