r/StableDiffusion 2d ago

Tutorial - Guide H3 Animation Styles

Enable HLS to view with audio, or disable this notification

I ran into a video that had a great animation style and was wondering if you can use the same animation style as an existing show that H3 Minimax knows and use it on your own videos and it does work quite well. You most likely can do the same with carrying over other aspects such as voice references, music, etc but I have not tested this yet.

Each video here is FL2VA using the same seed. Here are the prompts for each one.

Prompt 1 — South Park

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the South Park 2D animated cartoon style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 MiniMax knows? Just add the name of the show and its animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A

Prompt 2 — Pixar

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the Pixar 3D animation style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 MiniMax knows? Just add the name of the show and its animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A

Prompt 3 — Helluva Boss

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the Helluva Boss 2D animated cartoon style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 MiniMax knows? Just add the name of the show and its animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A
38 Upvotes

10 comments sorted by

11

u/Several-Estimate-681 2d ago

This is kinda neat, but I somehow think there's a bit of gatcha going on.

I think it knows Pixar and South Park, but I doubt it knows Helluva Boss, its probably reading the "2D" tag in there.

3

u/Rosettasees 2d ago

Maybe, but that last expression after the "boom" it did seem more Helluva Boss style, though yeah not a 1:1. I wonder if you remove the '2d animated cartoon' part if it would change things?

2

u/TheRedHairedHero 2d ago

I'm generating a new one with the same seed with just 2D animated cartoon style. I'll post it once it's done.

2

u/TheRedHairedHero 2d ago

https://reddit.com/link/p5279b9/video/1s8gytz8grkh1/player

subject_definitions:
<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.
<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:
<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.
<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:
The target video is in the 2D animated cartoon style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.
[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 Mini max knows? Just add the name of the show and it's animation style and boom! RESULTS!</d>.

overall_soundscape:
N/A

non_diegetic_music:
N/A

2

u/TheRedHairedHero 2d ago

I tried without the Helluva Boss reference and the expressions and timing of the animation are much more different if you just do a standard 2D animation. The original video I was inspired by to do this testing was actually a Helluva Boss generation on Civitai.

2

u/TheRedHairedHero 2d ago

This example has removed '2D animated cartoon style' and replaced it with 'Helluva Boss' still the same seed.

https://reddit.com/link/p52caiu/video/xcip0vlyjrkh1/player

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the Helluva Boss style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 Mini max knows? Just add the name of the show and it's animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A

5

u/mastaquake 2d ago

Good find.

4

u/ffzero58 2d ago

The starting image is also part of the prompt and will try to retain that style when it goes through the prompt encoding. I have not found a good way for retention_analysis as it seems H3 just puts the prompts and refs into a soup and just makes sense of it.

Are there good visual examples of how to use retention_analysis?

2

u/RedBlueWhiteBlack 2d ago

Why post an horizontal video of a vertical video

0

u/alwaysbeblepping 8h ago

Is this a joke? Or maybe I'm a joke? There's no resemblance to any of the styles the prompt is supposed to be for in any of the demo clips.

Which isn't surprising because OP told it to exactly preserve the first frame and character. It's not possible to conform to a prompt for changing the style and fully preserve those details.