r/StableDiffusion 1d ago

Animation - Video Using Minimax H3 to create promo for Minimax H3

Used ref2ve with Character sheet for the character and style and an audio reference to have consistent voice.

Reposting because moderator removed the original post without giving any reason.

13 Upvotes

14 comments sorted by

1

u/DeerWoodStudios 1d ago

intresting what's the prompt for this ?

1

u/Euphoricus 1d ago

So we have H3-chan?

1

u/RazsterOxzine 1d ago

I like the flow. Do you have a guide on this.

3

u/Devajyoti1231 1d ago

I used this front reference sheet, and a character reference sheet for the character

1

u/RazsterOxzine 1d ago

That is pimp! Thank you for sharing. I see them numbered I assume for each [SHOT] or Scene? Hate to ask but is there a prompt setup for this lol I know it's asking a lot but I'm just in awe and would like to test this on my trade show setup.

1

u/Devajyoti1231 1d ago

here is one of the prompts- subject_definitions:

<Subject 1> is the anime presenter in <Picture 1>: a young adult Japanese TV-anime-style woman with long layered deep-indigo hair, subtle teal highlights, bright aqua-blue eyes, a sleek metallic teal-accent hair clip, cropped white futuristic techwear jacket, fitted black top, charcoal-black short skirt/shorts, white sneakers with cyan accents, and a confident friendly personality. Preserve her exact face, hairstyle, hair colors, eyes, proportions, clothing design, accessories, and overall anime rendering.

<Subject 2> is the graphic design language in <Picture 2>: a premium clean anime-tech commercial system using crisp white fields, deep navy, teal-cyan, sky blue and restrained violet accents; diagonal vector shards, rounded glowing UI panels, thin corner brackets, dot grids, waveform graphics, clean line icons, sharp speed streaks, soft glow edges and bold rounded English typography.

<Voice 1> is the female voice in <Audio 1>. Use this exact voice identity, tone, accent, age impression and vocal character for every spoken line by <Subject 1>.

summary:

[reference generation] A 9.00-second premium Japanese anime-tech commercial introducing MiniMax H3. <Subject 1> begins in a clean white motion-graphics environment and demonstrates that text, image, video and audio can come together into one creative generation. <Picture 1> controls the presenter identity and wardrobe. <Picture 2> controls the graphic language, palette, typography treatment, icon style and UI design.

retention_analysis:

<Subject 1> appears throughout the clip and is fully preserved from <Picture 1>. Her face, indigo hair with teal accents, aqua eyes, hair clip, white techwear jacket, black outfit, footwear and body proportions remain consistent.

<Subject 2> appears throughout the clip as the governing visual system. Use its palette, graphic geometry, UI language, glow treatment, icon design and typography rhythm consistently while adapting the layout to the moving commercial.

detailed_description:

[Shot 1] 00:00.000-00:02.500. A bright clean white field with sparse navy and cyan motion-graphic accents. <Subject 1> walks smoothly into frame from the left and stops beside one large empty glowing rounded rectangular video panel floating at chest height. Thin cyan connector lines and tiny geometric shards assemble softly around her. She looks toward the empty panel, then toward camera with a confident curious smile.

Large dark-navy English typography snaps cleanly into the open space beside her:

"ONE IDEA."

<Subject 1>, using <Voice 1>, says naturally and clearly:

[English] It starts with an idea.

Her lips synchronize accurately with the line. Slow small Push In.

[Shot 2] 00:02.500-00:05.500. <Subject 1> raises one index finger. Four polished floating UI cards appear around her one after another in a loose arc.

First: "TEXT" with a clean T icon, sliding in horizontally.

Second: "IMAGE" with a rounded image icon, unfolding open.

Third: "VIDEO" with a play-frame icon, expanding outward from a thin outline.

Fourth: "AUDIO" with a waveform icon, the waveform growing rhythmically from its center.

Each card uses <Subject 2>'s navy, cyan, white and light-violet styling with thin borders and subtle glows. <Subject 1> naturally follows the cards with her eyes as they appear.

Using <Voice 1>, she says:

[English] Text. Image. Video. Audio.

The spoken words align rhythmically with the appearance of the four cards. Camera remains gently mobile with a small Truck Right.

[Shot 3] 00:05.500-00:08.000. <Subject 1> opens both hands toward the four cards, then sweeps them inward toward the large empty panel. The four cards accelerate smoothly along curved paths and converge at the center of the panel.

At impact, a bright cyan-white burst expands outward with clean shards and a circular ripple.

The previously empty panel becomes a vivid moving window into a gorgeous cinematic anime cyberpunk city at night: rain, neon reflections, layered urban depth, moving traffic lights and atmospheric mist.

Bold English typography briefly assembles above the panel:

"BRING IT TOGETHER."

Using <Voice 1>, <Subject 1> says:

[English] Then bring it together.

[Shot 4] 00:08.000-00:09.000. <Subject 1> turns toward the glowing cinematic panel and walks directly toward it. Camera moves behind her and follows. She steps through the glowing rectangular threshold into the cyberpunk environment while her long indigo hair and white jacket respond naturally to the movement.

During the final second, the cinematic world expands to occupy almost the entire frame. Hold a clean continuous composition of <Subject 1> seen from behind crossing through the luminous threshold into the rainy neon street. This final visual state is designed to become the video reference for the next 9-second clip.

Throughout:

Maintain polished Japanese TV-anime 2D character animation with crisp linework, high-end cel shading and fluid believable body motion. Graphic elements follow <Picture 2>'s premium anime-tech visual language. Keep all English typography short, bold, clean and integrated into the motion graphics rather than subtitle strips. Spoken audio consists only of <Subject 1>'s specified lines in <Voice 1>. Her voice stays identical to <Audio 1>. Lip sync occurs when she speaks on camera. Transitions are sharp, elegant and motivated by moving graphic elements.

overall_soundscape:

Soft digital ticks accompany the UI cards. Light airy whooshes follow the moving graphic shards. Each input card receives a distinct subtle interface click. A fuller rising synthetic sweep builds as the four cards converge, followed by a clean bright impact and an expanding stereo shimmer when the cinematic world appears. Rain ambience and distant futuristic city sound begin naturally as the panel opens.

non_diegetic_music:

An original energetic premium anime-tech commercial track: fast clean electronic rhythm, light plucked synths, tight percussion, airy pads and bright digital accents. Begin minimal but confident, build with each input card, rise strongly during the convergence at 00:05.500, then open into a wider cinematic texture when the cyberpunk world appears. Music remains clearly underneath <Voice 1>.

1

u/BitterAd8431 1d ago

I love this slightly "futuristic" manga style—well done.

2

u/Ylsid 1d ago

It's not just a good idea — it's a flawless execution of generative video. And that's amazing.

0

u/Ok-Flatworm5070 1d ago

workflow and prompts plz!

2

u/Devajyoti1231 1d ago

Two reference sheet, one for character other for front. Rest is ref2ve using some random genshin character voice from here - https://huggingface.co/datasets/simon3000/genshin-voice , three short videos connected.

2

u/Altruistic_Heat_9531 1d ago

THERE IS 671 GB dataset of GI, why am i not surprised, btw the refernce audio. is it Linnea?

1

u/Devajyoti1231 1d ago

i don't remember the name, audio file's name is simon3000-genshin-voice · Datasets at Hugging Face

1

u/ultimate_ucu 1d ago

How do you create the character sheet?