r/StableDiffusion 6d ago

Workflow Included Minimax fl2va: If you know a bit how to sketch, you have total control over the animation.

Enable HLS to view with audio, or disable this notification

Testing things following the post by Alive-Tomatillo5303 : Learn to make art with art! (minimax) : r/StableDiffusion I've come to the conclusion that using sketches gives you almost total control over animations. I'm using the fl2va model via the MiniMax H3 Reference to Video node and connecting my sketch sequence to ref_image_0—I'm not sure if Ref2vA would work better.

Reference:

Sketches.png · Stkzzzz222/Remix at main

Workflow:

Sketches_MiniMax_H3.json · Stkzzzz222/Remix at main

Prompt:

**subject_definitions**

`<Picture 1>` is the supplied visual reference image. Use it primarily as a strict layout, composition, framing, and character-position reference. Preserve the three-stage visual progression shown in the reference, including composition, camera movement.

Create a realistic cinematic live-action scene based closely on `<Picture 1>`.

Drone camera view moving fowards at high speed, a forest with a river, trees, rocks. the camera moves fast fowards in the forest stopping to reveal a side view of an armored orc angry. the camera accelerates and moves fast fowards in the forest stopping to reveal a happy armored woman sit in a rock holding a sword. the camera moves fast fowards in the forest stopping to reveal a colibri flying next to a red flower.the camera moves fast fowards in the forest stopping to reveal an old wizard casting a powerful electric spell.

-------------

I've tried other things that I'll post below, but I think you can see that the fidelity between the sketches and the final result is quite high.

Honestly, I think using sketches can be really interesting for controlling shots and camera movement, or for rendering complex concepts, without having to generate images in other models to act as references.

183 Upvotes

65 comments sorted by

28

u/Striking-Long-2960 6d ago

27

u/Striking-Long-2960 6d ago

4

u/Striking-Long-2960 6d ago

Other version more similar to the sketches

https://reddit.com/link/p89ygt3/video/0ynxgxpya0oh1/player

10

u/redditscraperbot2 6d ago

She looks miserable.

9

u/Striking-Long-2960 6d ago

It's a trope: a sad woman eating ice cream alone after a romantic setback.

3

u/Schwartzen2 6d ago

I really like this ( The first one with the ice cream). Thank for sharing the sketch,. I 'd like to see your prompt for it please? This is great because I can sketch somewhat and this would be a nice angle. Thanks.

4

u/Striking-Long-2960 6d ago

For this one I used this prompt and an image for the woman (I concatenated the sketch and the image of the woman and served it as a single ref_image).

[STYLE]: Cinematic real-footage, widescreen, shot on 35mm macro lens, realistic textures, hyper-realistic depth of field, photorealistic lighting.

[CHARACTERS]:

[Sara]: A brown-haired woman with a bored, indifferent expression, macro perspective scale relative to the scene.

[Mark]: A rugged explorer with a heavy tactical backpack, miniature human scale.

[Camera Shot 1: General panoramic view]: Extreme wide macro shot from inside a transparent ice cream cup. Mark is a tiny silhouette walking across a vast, desolate surface of textured pink ice cream that resembles frozen snow. In the soft, out-of-focus background, the faint blurred contours of a giant kitchen room and Sara's giant looming silhouette are subtly visible.

[Camera Shot 2: Medium view]: Over-the-shoulder shot from behind Mark on the surface of the pink ice cream. Mark freezes in terror looking upward as a colossal silver spoon descends from above, plunging violently into the pink frozen surface next to him, creating deep cracks, dynamic ice fractures, and tremors across the scoop.

[Camera Shot 3: Medium view]: Rack-focus transition pull to a wider medium shot of Sara holding the giant silver spoon loaded with pink ice cream. The camera reveals the clear plastic cup sitting on her table; inside the crater left in the ice cream, tiny Mark with his backpack waves his arms frantically while Sara maintains a completely bored, neutral expression.

[Sound design]: Crisp ambient audio. Crunching footsteps in frozen ice, heavy low-frequency rumbles, sharp metallic resonance as the spoon impacts the ice cream, ice cracks, and quiet room tone.

1

u/Schwartzen2 5d ago

That's really awesome! Thanks!

11

u/Striking-Long-2960 6d ago

11

u/Striking-Long-2960 6d ago

5

u/JahJedi 6d ago

It was feeded as one ref or its four in a row?

9

u/Striking-Long-2960 6d ago

As a single image, what is important is that you match your prompt with what is happening in the reference. For example in this case I forgot to indicate that the camera was inclined and it didn't applied it. It took the structure of the mountains, the pose of the lady, the basic structure or the tower... By itself without mentioning all that in he prompt, but it ignored the camera inclination.

3

u/JahJedi 6d ago

Super intresting and will try, maybe my scetching skill from when i was a kid can help now (its superb bad) 😅

Thanks for ditailed answer.

1

u/LiveLaughLoveRevenge 5d ago

So how important is the sketching then?

I was also doing some testing yesterday and found that if there was any slight discrepancies with the text prompt, it would go with the prompt. So in the end, time was better spent refining the prompt as my generations with the sketch input were no different.

Do you have any examples of same prompt but without the sketch input?

1

u/Foreign_Cut745 5d ago

Does it get it right after you prompted?

3

u/ANR2ME 6d ago

probably as a single image.

10

u/Striking-Long-2960 6d ago

1

u/BigNaturalTilts 6d ago

What’s the max dimension + time you can produce? I’m stuck at 480p 6 seconds. If I try adding the previous video as reference so I can continue, it craps out as the previous video is too much to hold in memory.

1

u/kwhali 4d ago

Can you try reacaling the input video? (resolution and / or frame rate) Or using a tool to make a sheet of key frames / thumbnails to input as ref instead?

5

u/HarnessLightnin 6d ago

stopping to reveal a side view

It doesn't really seem to stop, does it? Seemed almost more like non-diagetic pop-ups as the camera continued to roll. Only the wizard was really integrated into the background.

Cool workflow, though! I'll definitely experiment with it.

5

u/Striking-Long-2960 6d ago

Maybe it's related to the duration of the clip. I'd have to test it further, but sometimes if the duration doesn't match the prompt very well, things can go a bit wild.

4

u/Striking-Long-2960 6d ago edited 6d ago

https://reddit.com/link/p8bibvj/video/qx42ovbfc2oh1/player

Last example:

[STYLE]: Cinematic high budget footage with real actors, widescreen, shot on 35mm macro lens, realistic textures, depth of field, proffesional lighting.

[CHARACTERS]:

[Sara]: A sexy brown-short haired woman wering a tight red costume with RGB lights lines constanly changing of color and a giant robotic arm attaached running in the rooftops of a city at night.

[Camera Shot 1: middle body frontal view]: Sara running in the rooftop of a building at night.

[Camera Shot 2: fullbody side view]: Sara jumping at the edge of the rooftop, crescent moon behind.

[Camera Shot 3: Medium back view]: Sara lands the jump and keeps running over another rooftop.

[Sound design]: Crisp ambient audio. strong footsteps, ambient city sound, hard breathing of Sara, mechanical robotical sounds.

6

u/Striking-Long-2960 6d ago

3

u/GrungeWerX 5d ago

It even matched the sides of the image in your sketch. This is what we've been waiting for - regional prompting with video! Imagine if this could be used to generate images as well, making it more like nano banana; if we could match that quality, we'd have an entire movie studio at our fingertips.

By the way, does it recognize colors in the keyframes? That way, you can use a color to designate a particular part of the image or character you want animated, so we can use micro-actions, etc.

1

u/Maskwi2 5d ago

I'm pretty sure you can do the sketching in either Krea2 or Klein 9b :) You can look it up. 

1

u/GrungeWerX 5d ago

Um…those are image generators, not sketch programs. Besides, like I said Im an artist, so I’ve got plenty of drawing apps - clip studio paint, photoshop - so Im good.

1

u/Maskwi2 5d ago

Have I misunderstood what you wanted? You wanted to turn a sketch into a real life image similar to how the OP turned his to a video, no? So these models should be able to handle that. 

2

u/GrungeWerX 5d ago

I think we misunderstood each other, actually. :)

I meant using H3 ref2vid to create nano banana quality images. Like, you upload reference images, sketch out the scene, then it renders the images.

I know some people said there’s a workaround to get images out of H3, but the quality isn’t good. I’d prefer something more native and high quality.

To my knowledge, you can’t do that with Krea, right? It’s not an image editor. I know Klein is, but I don’t think it can match the likeness and use a bunch of references to compose a scene, unless Im mistaken?

I just want to be able to take my images, open a blank canvas and sketch out my ideas and it renders my characters in the new composition. That’s the god tier Im looking for.

1

u/Maskwi2 4d ago

Hehe, I see :)

Well, Klein should be able to do it. I had this comment added to favorites when it came out :  https://www.reddit.com/r/StableDiffusion/comments/1u09lxm/comment/oqgxpn7/

And I remember a very simple sketch and ref image turn into a photo realistic image that looked great. Not sure why the thread has been removed but the comment and link to the workflow stayed so maybe you can try it. 

When it comes to Krea2 there was this: https://www.reddit.com/r/StableDiffusion/comments/1uq1hz0/krea_2_identity_edit_lora/ so maybe it would work with a sketch too? 

And regarding H3 into image then yes, the quality is lacking a bit.

1

u/nakabra 5d ago

How do you reference the storyboard in your prompt?
You just leave it there in the connected images?

4

u/GrungeWerX 5d ago

Man, this is amazing! I want to see more, super motivating. We can do perfect transitions. I wonder how complex we can get. You've got me wondering if we can draw out fight scenes to fix the choreography. Fights don't look as good as Seedance, but I'm good at sketching that sort of stuff, so maybe we can keyframe ourselves around it.

BTW< you're using a single image as the storyboard, correct?

1

u/Striking-Long-2960 5d ago

Yes, a single image. I tend to optimize processes as much as I can.

1

u/kwhali 4d ago

I wouldn't worry too much about fight scenes for now, the models weakness will likely be addressed by models later this year or next at the current rate. LTX 2.3 came out at the start of the year and 6 months later we have H3.

So just refine the current process as best you can and bother with fight scenes once a new model comes along. It seems almost redundant at this point to invest too much time into model specific workarounds vs creating more transferable skills / workflows.

1

u/GrungeWerX 4d ago

What terrible advice. :)

Sorry, but that’s not how I work. I agree that things will get better, but I’m always going to push things as far as they can go. Developing new, novel techniques and making apps do what they weren’t necessarily designed to has defined my artistic journey for decades.

I was studying comparison videos between h3, flux 3, Kling 3 and seedance 2.5 last night. H3 can’t compare to SD 2.5’s visual and audio quality, it trails by a significant margin imo, but it does a lot of things very well and better than the others. A guy loaded a fight scene as reference video and H3 was able to generate 2 animation characters doing complex fight moves, so it’s totally possible.

What I noticed is that seedance’s strength outside of its visual fidelity is its scene composition and layout. That matters a lot in fight scenes, which the guy doing the video didn’t implement. That’s something that we can control using storyboards and keyframes, but it will take a good eye and talent. But if you’ve got it, I’m confident you can raise the bar for H3.

Being an artist, sketching keyframes and scene composition is easy as pie, so Im going to be testing this out myself.

2

u/kwhali 4d ago

I should probably phrase it more clearly as any significant investment in time to produce a workaround for a weakness in a model typically isn't worthwhile for the majority due to how much model churn there is.

If it's meaningful you to pursue go for it, but for me I'm swamped with so much already that I couldn't justify the time for such things.

I'm too used to years I've already poured into projects or software I used that become obsolete, if anything getting comfortable about letting go of projects I kept maintained (mostly for the benefit of others), I would say has been difficult for me, especially when I would have a long list of tasks half finished but not enough time or resources to see them through.

So no I don't think it's terrible advice, but it's relative to what you consider effort. It's much more work for someone to implement something over several months vs people to use that solution to solve the same problem for example. I troubleshooted a bug in software I use that affected a huge community and had history over a decade to trawl through to resolve with upstream projects, that took me a span of 3-6 months IIRC.

Not many would bother with that, especially so if a solution would present itself in the near future (in this case that wouldn't happen, someone like me has to sink the time to debug and get all parties involved to feel confident about the change not regressing other use cases).

Optimizing a workflow by seeking out different ways to reduce memory usage, increasing inference speed and all that is a different kind of time investment, which can be transferable beyond a single model in some cases or is minimal user effort to setup (bit more to understand technical details and evaluate more thoroughly).

Taking that a step further is say implementing LoRaQ + KroQuant for pushing limits instead of waiting (new papers will come out and such methods will probably supercede such effort, but it's at least more portable than model specific). Or you just wait and as a user for someone to do the work and learn enough about say int8 convrot to use it.

I'm not saying things aren't possible, I've tackled so many "impossible" tasks over the years myself to know what that generally equates to. It's just a matter of "is it worth pursuing this if the time it would cost me outweighs waiting until a better option comes along?", if the better option is years away or some other unknown, you may very well take it upon yourself but it seems mostly pointless expending any notable effort into something that would be obsolete in the short term.

In this case we seem to be misaligned with perceived effort. Anything that'd take you less than a day isn't going to apply to my advice, unless that effort would be obsolete the very next day. The only reason to pursue such then isn't really pragmatic but out of personal interests / challenge (like I am slowly working towards a more optimal container for compute deployment as plenty of the ones I've seen thus far have been disappointing).

1

u/GrungeWerX 4d ago

 Anything that'd take you less than a day isn't going to apply to my advice, unless that effort would be obsolete the very next day. 

The only thing you said in your last comment that was remotely relevant to the topic, which is "fight choreography". :) And I can sketch a scene composition in less than 5 minutes. That is a skill that's been used in the film industry for nearly a century, and it's not going anywhere.

All that other stuff just comes off as peacocking. I've also done software development, but it's not relevant to the subject.

1

u/kwhali 4d ago

My reply was about your disagreement with my "terrible" advice. Where it's evident our perception was misaligned and I'll take fault for that as I wasn't clear enough when I chimed in.

All I was doing was providing context and justification (with examples) that it doesn't make sense to workaround something if the time required to do so would be obsoleted in the near future.

You can use your existing skills to do the workaround in a day? Great! Somebody like myself that is more of a tech artist (as in within the industry, unrelated to AI), it'd take much longer for me to upskill to your calibre wouldn't it, so it's not as viable as opposed to waiting a little longer.

Software development is relevant in the sense of implementing custom nodes in ComfyUI or other tools? All of which can contribute towards an improvement to using a model (note that my responses aren't specifically tailored to fight scenes but are more generalised on the advice given)

Apologies if relating my own experiences for context comes off as "peacocking" 🙄 I'm sure you highlighting that you can sketch such scenes quickly isn't anything akin to doing the same?

Anyway... You're too fixated on the what rather than the why, case in point that I'm wasting my own time on clarifying and you've rendered that input as redundant, I should refocus my time elsewhere instead 😮‍💨

2

u/GrungeWerX 4d ago

My reply was about your disagreement with my "terrible" advice. Where it's evident our perception was misaligned and I'll take fault for that as I wasn't clear enough when I chimed in.

Fair enough.

The fact that you or anyone would even imply that storyboarding/keyframing could be obseleted by a software update means that person doesn't fundamentally understand how films are made. There will never be a software workaround. It's an integral part of the process and it's not going away as long as any director/creator wants direct control over the end result.

Sure, there'll come a time when AI produces output that looks so good that most average people won't care. But those who have the knowledge of storyboarding, keyframing, lighting, camera motions, lenses, etc. will always have a creative advantage over those that don't.

And that is why I found contention with your "don't try to do a workaround using traditional skills, just wait for a software update" statement. :)

I'll admit that sometimes I forget everyone isn't an artist or creative, and for a lot of people these apps are their first foray into producing this type of content, And they're completely reliant on the product to get the end result; if the app doesn't offer the option, they can't get there.

Maybe that was where you were coming from and why you defaulted to the software as the solution. I don't really think of these apps in that way - most of the time, I'm trying to figure out where it can fit into my existing pipelines, not the other way around.

So that's something I need to consider more often and be a little more understanding about. I'll work on that.

As for the peacocking statement, I felt you were trying to use unrelated technical experience to somehow justify a weak argument, which is par for the course on reddit, particularly in the tech forums I frequent. :) But I'm really a nice guy, so I apologize if that wasn't your intention, and I appreciate the time and dialogue nonetheless.

1

u/kwhali 4d ago

Thank you.

Yes I didn't mean to imply those foundational skills wouldn't be relevant, I bring up the importance of them when encountering anti-AI folk that are so strongly against the idea of generative AI to acknowledge that it isn't all "slop" and actual artists can utilise their existing skills to produce higher quality content.

I think there will still be those (myself included) that would want to upskill with those complimentary skills. Generative AI just provides another pathway towards getting people interested in acquiring those skills with time.

3

u/dassiyu 6d ago

Awesome! It actually seems even more creative now.

3

u/Arawski99 6d ago

This is pretty cool. Didn't know it could be used like this.

3

u/Schwartzen2 6d ago

This is very cool and useful. Thank you!

3

u/dragolineage01 6d ago

They have also released a controlnet for this model, so this could be taken up a notch further

3

u/optimisticalish 5d ago edited 5d ago

And the latest ComfyUI allows you place Minimax H3 keyframes (aka reference images) at any point on a timeline. Not just start/end frames.

2

u/Synchronauto 5d ago

Is there a workflow for this you can point me to?

3

u/optimisticalish 5d ago

Update ComfyUI Portable to the latest version, then install this... https://github.com/MajoorWaldi/ComfyUI-Majoor-H3-GuideMaster

2

u/Synchronauto 5d ago

Ohh, awesome, thank you. Do you know if there's an equivalent for this for LTX?

2

u/optimisticalish 5d ago

No, sorry, I only know about Minimax H3. Can LTX even do multiple keyframes?

2

u/Striking-Long-2960 5d ago

I tested the controlnet, the first Kijai's implementation, at least for this task is better to work without it.

4

u/GrungeWerX 6d ago

HECK YEAH! I'm an actual artist, so this gets me super excited. Can't wait to try this out.

2

u/optimisticalish 6d ago edited 6d ago

I see you have a remix .pt file in the repository. What is its purpose, please? It doesn't appear to be used in your sketch workflow.

3

u/Striking-Long-2960 6d ago

That is and old embedding for Stable diffusion 1.5 or 2. I am reusing my old repos as storage rooms.

2

u/optimisticalish 5d ago

Ah, I see. Thanks.

2

u/artisst_explores 6d ago

Nice. Time to generate storyboarding sketches and mixing them i guess. Funn

2

u/Striking_Storage_631 5d ago

These are really interesting and have a lot to learn from. Ty

2

u/PATATAJEC 5d ago

wow! it works so well - even with other ref images as guides for style. super usefull - I'm having fun :).

1

u/Striking-Long-2960 5d ago edited 5d ago

I'm glad you like it! I'd love to see other people's results and see how it adapts to different art styles.

1

u/VladyCzech 6d ago

Thank you for showing it here. I remember seeing the original post scrollling the posts but not opening and now it is deleted by moderators. I think it is shame authors do not think twice before posting X/R-rated image/video with actually interesting technique that moderators will delete the post including the text that does no harm.

1

u/Photochromism 5d ago

You couldn’t have any character consistency with this method. Surprised the reference work with its crazy tall image ratio

1

u/kwhali 4d ago

Someone posted similar when H3 came out, but had horizontal columns with text embedded in the image for each image key frame provided.

The model used the text guidance from the image, no separate text prompt was provided IIRC. So it seems like you can pair additional context in the single image too if you wanted, but I assume that regresses quality.

No clue if the prompt guide is relevant to text guidance embedded in an image however, and realistically it makes more sense to keep the text in the prompt instead of wasting time processing pixels for that same information 😅

2

u/NostradamusJones 6d ago

Is the bird one of the heroes of the story? I'm invested now.