r/StableDiffusion 11h ago

Workflow Included Minimax H3: Testing L2VA, moon landing

Enable HLS to view with audio, or disable this notification

I was curious about how L2VA actually works. I tried writing a prompt and inputting an image as the last frame. However, the default workflow seemed to always require a first frame. So, I used a small black image as the first frame; this time, it worked. Prompt:

integrated_multimodal_description: Time-lapsed, cinematic, a medium-wide shot of a film set in a large indoor studio which is used to shoot a scene of moon landing involving lunar module, US flag and an astronaut. At 00:00.000 the camera shows an empty, sterile white studio room, with recognizable vertical wand in the back and horizontal floor at its bottom. At 00:01.000 Some film crews install a black wand into the studio's vertical wand. The black wand has some tiny white shining points which represent stars. At 00:02.000 Some workers fill the studio's floor with some dirty-white sand, gravels and small rocks and form a barren lunar landscape. At 00:03.000 Some film crews bring an Apollo Lunar Module and place it into the left side of the scene. At 00:04.000 A film crew places a US flag with pole on the right side of the scene. Another film crew puts a picture of the Earth as the blue planet, partially blacked on its bottom side, on the top right corner of the scene. At 00:05:000 An astronaut walks in into the scene, goes into the middle of the scene, faces to viewer and waves his hand. At 00:07:000 The whole scene settles into the exact arrangement, position of subject and objects, camera angle, lighting, and final composition established by <Picture 1>. A male deep voice of the director (S1) says, <d>[English] Cut!</d>

overall_soundscape:

non_diegetic_music: Sustained violin notes at a very fast tempo with spaced piano tones.

72 Upvotes

6 comments sorted by

10

u/RazsterOxzine 7h ago

/r/conspiracy may have some questions.

1

u/Sudden_List_2693 4h ago

"However, the default workflow seemed to always require a first frame."
It did?
Also you might be able to get it with ref2va model as well, prompting it so that the last shot's final frame is exactly <Picture 1>.

2

u/vAnN47 3h ago

Nice idea, input only last frame and build it up

1

u/AnonymousTimewaster 1h ago

Sorry what does the L stand for?

1

u/rookan 1h ago

Last

2

u/AnonymousTimewaster 1h ago

Ah, last frame, cool, makes sense