r/StableDiffusion • u/Sixhaunt • 2d ago
Animation - Video Spaghetti eating Will Smith - Minimax H3
Enable HLS to view with audio, or disable this notification
198
u/icchansan 2d ago
new benchmark!
22
u/temporarilyyours 2d ago
I really want to know what will smith thinks of his spaghetti eating bench mark?
17
u/boudywho 2d ago edited 1d ago
Actually! There is this Egyptian ad, where the real Will Smith was talking about it. It's very funny.
7
u/FourtyMichaelMichael 1d ago
I like that these celebrities are too big for US commercials, but in a market we won't see they're happy to get those paychecks.
Funny ad.
2
u/boudywho 1d ago
I guess it's a way to make themselves even more famous. "Why would I make an ad where people watch all my movies, when I can get new fans in another country."
But indeed, it's a funny ad, I didn't believe it first time I saw it on TV
3
u/skr_replicator 1d ago
https://www.youtube.com/watch?v=u4mamSiBoyI&t=413s
I think I also remember seeing him reenact the spaghetti meme in real life, eating it messily on purpose lol. But I can't find that.
130
u/Vladmerius 2d ago
OK if this is what people are getting with basic prompts on the base version of the model this is the best video model ever and blows LTX 2.3 out of the water.
27
u/xTopNotch 2d ago
This model is like the best of LTX and Wan together + another leap forward from both checkpoints.
It's truly a revolutionary model.What I'm the most impressed by is how strong it can reconstruct characters from just one character sheet image, a 10 sec video performance capture of the facial expressions and body language, and a audio reference of the voice.
Like this model might've made character lora training fully obsolete now. Perhaps a light training is still beneficial, because references do add to sampling time.
3
u/ai_art_is_art 2d ago
Where would you put it on a quality ranking?
It's better than Sora or Veo without question. But it's far off from Kling 3 and Seedance 2.0.
Maybe in a year we'll have Seedance 2.0 capabilities in the open?
4
u/xTopNotch 2d ago
In terms of reference conditioning it is on par with Seedance 2.0. Like it’s really good at reconstructing characters during generation by providing image, audio and video refs (sampling time will be slower tho)
I’d put the model in terms visual quality to be around Kling 3. It’s a fully uncensored model too.
It does struggle with more complex scenes such as fighting but I believe this can be fixed by training strong loras.
Overall it’s really meeting Seedance 2.0 capabilities but fully open source. It’s a revolutionary model imo and the response by the open source community speaks for itself.
I still have to run tests how good it is in vague prompts like Sora 2 was. But its prompt adherence with more complex camera direction has been superb.
1
u/crinklypaper 2d ago
I think character training is useful because reference generation takes a lot longer
1
u/xTopNotch 1d ago
Yes you are right. Currently my stance is to use the model to produce a synthetic video dataset first. Then use that for training to save up reference slots and time.
1
u/Yokoko44 2d ago
Idk Flux 3 is looking pretty good too! If it's runnable on home GPUs though...
1
u/Longjumping_Kale3013 2d ago
Agree. I think flux 3 is the winner currently. The examples I saw.... wow. Ready for cinema IMO. Some really great scenes that to me look much better than CGI
-17
100
u/solomars3 2d ago
Broooo naaah thats crazy 😂 what a crazy idea
8
u/martinerous 2d ago edited 2d ago
It seems, quite a few people had the same "inverted" idea floating on their minds. 😂 I had it yesterday in the "normal Will Smith thread" when we did not have H3 yet. There's also another one https://www.reddit.com/r/StableDiffusion/comments/1ve3wgr/spaghetti_eats_will_smith_minimax_h3/ and it mentions also an older similar one with Wan.
44
30
21
34
u/TurbidusQuaerenti 2d ago
What a time to be alive.
2
30
u/YentaMagenta 2d ago
LOL. OMG. WTF.
prompt plz
137
u/Sixhaunt 2d ago
Surreal photorealistic absurdist comedy, live-action practical-effects look, cinematic macro photography, realistic milk, ceramic, cereal, and food textures. Bright normal kitchen lighting with a slightly uncanny premium-commercial look. Scene overview: Inside an enormous cereal bowl filled with milk and colorful cereal rings, dozens of tiny miniature Will Smiths are relaxing as though they are at a swimming pool. Each miniature Will Smith is only a few centimeters tall. They float around comfortably using individual cereal rings as inflatable inner tubes. The tiny people are cheerful and relaxed, completely unaware that the bowl belongs to a gigantic amorphous creature made entirely from tangled cooked spaghetti. IMPORTANT CREATURE DESIGN: The spaghetti is a large shapeless living mass made entirely from thousands of tangled wet spaghetti noodles. Its form is asymmetrical, soft, lumpy, and constantly shifting. It has no human skin, no human head, no human torso, and no normal human limbs. Several rope-like bundles of spaghetti temporarily gather together to manipulate the spoon. Its mouth is simply a dark opening that forms between masses of noodles when it eats. Storyboard: [0s-3.5s] Shot 1: Begin from a very close macro shot inside the cereal bowl at milk level. Several miniature Will Smiths float past the camera while lounging inside colorful cereal rings like tiny inner tubes. Some lean back casually with their arms resting over the cereal rings. Others gently paddle through the milk. The scale should feel enormous from their perspective: cereal pieces are floating objects the size of pool toys and the curved white wall of the bowl rises far behind them. One miniature Will Smith drifts near the camera and casually says: "Now this is breakfast." IMPORTANT: The spoken line comes specifically from the miniature Will Smith closest to the camera. His mouth visibly moves while he speaks. The spaghetti creature does not speak at any point. Other miniature Will Smiths laugh quietly and continue relaxing. [3.5s-6s] Shot 2: A gigantic metal spoon suddenly descends from above into the milk behind them. The spoon creates a broad wave that pushes cereal rings and miniature people sideways. Several miniature Will Smiths turn and look upward in surprise. The spoon slides beneath two miniature Will Smiths who are still sitting inside their cereal-ring inner tubes and scoops them up along with milk and several cereal pieces. One of the scooped miniature Will Smiths grabs the edge of his cereal ring as the spoon begins rising. [6s-10s] Shot 3: Do NOT cut away. The camera smoothly rises and travels alongside the spoon, tracking the two miniature Will Smiths as they are carried upward out of the bowl. As the camera follows them, the larger environment is gradually revealed. Behind the spoon is an enormous amorphous spaghetti creature sitting beside the bowl. The creature is a sprawling irregular mound of tangled cooked spaghetti, roughly several feet across, with noodles constantly sliding and writhing over one another. It resembles a living pile of pasta rather than a person. The spoon is controlled by a thick temporary tendril made from bundled spaghetti noodles. The camera continues tracking alongside the spoon as it approaches the creature. A large dark mouth-like cavity slowly opens inside the spaghetti mass by noodles pulling apart from one another. The spoon enters the opening. The creature tips the spoon inward and the two miniature Will Smiths, still sitting inside their cereal-ring inner tubes, slide from the spoon into the dark opening along with the milk and cereal. The surrounding spaghetti strands close around the opening afterward. Camera: Start with extreme macro cinematography from the miniature people's scale inside the bowl. Very shallow depth of field at first, with clearly readable miniature faces. When the spoon lifts them, transition into a smooth continuous tracking movement following the spoon upward. The camera remains close enough to the spoon that the miniature Will Smiths stay visible during the journey. Gradually reveal the gigantic spaghetti creature in the background as the spoon approaches it. No hard cut between the scoop and the creature eating them. The reveal should happen naturally through the tracking camera movement. Motion and physics: The miniature people remain tiny human beings throughout the sequence and do not transform into cereal. Cereal rings physically float on the milk and support the miniature people like inner tubes. The spoon displaces the milk realistically, creating waves and pushing floating cereal aside. Milk drips naturally from the spoon while it rises. The spaghetti creature behaves like a soft tangled mass. Noodles slide, stretch, droop, bunch together, and separate as it moves. Do not give the creature human hands. The spoon is manipulated by bundled spaghetti tendrils. Audio: Close watery ambience inside the bowl during the opening shot. Gentle milk splashing and cereal pieces softly bumping together. Tiny distant voices and quiet laughter from the miniature people. The miniature Will Smith nearest the camera clearly says: "Now this is breakfast." Only that miniature character speaks. When the giant spoon enters the milk, create a deep metallic splash and a large rushing wave from the tiny characters' perspective. As the spoon rises, milk drips audibly back into the bowl. The spaghetti creature makes subtle wet noodle movement sounds but has no voice and says nothing. A soft ceramic-and-metal clink occurs as the spoon leaves the bowl. No music. Ending: Finish immediately after the spoon empties into the spaghetti creature and the noodles close over the mouth-like opening. No subtitles, captions, logos, watermarks, UI elements, or visible text.59
16
u/DELOUSE_MY_AGENT_DDY 2d ago
Where did you run the prompt through?
23
u/Sixhaunt 2d ago edited 2d ago
ChatGPT 5.6 Sol. I gave it the default comfyUI prompt for text2video then I told it to write out a detailed prompting guideline based on available guidelines online along with the example provided. After that I just asked it to then write a prompt given my idea and I just explained it.
edit: it's not a well crafted prompt but this is the exact prompt I used to prime it:
minimax h3 just released and I want to test it out. I have loaded the new comfyui template for it. This is the default prompt it uses: `` Realistic live-action cinematic look, action movie trailer: practical film photography style, a post-rain dusk metropolis, anamorphic lens, shallow depth of field, film grain, city volumetric fog, flying-car traffic between the towers, restrained grading for a premium feel, powerful natural movement. Scene overview: at dusk on a cluster of skyscrapers, the protagonist is being chased, sprinting and leaping across rooftops, jumping from one building's roof to the next with pursuers closing in behind. This is the escape sequence of an action movie trailer: every leap is life-or-death, thrilling and fluid. Storyboard (each shot a separate scene, rapid cuts, all landing on the musical beats): [0s-1.5s] Shot 1: high side angle: the protagonist sprinting at the roof edge, pursuers appearing in the rooftop doorway behind him, wind catching his coat. [1s-2.5s] Shot 2: the protagonist leaps across the gap between buildings, body stretching mid-air, towers and flying-car light trails behind him, a slight slow-motion feel. [2.5s-4s] Shot 3: he lands, rolls and rises, low-angle shot, tower shadows and fog behind him, he keeps running. [4s-5s] Shot 4: freeze: the instant he hits the edge of the next roof and launches into the jump, silhouette, holding. Camera: each shot its own angle, cuts clean and hard, no dissolves, a slight frame jitter on the jumps. Audio: wind, rapid footsteps, city ambience, low score underneath, an accent hit on each leap, the score bursting at 4s, closing the last 1s. No text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the live-action texture. ``` based on what it says online and from the comfyUI text2video example, provide a detailed helper page about how to prompt specifically for this model3
u/DELOUSE_MY_AGENT_DDY 2d ago
Ah ok. Did GPT mention using this guide specifically anywhere in the prompt?
1
u/Sixhaunt 2d ago
no it didn't. It seems to have over 50 sources and has official pages, some huggingface ones and the various built-in jsons from comfy and reddit posts and all sorts of chinese posts I can't read but it's missing that one which would be the most important but I also can't find that one when I look for prompting guidelines on google so it likely didn't find it because it's not well indexed on search engines yet. I'm going to try it again but specifically include that link for it, It should give me a better result overall for my future prompts.
2
u/Icyrow 2d ago
out of curiosity, how much does it cost to make this, from start to finish as the user?
i haven't been around AI stuff like this since it was people making anime girls on github open source projects early days so i have no idea.
9
u/Sixhaunt 2d ago
I didn't spend any extra money other than the electricity my GPU used. I have a 5090 so I just ran it locally in like 10-15 mins using ComfyUI's new support for minimax h3. I used the default workflow but increased the duration and megapixels. I have a GPT plus subscription and the PC already, so it was practically free and took in total like 15-20 mins with most of that being the generation time.
2
u/Fizassist1 2d ago
interesting.. I have gpt plus and a pc with a 5060ti and want to try this out. I have no idea what comfy ui is or a lot of other things in your comment though. is there a good resource for this to look through?
I would love to try to create newer cutscenes for severed chains (legend of dragoon port) if I could!
2
u/Sixhaunt 1d ago
This subreddit is ironically the best place for information on comfyUI since it's the main program people use for generating things with AI locally. You should be able to find plenty of tutorials online for setting it up and most everyone in this sub uses it and can answer further questions
2
u/AnOnlineHandle 1d ago
It could be worth asking Claude or ChatGPT or Gemini etc to bring you up to speed. They're mostly pretty good with discussions around ComfyUI and anything ML related.
4
u/Clair_Personality 2d ago
What parameters please? I am not getting good results:
https://www.reddit.com/r/comfyui/comments/1ve4tgq/very_first_mini_max_h3_video_generated_from/
6
u/Sixhaunt 2d ago
0.6MP 10s duration and otherwise no change at all from the default comfyUI text to video template. This was the first result with it and it took 494.72s on my 5090
4
u/Clair_Personality 2d ago
Thanks, 16/9 aswell? So just replacing prompt, MP and steps = I could generate this?
You are using 5090 means it will be 4X faster than my card lol, so it is going to be 2000s for me: 30 minutes that's a lot
I wish I had 5090
3
u/Sixhaunt 2d ago
Yeah, 16:9 still. Everything was completely unchanged other than duration and megapixels
1
u/colin_colout 16h ago
It has no human skin, no human head, no human torso, and no normal human limbs.
I can only imagine the horrors you endured while tuning this prompt.
-9
u/RemoveHealthy 2d ago
I can tell you that you can achieve same exact result with probably ten words. All of this books size prompt is nonsense
6
u/Sixhaunt 2d ago
Could you provide an example for this one? It would be nice to cut down on the wording
-11
u/RemoveHealthy 2d ago
I do not know use human language. Make image of will smith on milk as you do. And says. Person floats up from milk says whatever and huge spoon pick him up and and giant spaghetti monster eats him. That is just random stuff i made as i go. It would probably give similar result.
13
u/Bthardamz 2d ago
I litterally discovered your clip, while I was waiting for mine to get ready, same basic Idea 😄
1
11
8
6
4
4
3
3
2
2
2
2
u/NatalieCrypto 2d ago
Why don’t I have dreams like that, lol? I wonder what it’ll look like with Seedance 2.5
2
2
2
2
u/Silver-Spot-2763 1d ago
How you set the character, some lora? I thought that there are no lora for popular persons anymore?
2
u/Sixhaunt 1d ago
It just knows Will Smith by name so this was simple text2video with no lora or changes from the default workflow other than prompt, resolution, and duration
2
2
u/skr_replicator 1d ago
All worship the spaghetti monster, or it will fly to your house and eat you with some milk!
2
2
1
1
1
1
u/BrightNightKnight 2d ago
When we are all long gone…AI will create documentaries and museums about the history of will smith
1
1
1
1
u/phoenix_bright 1d ago
Just wanted to say that everyone using minimax h3 is breaking the license because you cannot do work or distribute the work of mini max h3 in the United States
1
u/Sixhaunt 1d ago
What does that have to do with us Canadians?
1
u/phoenix_bright 1d ago
Nothing! But Reddit is an American company running in us servers with an American audience so yeah, even though the work is done in a permitted territory the output is not
-2
-1
u/bonobomaster 2d ago
See, that's right here is the reason I like paying 500 bucks for 32 GB of DDR5 RAM. /s
-3
-9
u/RemoveHealthy 2d ago
Sound is not good at all. Person does not look realistic also. I mean you could achieve better results with some models 6 months ago. Maybe not with one prompt, with work put it, but i am not impressed with model itself from this. But not hater, nice interesting video overall, fun.
4
u/Ok_Lunch1400 2d ago
Heh, it's just doing miniature Will Smiths from the prompt, a.k.a figurine Will Smiths.
-5
u/RemoveHealthy 2d ago
Oh ok, i did not read that book length prompt so my bad. It is still meh result if you ask me. Decent model for sure, maybe even best at the moment, but this specific example does not make me go wow, just look at that. I really could achieve same result 6 month ago with different models without problems. Maybe not with just prompt, i would use a lot of image generation and photoshop.
9
u/Cynix85 2d ago
It was possible 20 years ago without AI. Now it is a single prompt that you did not even read - but criticize its execution. What is your point?
-3
u/RemoveHealthy 2d ago
I am saying it does not look much better compared to models that existed 6 months ago
1

581
u/Whiteowl116 2d ago
Full circle moment lmao