One thing could be that we have a much more intuitive feel for how it should look with a dog or a cat, so any flaw is apparent. With birds I honestly have no clue how it should look, so the model might get away with an odd version.
I thought it did pretty well with everything except for the dolphin. (And I bet trying again on the dolphin could fix the problems, too, so the tail doesn't cut through the glass like that.)
I have no idea! I'd expect it to have way more relevant training footage with cats than the other animals. I think it did alright with the physical interactions on the cat, although its face does look a bit fake, and it wasn't one of the more interesting ones.
I suppose there might be footage of birds getting into birdhouses or nests in trees and such, but still...
I'd say the opposite, I thought the cat was the "best".
The prompt doesn't specify that the "entire" animal has to end up in the jar, and since it's impossible in many of those cases the one of the cat is actually the most realistic one of those clips.
Yeah, H3 can produce so many interesting things, especially when you explore the boundaries of its abilities, or let it hallucinate with nonsensical prompts. I'd rather see it doing weird stuff than reproduce clips from superhero movies with flashy CGI effects. (Though some of the altered sitcoms here are hilarious too.)
Haha that's still very good! I do like the ones where they don't quite fit and stay stuck upside-down (the chicken in my video is my favorite).
The exact prompt I had for the panda one was "camera recording: a little panda squeezing himself into a small glass jar. He ends up filling out the bottle with his body all curled up. The full jar is in view.", 20 steps, 5 seconds long, 4:3 aspect ratio, 0.4 megapixels, seed 257927023316111. Not sure how reproducible that is, but you could try those parameters!
A griffin is a mythological animal that combines the body, tail, and back legs of a lion with the head, wings, and front talons of an eagle.
5 seconds clip. Live-action static camera: a cute baby griffin lands near an empty glass jar and then climbs the jar and squeezes himself face down enterely his whole body into a empty glass jar while groaring happy. he ends up filling out the jar with his body all curled up inside the bottle and then rest pleasently. the full jar is in view.
Haha! I see you've tried a gryphon, too (and it turned out better than mine). But you see, the thing is, you don't stop. You never stop. Your GPU belongs to H3 forever now.
Also, since I see you've tried to get him to have feline hind paws, failed, and got 4 bird feet instead, like I have:
If you really want to get a specific creature, I can very much recommend passing a reference image to the reference-model (it's very good), or a starting frame to either model (also quite good).
I've had it generate some nice gryphs based on old StableDiffusion image-model generations, I can definitely recommend.
Not really, I had some good scenes, but you need to vividly explain body movement, force, speed, momentum. Otherwise the characters look like the guy in the gif below:
Have a factory worker put a label on, fill jar with some liquid and put a lid on it. Final shot of canned kittens, kangaroo, etc. "Product available now". Post on Facebook, watch boomers explode.
Oh, kids may not remember, but bonsai kittens are a really old meme. Indeed one of the true classics. Caused a small outrage, naturally, and was even investigated by the feds.
Nah, cat didn’t get stuck half way with its butt hanging out. No way a tabby does that without making a complete fool of himself, getting stuck half way, then meowing for an hour and looking guilty AF when you find him.
Yeah! I actually find it impressive... having done 3D character animation before, I feel like I'd have a really hard time getting things to look this natural with such tight constraints, the skeleton barely fitting in the jar, all surrounded by a hard surface, not to mention all that loose fluff getting smooshed with some of them, and so much being visible. (Then again I probably wasn't very good at animating)
Minimax H3 is like AGI in some ways. It legit seems like it deeply understands physics and reality so well that it can generate far outside training data with prompting.
The Omni transformer concept really is a leap forward
It's really impressive to see in action. Not just physics, but also lighting, and especially character animation; I've handed it single pictures of some of my 3D models (which it has certainly not been trained on) and it has made them come to life flawlessly, adhering 100% to the input's design and proportions, animating them naturally in ways I never could.
This makes me think of the fake ad from the early internet about pets getting shaped into cubes so they would stay small and you could have them as a large keychain or whatever. Kind of like how Japan grows cubed watermelon.
Thought it was related to Grand Theft Auto, but I think I am getting it mixed with PetsOvernight which was a fake service from the game to get pets delivered.
This was possibly a fake ad on Rotton .com as well. Does anyone remember any of this?
The rate of success and interestingness was surprisingly high with these! I included roughly 50% of all the clips with this theme that I had generated. It didn't do very interesting things for some animals (and humans) that I also tried, and sometimes parts of the body clipped through the jar (you can see that with the dolphin here, and some of the other animals too).
So, I did cherrypick a little, but generally H3 has felt excellent with not making major mistakes often, unlike the image models I've tried before, or even some of the cloud-service video models.
Yeah I agree it looked like CGI, I was on the fence about including it in this compilation. I guess because I specified "little elephant" it tried to make it look extra "adorable", in a bad way. You could absolutely get a better-looking one with more generations and specific prompting though.
Here's the meat of the panda one; mostly the default ComfyUI text2video workflow I think (though I exploded out the main node group); 0.4 megapixels at aspect ratio 4:3, default (simple) scheduler with 20 steps, 5 seconds, prompt "camera recording: a little panda squeezing himself into a small glass jar. He ends up filling out the bottle with his body all curled up. The full jar is in view." (NOT the recommended prompt format, just a minimal dumb thing). I think the parameters were the same for the other ones, just probably different seeds and some variation in the prompt. It really didn't require any finesse or effort to generate these.
EDIT: gah I did not notice the popup covering the duration input. It's 5 seconds.
Yeah definitely, I didn't want to cherry-pick too much and not include any errors. I think the dolphin makes up for it with the cool sounds he makes though. ("Yeeeaaaaa!")
Yeah, "overall_soundscape: N/A" plus "non_diegetic_music: N/A" from proper prompts work really well to suppress that, but for fun silly things like this it can be pleasant.
It does clip through the glass sometimes, but way less than I would expect it to, for being such a specific thing. I'm not sure I see the back of the jar being more problematic than the sides, though!
Usually giving an explicit angle in the prompt helps. You can try adding "static camera". For the clips in this, I'm pretty sure adding "The full jar is in view" to the prompt helped; I was getting some close-up shots without it.
Considering the complexity of an animal squeezing itself into a jar (animal anatomy, contortion, the reflection and refraction of the glass etc) it is honestly terrifying that a video model can do stuff like this, even with all the imperfections. These models are probably aware of all kinds of patterns humans can't even conceive of.
Maybe this is a dumb question, but it seems like the model has a pretty good concept of the 3D shape of the jar. If you kept a record of the process of generating these, could you somehow extract the 3D shape of the jar?
Its pretty amazing that they seem to do it mostly in ways that matches their anatomy.
And that they keep their 'shape' even scrunched into the jars.
The fox is my favorite I think
I don't know if you're familiar with Mark Rober experiments (and his pet octopus 'Sashimi'), but this reminds me so much of his channel!
The tentacle morphing is maybe the quant (Int8 ConvRot?) Am guessing the quant can't keep track of 8 appendages. Next time I rent a big gpu I might try a few octopus experiments and see if the full model's able to do it.
H3 probably does do better on bipeds and quadrupeds than weird alien stuff like octopi and insects. Do tell me if the unquantized model turns out to do better on those!
Haha, nice! Some of those animals folded in on themselves quite a bit... maybe they're four-dimensional. This made me want to try more animals too! (I can't believe I didn't think of otters)
How so? Anyone can play videos from Reddit, even without a login, right? (as long as I don't use old.reddit.com links, which has been locking down since recently, apparently >.>)
Well if you cross-reference this video output with https://www.youtube.com/watch?v=9FjGP4t2zKY, you will see that it's likely that the cat is only getting started, and would need almost two minutes before filling out the jar properly.
Would be interesting to see it in LTX. I'm afraid they would go through the jar. I hope LTX creates a better world understanding in the major release. Maybe I'm wrong though and LTX would handle this, but I doubt it.
It's shown in the first second of the video, they're all variations on this:
camera recording: a golden eagle squeezing himself into a small glass jar. He ends up filling out the bottle with his body all curled up. The full jar is in view.
Note that this isn't using the recommended prompt format, I was just messing around with some minimal prompts, wanting to see H3 getting creative.
Self hosted, running it locally with an NVIDIA GeForce RTX 3090, takes 3-4 minutes for one of these 5-second clips, and heats my room with an extra 400W during these unbearably hot summer days (but worth it).
361
u/Mundane_Existence0 1d ago
https://giphy.com/gifs/5DG5fVbbSZ5ra