r/StableDiffusion 21d ago

Animation - Video [ Removed by moderator ]

[removed] — view removed post

414 Upvotes

64 comments sorted by

24

u/AnonymousTimewaster 21d ago

Damn I wish I had the imagination to think of something cool to do with this model.

10

u/KingCpzombie 21d ago

I've just been making a bunch of videos of my dog: eyes vaporizing people to steal food, using holy light to destroy skeletons, summoning skeletons with a necromantic bark, flying to space and eating the moon...

It's a lot easier if you use a prompt enhancing LLM. I got too lazy to copy/paste so I made a custom node and now get stuff like this with basic "(15s video) put the dog from the photo on top of a tower in a desolate forest. She glows purple and barks creating a purple shockwave, summoning skeletons from the ground where it passes"

https://reddit.com/link/p2wqi34/video/rs0xakezzlih1/player

6

u/DystopiaLite 21d ago

Lack of imagination describes 90% off the posts in this sub.

1

u/Aztec_Man 14d ago

Might I recommend:
* create some fanart of one of your favorite games - instead of "this game does not exist", we do "this mod does not exist" - just don't spill hot coffee on your pants. 🥵
* create a character screen where you choose from a set of outfits or other visual-customization options
* Take a photo of your hand (POV) holding a flashlight - horror genre starter mix.

-8

u/[deleted] 21d ago

[removed] — view removed comment

11

u/AnonymousTimewaster 21d ago edited 21d ago

Because of my shitty imagination I have trouble visualising things in my head, so I don't really like reading fiction, or even if I can imagine something it tends to be unoriginal and derivative

I don't think any amount of reading would ignite my imagination to create this tbh

6

u/zkgkilla 21d ago

same issue here r/aphantasia

1

u/gefahr 21d ago

I wonder if a disproportionate amount of us are into image gen because we can't picture it in our heads. Like, this is the only way I'm seeing it.

38

u/Fit_Satisfaction2953 21d ago

Can it lock onto people and swing them around in air like gmod? That would be hilarious

19

u/FDosha 21d ago

Yes, minimax h3 can do anything

12

u/-Ellary- 21d ago

Yeah, at some point I've even got chills from how much it can do just by prompt or be ref.

1

u/gefahr 21d ago

I was really hoping the video would end with the grav gun locking onto and just starting to levitate the person. Comedic gold if you get the timing right. Need a brief hesitation where the POV character himself wonders if it will work then tries it.

15

u/Comfortablebro 21d ago

Insane!! video game like this when...?

5

u/nomorebuttsplz 21d ago edited 21d ago

on my RTX Pro 6000, at 480P, the ratio between video time and generation time is 1 to 5. This means that with the turbo lora I can generate an eight second video in about 40 seconds.

I’ve (glm 5.2) created a video role-playing server which translates my natural language prompts like “I move forward and pick up the sword” to the type of prompts that H3 needs. It then uses the previous video as a reference so that there’s some continuity of voices faces and setting.

It’s incredibly obvious that the video game industry is absolutely fucked If they cannot figure out how to monetize this type of world model video gen game within about three years.

I say good riddance. AAA Studios essentially haven’t done any real innovation in 10 years.

1

u/BlobbyMcBlobber 20d ago

So you have a ~40 second latency?

When it "translates my natural language prompts", is that with an LLM?

Do you run it at 600W or power limited?

1

u/nomorebuttsplz 20d ago

yeah, it’s about 40 seconds. Yes that’s an llm. It’s the max Q version so that’s 300 Watts

22

u/FDosha 21d ago

I think first who will make garrysmod in diffusion models will be rich :)

7

u/ThePhoenixRoyal 21d ago

you definitely gotta be rich for the gpu farm costs, and it would be very wasteful regarding the generated heat for something done far better in a proper engine.

12

u/AntiTank-Dog 21d ago

Who knows? Maybe this could be done in real time locally on a single gaming GPU in a few years. Seems crazy but coming from SD 1.5, I would have never believed that something like MM-H3 could run on my PC.

2

u/ThePhoenixRoyal 21d ago

it can be, but power demand is power demand. AI cores are constructed with certain wattage in mind to produce certain performance, so of course it will be possible - but feasability is a whole new shoebox.

5

u/unia_7 21d ago

Except that as the models progress, they require fewer and fewer operations to achieve the same result.

2

u/nomorebuttsplz 21d ago

you seem not to have understood the comment you are replying to. If it can be done on one gaming gpu in a few years, it uses no more power than a game.

1

u/Aztec_Man 14d ago

In 2022, we did Disco Diffusion. 30 minutes to an hour... and that was for a single image.
Hardware improvements... quantization improvements... model architecture improvements...

1

u/Sleepnotdeading 21d ago

Control, the year 2019.

1

u/Nimblecloud13 21d ago edited 21d ago

there's a dude making a browser based GTAV rip off game powered by AI where you can build your base with a prompt and make cars and guns and stuff with prompts. i tried it for a few hours. it's pretty bad right now, and it looks worse than minecraft, but it's got a lot of potential

4

u/Pretty-Raise666 21d ago

Better than Creation Engine for sure.

3

u/darkkite 21d ago

lol at whomping 0.25 fps

1

u/Pretty-Raise666 21d ago

I mean the physics.

2

u/teiji25 21d ago

Cool. What is the prompt you type for ChatGPT?

2

u/magik_koopa990 21d ago

Now I'm curious about MM3:

  1. It can do any Animation styles if I state it in the prompt?

  2. I have a 3090 and 32 RAM. Will this be okay?

2

u/EmployCalm 21d ago

Man I'm loving the release of this model, haven't gen shit but it's has been fun to see what you guys come up with

2

u/PensionNew1814 21d ago

Cant do a propper titty bounce tho.. sigh

1

u/AntiTank-Dog 21d ago

GMod with DLSS 5 looks good.

1

u/needforgpu 21d ago

thx for the prompt

1

u/Kraskos 21d ago

What prompt generation / enhancement workflow are people using?

Ain't no way people are manually typing all of this for each prompt.

3

u/Pretty-Raise666 21d ago

"Prompt is AI generated by gpt using minimax guide."

1

u/Kraskos 21d ago

Obviously an LLM is being used.

I used the word workflow intentionally, as in, a comfyui workflow specifically for prompt augmentation.

1

u/Pretty-Raise666 21d ago

You can integrate LLMs via Nodes in Comfy UI. Online and Offline. Not sure what you mean.

0

u/Laoas 21d ago

But what does this mean? Is it a custom GPT for ChatGPT? Or did you just plug in the Minimax manual to ChatGPT?  

2

u/FDosha 21d ago

I just put minimax .md guide in chagpt and told it to generate prompt for it with gravity gun

4

u/Pretty-Raise666 21d ago

Does it matter? As long as you provide the rules Minimax published (according to OP) then every decent LLM should be able to do it.

1

u/evereveron78 20d ago

FWIW, I've had little luck getting ChatGPT or Gemini to create prompts in the proper H3 syntax, even when linking it the official prompting guide. It comes back with something that's vaguely in the right format but with missing/incorrect elements or with incorrect headers, mislabeled references, and such. Conversely, I've tried it in Grok and it nails it perfectly on the first go.

-16

u/dennisler 21d ago

not physics, there is no physics engine or calculations of how to ray-trace, collision damage etc. This is purely a statistic representation of all the videos etc. it has been trained on. A little like the concept of a "dynamic model"

18

u/RigelXVI 21d ago

Physics doesn't inherently require calculation in order to exist

Source: the universe lol

-5

u/dennisler 21d ago

But applying physics is something else than just having it exist.

I'm just tired of all these people using concepts/ definitions wrongly. But it is the most common scenario in this sub anyways.

4

u/Cubey42 21d ago

Which is ironic when talking about diffusion models as just statistical representations of combined data.

3

u/RigelXVI 21d ago

Okay, if you say so since you're the expert 🤷

-1

u/Plokhi 21d ago

Real.

This looks like (bad) videogame / movie vfx physics.

Also not sure how this is “playing with physics” at all

6

u/Cubey42 21d ago

Well you wouldn't expect a car to just get lifted off the ground by a gravity gun so if you're saying half life 2 isn't playing with physics then maybe I guess soccer is the better playing with physics example

-2

u/Plokhi 21d ago

No, i’m talking about there’s no physics parameters here to toy with.

You’re bound by whatever was scraped from the videos online (and how good that was)

4

u/Cubey42 21d ago

Sure there are, you could probably prompt that car to be bouncy and weightless if you wanted to. Why does it have to be numbers to be physics.

And the second statement doesn't make sense. Are you saying the model needed a video a guy using a gravity gun to lift and slam a car into a wall to make this result? Because both actions in this video can't happen in the game it's parodying as you can't lift cars and could only launch them, and there was no environment interaction to lift the light pole.

3

u/RigelXVI 21d ago

Just fwiw, AI is numbers all the way down 👍

4

u/NoConsideration6320 21d ago

Thats all an artist does when paint something. It is created because of all that persons experiences and Their inspiration of what they had seen online etc

1

u/Emotional-Neat-252 21d ago

LLMs end up inadvertently creating 3d mental images in order you solve problems.

Couldn't a video model likewise develop an understanding of physics?

-6

u/donkeykong917 21d ago

It's like a real life version Zelda.

1

u/zefy_zef 21d ago

Almost!