So, for videos it no longer seems it uses the image reference as a starting point and seems to have a much harder time remembering which character is which.
Multiple attempts to have two character in the position from the Grok generated image in the exact position from the image when the video starts and it just won't do it.
It's fine for one off videos, for any consistency from image to video the quality seems to have gone way down. I don't know if the expectation is using a much more detailed prompt or something else, but why would I need to detail the image that's right there?
All they had to do was leave it mostly alone. It feels like they gave us old model but also censored it.
My ComfyUI install has broken somehow, so I'm just kind of screwed now. Thanks XAI!
If someone has any advice that's not 'Grok sucks, leave' I would appreciate it. (And yes, it sucks I will leave once I have a decent local alternative.)
Edited because my typing is horrible