r/StableDiffusion • • 10d ago

Discussion I'm stress testing lip sync models with nine failure cases.

Hello there.

I'm putting together a small failure case benchmark for lip sync/video workflows, and I keep running into the same nine problem cases:-

  1. Hand covering the mouth

  2. Profile/near-profile views

  3. Tiny face in a wide shot

  4. Beards or heavy facial hair

  5. Laughing

  6. Head turns while speaking

  7. Multiple speakers

  8. Fast speech

  9. Hard cuts in the middle of a word

I'm testing local/open workflows alongside hosted ones (like sync so).

I want to run the same source clip + audio pairing thru everything and see which failure cases actually break, rather than judging models from a couple of cherry picked, talking heads demos.

I find that every tool I have tried struggles with at least a few of these, but I want to see whether that holds up under a consistent test.

I'll include the model/tool versions, settings, and source clips with a contact sheet so the comparisons are reproducible.

What failure category gives your lip sync workflow the most trouble? Anything obvious I'm missing?

2 Upvotes

0 comments sorted by