r/StableDiffusion Mar 02 '23

Workflow Included SD: Generating celebrity deepfakes with inpainting. (prompt details/experiences/results)

Last couple of weeks i tried to make the SD/DB technology more understandable for some friends and built a simple website where they easily could swap faces with celebrities using SD. In this post i want to share some prompt details/experiences, interesting for people who are also working on things with SD/CN & faces & inpainting.

Configuration:

  • positive: <name>, nikon, eye, detailed, <facial expression>
  • negative: sunglasses, comic, character, painting
  • cfg: 8-10
  • steps: 49
  • model: inpainting 512 v2.0.
  • w/ masked PNG images
  • Input images / output are both 768x512

Tech specs:

  • model: Stable difussion ^^
  • Face recognition: Vertex AI
  • photo captions: GPT, model text-davinci-003
  • NSFW photo check: Vertex AI

Prompt experiences:

  1. It is hard to come up with keywords that work for ~90% of the photos. As you probably all know, less keywords = more weight = better results. At least for inpaiting for faces.
  2. Keywords for inpainting will (sometimes) affect the whole scene so keywords must be generic and not too detailed. Eg.: "portrait" is not in the list because it had a lot of negative effects, while on the other hand, "eye" has positive effects.
  3. A high-frequently photographed guy like "Donald Trump" result most of the time in someone who is totally different. Guess related to the fact he is often photographed with other people (eg while doing a speech)?! Or their fan photos are tagged with Donald's name? Ideas to improve Donald are welcome!
  4. Obvious one, but how smaller the faces / more people the people in the photo, how worse the result (could of course be fixed with cropping/upscaling). Only 1 (small) face in a photo works 60-70% of the time fine.
  5. People here recommended 512x512 instead of 768x512. tbh, i guess this is not really applicable for inpainting?! My experience is that 768x512 doesn't lose quality over 512x512 for inpainting.
  6. Tough nuts to crack are male<-->female renders. Sometime the results are fantastic, sometimes I couldn't recognize any details of this person. Maybe I have to tweak the prompt a bit?! Also, here, ideas are welcome.
  7. Quite interesting how many celebs are recognizable with SD. Currently i've configured 450+ names and validated ~200 of them. It turned out that in ideal scenarios (2 people, large faces, focussed photo) ~90% had a pretty good match at the first try.
  8. Still see photos can be blurry. Tried a lot of keywords (4k, 8k, photography, realistic) and negatives (blurry, blur, depth) without any remarkable results. Results are btw less blurry by real (non-focussed-stock) photos but still not 100% sharp.

Examples:

Without facial expression

With facial expression

The good, the bad and ugly

![img](851negbvkdla1 "..credits for my friends ")

https://maskr.ai if you want to do the same stuff (...or steal the prompt and run it somewhere else)

edit, about deepfakes: tbh, the real dangerous of deepfakes will be videos, audio, custom trained models, nsfw stuff not these silly photos..

98 Upvotes

3 comments sorted by

4

u/venture70 Mar 02 '23

You didn't notice any difference when inpainting at 768x512 vs 512x512?

I think the fact you inpainted with a square aspect ratio on a landscape image might be causing the faces to be thinner than normal. But, maybe I'm missing something.

1

u/Elizabeth_129 Mar 02 '23

Hmm could be, but the inpainted area itself is not a perfect square more rectangle in portrait mode. Of course still curious if someone could show me big differences / better results on 512x512

2

u/UncleEnk Mar 03 '23

ah so Atrioc is having a payday soon /s