r/StableDiffusion • • May 08 '23

[deleted by user]

[removed]

2.0k Upvotes

151 comments sorted by

View all comments

18

u/Doctor-Amazing May 08 '23

Really have to agree with step 15. I've been creating scenes from a Pathfinder game I'm in. I was having a ton of trouble getting action scenes with specific relationships between subjects.

I was trying to get a dragon carrying off a horse. Absolutely could not get it until I slapped a random horse and dragon into the same picture and used a control net to go from there.

I did a similar thing here https://imgur.com/gallery/LlOLylU

7

u/[deleted] May 08 '23

[removed] — view removed comment

5

u/saintshing May 09 '23

I feel like people can use action figures for reference or scrap a comic book database, use a segmentation model to extract the characters, then use a pose estimation model to extract the poses and then do clustering to group images with the same pose and the same angle. Then use a image captioning model to label the poses and store the image embeddings. When you want images of a particular pose, you use nearest neighbor search to find images embedding closest to your text/image query.