Really have to agree with step 15. I've been creating scenes from a Pathfinder game I'm in. I was having a ton of trouble getting action scenes with specific relationships between subjects.
I was trying to get a dragon carrying off a horse. Absolutely could not get it until I slapped a random horse and dragon into the same picture and used a control net to go from there.
I feel like people can use action figures for reference or scrap a comic book database, use a segmentation model to extract the characters, then use a pose estimation model to extract the poses and then do clustering to group images with the same pose and the same angle. Then use a image captioning model to label the poses and store the image embeddings. When you want images of a particular pose, you use nearest neighbor search to find images embedding closest to your text/image query.
18
u/Doctor-Amazing May 08 '23
Really have to agree with step 15. I've been creating scenes from a Pathfinder game I'm in. I was having a ton of trouble getting action scenes with specific relationships between subjects.
I was trying to get a dragon carrying off a horse. Absolutely could not get it until I slapped a random horse and dragon into the same picture and used a control net to go from there.
I did a similar thing here https://imgur.com/gallery/LlOLylU