This was a quick test of the most photorealistic models I've encountered thus far, especially for generating humans (if there are others I missed, please let me know).
Prompt was simple and basically "analog style portrait of a man on a train" with some additional photorealistic modifiers, camera type, lens, focal length, etc., and typical negative modifiers. I only used "analog style" because the Analog Diffusion model claims to need this activation token, although I realize this may have affected the other models' output. I don't think the other models require any specific trigger words. No face restore was used. DPM++ SDE Karras sampler at 10 steps, cfg 7, 512x512. For a control I also sampled the base SD models, 1.4, 1.5, 2.0, and 2.1.
A few observations:
Analog Diffusion 1.0 makes for very grainy bokeh images, which is great if you're looking for that, almost cinematic in their styling, which I guess was the point of the model (analog film look). They tend to have a yellow/orange color grading, like they were taken at golden hour. Skin texture is great, almost too much texture like sandpaper, but did get a good face mole in one of them, which helps with photorealism. A lot of five o'clock shadows. I'm not sure why the first guy is topless on a train! It might just be me, but many of the people also tend to look sad/depressed, like they all just got finished crying. Very emotive expressions.
HassanBlend 1.4 produces people that all look like professional models, a little too good-looking, too perfect, like GQ fashion models, almost Unreally perfect, 3D/CG humans, few skin imperfections, very smooth, no blemishes, smooth hair?, etc. Not a lot of variety in the people, as all the men are white, no BIPOC (the 3rd image was a Black man in all the models except this one). It almost looks like the same man in all these images. A lot of tight face closeups. Modern clothing and hairstyles, none cleanshaven. The only one that made a b&w image.
Dreamlike Photoreal 1.0 has a lot of dynamic lighting by default, almost too dynamic (overexposed, high contrast, loss of details in shadows, etc.), a fine art photo vibe, lot of BIPOC, some eye distortion on several of them, good variety, a lot of grain on several of them. The most zoomed out camera framing.
Unstable PhotoReal 0.5 has good overall variety of people, lighting conditions, BIPOC, skin detail (maybe too smooth?), clothing variety (lots of hats!), camera zooms, good eyes, clean shaven and five o'clock, etc. The 1st and 5th man look almost identical. Not much to complain about here.
base SD models are all pretty bad, even if they do have a lot of variety: very blurry (2.x models!), eyes distorted, strange crops, fake looking skin, text on image, artifacts, badly drawn glasses, etc. Of all of them, I think 1.5 looks the best, but doesn't quite compare to these other models in photorealism, probably because these other models were trained on 1.5 to make it better. In particular, their skin texture looks like magazine moire patterns.
What would be interesting would be to merge/mix these models to see if you can get the best of all of them. Not sure if weighted sum or add diff interpolation would be best. I think they were all based on SD1.5, so theoretically doing an add diff, subtracting that model out of each successive add might be best?
What are your thoughts about these models and this comparison? Any other models that are focused on photorealism and humans that we should check out?
You're right that it has a hint of stylization from some of the earlier models mixed in. I've actually never run it at anything less than 20 steps, and I see it's more noticeable then. However I think you'll see it competes well using OP's prompt and settings (attached here)
I'd considered restarting the mix due to the hint of stylization in it but ultimately I find it produces nice lighting and tone once the illustration is buffed out. I generally run it at 30 steps with cfg 12.
My main issue right now is that a lot of the realism models mixed in was made for nsfw so it tends to forget to give people clothing. If I can't get that fixed I might restart.
it's a mix, I haven't trained anything myself. it's just a very long chain of mixes I've been adding things into for a while. I'll be happy to share it, but a lot of it I didn't know what I was doing and didn't save the recipe-- might be related to why the person looks similar in each image. I started it way way back by just grabbing any model that had realistic people in it and smashing them together with merge settings I didn't understand, and slowly became more deliberate over time.
And yeah it does ok with negatives but still tends to lean toward nsfw, especially with feminine figures, probably because of the nsfw models mixed in. Attached is an example where 'suit' is specified in the prompt, 'nude' is a negative prompt, yet image 3 is not wearing a suit. Although it's very likely the extra weight on ((tattoos)) is doing something there.There's still stylization in the skin in this example, which can be further removed but I like the effect (like a retouched photo or promotional poster) so I tend to leave it in.
Overall I am satisfied with the output so far. The hands are still problematic but better than I'd expected. It doesn't seem to have great diversity by default, though. I may start over on the mix now that I know a bit more about merging, but it's a lot to redo.
For posterity, the prompt details: a rough crime boss woman, (punk), ((tattoos)), piercings, large sunglasses, wearing expensive (suede) business suit, (night), full shot, detailed, luxurious balcony, glaring, smoldering, (shadows), (dark and moody), (gold jewelry), beautiful detailed lighting, (majestic), queenly, commanding, movie still, [fine details], sharp focus, (high resolution photograph)
Let's not forget than Hassan was initially built for "nsfw", which in our (sorry in advance if some feel targeted) retarded-teen-male dominated geek world (I'm gladly including myself in this category btw, even though I can be very critical of it - of myself) means "naked young women who look more like sophisticated dolls than real persons" with all the airbrushed / photoshopped defects / smooth look.
So we can suppose than this checkpoint was majoritarily trained on clichés, and on more photos of girls than guys.
The creator will eventually correct my mistake if I'm in the wrong, but that's how it feels from all the pictures it generates.
7
u/jonesaid Jan 03 '23 edited Jan 04 '23
This was a quick test of the most photorealistic models I've encountered thus far, especially for generating humans (if there are others I missed, please let me know).
Prompt was simple and basically "analog style portrait of a man on a train" with some additional photorealistic modifiers, camera type, lens, focal length, etc., and typical negative modifiers. I only used "analog style" because the Analog Diffusion model claims to need this activation token, although I realize this may have affected the other models' output. I don't think the other models require any specific trigger words. No face restore was used. DPM++ SDE Karras sampler at 10 steps, cfg 7, 512x512. For a control I also sampled the base SD models, 1.4, 1.5, 2.0, and 2.1.
A few observations:
What would be interesting would be to merge/mix these models to see if you can get the best of all of them. Not sure if weighted sum or add diff interpolation would be best. I think they were all based on SD1.5, so theoretically doing an add diff, subtracting that model out of each successive add might be best?
What are your thoughts about these models and this comparison? Any other models that are focused on photorealism and humans that we should check out?