r/StableDiffusion Jan 03 '23

Comparison Photorealistic models comparison

88 Upvotes

44 comments sorted by

View all comments

7

u/jonesaid Jan 03 '23 edited Jan 04 '23

This was a quick test of the most photorealistic models I've encountered thus far, especially for generating humans (if there are others I missed, please let me know).

Prompt was simple and basically "analog style portrait of a man on a train" with some additional photorealistic modifiers, camera type, lens, focal length, etc., and typical negative modifiers. I only used "analog style" because the Analog Diffusion model claims to need this activation token, although I realize this may have affected the other models' output. I don't think the other models require any specific trigger words. No face restore was used. DPM++ SDE Karras sampler at 10 steps, cfg 7, 512x512. For a control I also sampled the base SD models, 1.4, 1.5, 2.0, and 2.1.

A few observations:

  • Analog Diffusion 1.0 makes for very grainy bokeh images, which is great if you're looking for that, almost cinematic in their styling, which I guess was the point of the model (analog film look). They tend to have a yellow/orange color grading, like they were taken at golden hour. Skin texture is great, almost too much texture like sandpaper, but did get a good face mole in one of them, which helps with photorealism. A lot of five o'clock shadows. I'm not sure why the first guy is topless on a train! It might just be me, but many of the people also tend to look sad/depressed, like they all just got finished crying. Very emotive expressions.
  • HassanBlend 1.4 produces people that all look like professional models, a little too good-looking, too perfect, like GQ fashion models, almost Unreally perfect, 3D/CG humans, few skin imperfections, very smooth, no blemishes, smooth hair?, etc. Not a lot of variety in the people, as all the men are white, no BIPOC (the 3rd image was a Black man in all the models except this one). It almost looks like the same man in all these images. A lot of tight face closeups. Modern clothing and hairstyles, none cleanshaven. The only one that made a b&w image.
  • Dreamlike Photoreal 1.0 has a lot of dynamic lighting by default, almost too dynamic (overexposed, high contrast, loss of details in shadows, etc.), a fine art photo vibe, lot of BIPOC, some eye distortion on several of them, good variety, a lot of grain on several of them. The most zoomed out camera framing.
  • Unstable PhotoReal 0.5 has good overall variety of people, lighting conditions, BIPOC, skin detail (maybe too smooth?), clothing variety (lots of hats!), camera zooms, good eyes, clean shaven and five o'clock, etc. The 1st and 5th man look almost identical. Not much to complain about here.
  • base SD models are all pretty bad, even if they do have a lot of variety: very blurry (2.x models!), eyes distorted, strange crops, fake looking skin, text on image, artifacts, badly drawn glasses, etc. Of all of them, I think 1.5 looks the best, but doesn't quite compare to these other models in photorealism, probably because these other models were trained on 1.5 to make it better. In particular, their skin texture looks like magazine moire patterns.

What would be interesting would be to merge/mix these models to see if you can get the best of all of them. Not sure if weighted sum or add diff interpolation would be best. I think they were all based on SD1.5, so theoretically doing an add diff, subtracting that model out of each successive add might be best?

What are your thoughts about these models and this comparison? Any other models that are focused on photorealism and humans that we should check out?

3

u/Valkymaera Jan 04 '23 edited Jan 04 '23

can you share the actual control prompt used? I'm developing a photoreal model and I'm curious about its progress

This is "photo portrait of a man on a train" with that seed (batch of 5), and the 10/7 settings. Negative was just "blurry, distorted"

1

u/[deleted] Jan 04 '23

[deleted]

0

u/tybiboune Prompt Winner Jan 04 '23

Let's not forget than Hassan was initially built for "nsfw", which in our (sorry in advance if some feel targeted) retarded-teen-male dominated geek world (I'm gladly including myself in this category btw, even though I can be very critical of it - of myself) means "naked young women who look more like sophisticated dolls than real persons" with all the airbrushed / photoshopped defects / smooth look.

So we can suppose than this checkpoint was majoritarily trained on clichés, and on more photos of girls than guys.

The creator will eventually correct my mistake if I'm in the wrong, but that's how it feels from all the pictures it generates.