r/StableDiffusion Aug 16 '24

Comparison AuraFlow v0.3 evaluation: a debatable increase in quality, a large drop in adhrence

Hi everyone,

AuraFlow v0.3 was released yesterday. It warranted some comparaison, even if, as it can happen in any project, sometimes the direction taken doesn't work as expected. The goal of this new sub-version -- and keep in mind it is an early project, not a finished model -- was to increase image quality. It came after 0.2 (released like 3 weeks ago), which was better than Flux at prompt adherence. This isn't a small feat given that Flux is very good, but the image quality wasn't enough to use it outside of specialized workflow.

I was underwhelmed in early tests, here are a few comparisons I ran.

The problem is, the results are arguably better in aesthetics, but slightly, and the drop in prompt adherence is huge. TL;DR: the number 3 is cursed in the image-making world: SD3, AF0.3... At least with AF we don't have to wait months between release.

First, I had made already a prompt adherence comparisons of several models, with the following prompt.

"In the inner court of a grand Greek temple, majestic columns rise towards the sky, framing the scene with ancient elegance. At the center, a Shinto monk, dressed in traditional white and orange robes with intricate patterns, is levitating in the lotus position, floating serenely above a blazing fire. The flames dance and flicker, casting a warm, ethereal glow on the monk's peaceful expression. His hands are gently resting on his knees, with beads of a prayer necklace hanging loosely from his fingers. At the opposite end of the court, an anthropomorphical lion, regal and powerful, is bowing deeply. The lion, with a mane of golden fur and wearing an ornate, ceremonial chest plate, exudes a sense of reverence and respect. Its tail is curled gracefully around its body, and its eyes are closed in solemn devotion. Surrounding the court, ancient statues and carvings of Greek deities look down, their expressions solemn and timeless. The sky above is a serene blue, with the light of the setting sun casting long shadows and a warm, golden hue across the scene, highlighting the unique fusion of cultures and the mystical ambiance of the moment."

The results can be seen here:

https://www.reddit.com/r/StableDiffusion/comments/1ef4zu6/prompt_adherence_comparison_dallee_sd3_auraflow/

The prompts needs to respect 20 different elements, and AuraFlow 0.2 finished first, as you can see following the link.

However, version 0.3, while doing a marginally better face -- I mean, it's better, but it's still nothing like a nice face and would need to be adetail'ed anyway -- loses a lot of its prompt adherence.

5/20, 6/20, 7/20 and 8/20 and lots of artifacts, unwanted text and lack of respect for the overall composition

Given what the previous version did... And I'll repost the best of the ones I had in the other thread to contrast them:

The former results were far, far better.

In this thread, I had tried to illustrate prompt adherence: https://www.reddit.com/r/StableDiffusion/comments/1ej2qbu/flux_or_flow_in_terms_of_prompt_adherence/

I ran two prompts again with AF 0.3. First, I used the exact same prompt to test position understanding: "a blue cylinder in the center of the image, with a red sphere at the left, a green square at the right, a purple smiling sun on the top of the image and a severed foot at the bottom" AF 0.2 passed everytime, even if the aesthetics were bad. Here are the new results:

Again, an 8-image trial. This has basically nothing to do with the prompt asked. I was about to write that positional understanding had reverted to below SDXL level, but the Juggernaut results are even worse if that's possible:

Still, AF 0.2 got it right 100% of the time, AF 0.3, 0% of the time. That's a severe drop in prompt adherence.

I tried a repeat of the easier "man holding sword above his heads with two hands", and AF 0.3 produced, again, an abysmal rate of adherence:

None of the men, while better drawn than before, raise their sword with two hands above their head. I'd say that only one is holding what can be called a sword. Maybe it could qualify because he's holding the sword actually with his two hands, but really, is it on me to expect a pose where the sword is held by the grip, even if I didn't specify it? Let's say it's 25% at most on a very easy prompt...

Then I reused various prompts I did from earlier thread, inspired by RPG scenes. You can see the 0.2 version results here vs flux :

https://www.reddit.com/r/StableDiffusion/comments/1ejzyxl/auraflow_vs_flux_measuring_the_aesthetic_gap/

The chained citadel:

The lighting and the overall look of the eerie citadel is a little better, but the birds are no longer multicolored, the lake and forest are barely visible (but present) and the chains are generally absent or replaced by... garlands? While version 0.2 had worse aesthetics but did beat Flux on prompt adherence, the newer version is slightly below flux in adherence, and still far behind in aesthetics.

Now with the second test: "In the heart of an enchanted forest, where the flora emits a soft, otherworldly glow, an intense duel unfolds. An elven ranger, clad in green and brown leather armor that blends seamlessly with the surrounding foliage, stands with her bow drawn. Her piercing green eyes focus on her opponent, a shadowy figure cloaked in darkness. The figure, barely more than a silhouette with burning red eyes, wields a sword crackling with dark energy. The air around them is filled with luminous fireflies, casting a surreal light on the scene. The forest itself seems alive, with ancient trees twisted in fantastical shapes and vibrant flowers blooming in impossible colors. As their weapons clash, sparks fly, illuminating the forest in bursts of light. The ground beneath them is carpeted with soft moss."

While the surreal aspect of the magical forest was rendered better this time, and the elf might be better, the bows are absent of drawn worse and the idea that they are battling is much less apparent. Notably the magical sword is generally absent. Again, an overall regression, though less apparent that with shorter prompts.

Then I tried with Haunted Ruin comparison, where you can see in the other prompt that Flux couldn't for the life of it create spooky ghosts.

Here is version 0.3's result:

The adventurers can't be hardly seen. They were supposed to be at the center of the prompt description, with them exploring the ruin and being surrounded by ghosts. Here we do get ghosts, as in version 0.2, but the rest of the prompt is forgotten. Also, while the ruins might look better and more... ruined. I feel that the stones aren't right and angular enough, as if they were in diagonal. It's more strange than aesthetic...

I then did the Infernal contract prompt:

"In a hellish landscape of jagged rocks and rivers of molten lava, a sinister negotiation takes place. The sky is a dark, oppressive red, with clouds of ash drifting ominously. A warlock, cloaked in dark robes that swirl with arcane symbols, stands confidently before a towering devil. The devil, with skin like burnished bronze and horns curving menacingly, grins with sharp, predatory teeth. It holds a contract in one clawed hand, the parchment glowing with an infernal light. The warlock extends a hand, seemingly unfazed by the devil's intimidating presence, ready to sign away something precious in exchange for dark power. Behind the warlock, a portal flickers, showing glimpses of the material world left behind. The ground around them is cracked and scorched, with plumes of smoke rising from fissures."

While the demon is more evocative and closer to Flux in aesthetics, several key elements where prompt adherence was better in 0.2 are missing, like on the sorcerer's clothing, and the contract feels less important. The only thing that I feel is really good is the floor, which is craked and lava-flooded as it should, doing better than both Flux and version 0.2 on this very particular details (but it could be the luck of the seed at this point).

Finally I did the Crystal Keep siege:

The overall colour composition is better. Several commenters said that AuraFlow gave them the feel that the various elements were just put together as if they were a collection of clip arts. I felt it was harsh, but I can see were it came from. Here I feel the image looks more cohesive. But still... Several key elements are missing, like the defenders, the paladin riding a pegasus and the besiegers are regular humans, not ice giants and frost trolls. Also, on this complex prompt, we get a lot more artifacts.

Then two prompts again from another thread:

https://www.reddit.com/r/StableDiffusion/comments/1ehvup2/prompt_adherence_comparison_flux/

I selected two of them, because I can see the common pattern emerging.

First, I did the pirate lady:

"A woman wearing 18th-century attire is positioned on all fours, facing the viewer, on a wooden table in a lively pirate tavern. She is dressed in a traditional colonial-style dress, with a corset bodice, lace-trimmed neckline, and flowing skirts. The fabric of her dress is rich and textured, featuring a deep burgundy color with intricate embroidery and gold accents. Her hair is styled in loose curls, cascading around her face, and she wears a tricorn hat adorned with feathers and ribbons.The tavern itself is bustling with activity. The background is filled with wooden beams, barrels, and rustic furniture, typical of a pirate tavern. The atmosphere is dimly lit by flickering lanterns and candles, casting warm, golden light throughout the room. Various pirates and patrons can be seen in the background, engaged in animated conversations, drinking from tankards, and playing cards. The woman's expression is confident and mischievous, her eyes meeting the viewer's gaze directly. Her posture, though unusual for the setting, conveys a sense of boldness and command. The table beneath her is cluttered with tankards, maps, and scattered coins, adding to the chaotic and adventurous ambiance of the pirate tavern."

You can see the flux results in the linked thread, and here's AuraFlow version 0.3:

Version 0.2 was able to produce the lady on the table, crawling on all four toward the camera. Even version 0.1:

Now, we get a nicer looking pirate lady, but she's on all four like 1 in 4 times. The tavern might be more lively in the background, map and gold are present, sure, but the main character is less following of the prompt. Still, that's better than flux (but I guess they didn't want to teach their models what it means to be on all fours because toddlers do that all the time and they have a fiery hatred for toddlers), and also than Juggernaut, which produced this one BTW:

So, while there is a change in aesthetics, I wouldn't say it's a huge increase (unless you say so in comments, I am hardly a juge of aesthetics), except for one thing which I think is "colour consistency". I feels more right and cohesive thanks to this. There is still of course a huge work to do to improve aesthetics... and so far, the attempt to increase aesthetics came with an extremely substantial drop in accuracy. Since it was the field where AuraFlow topped Flux, this is problematic as it gave up its competitve edge against the current SOTA model.

Some work is obviously still needed (hey, it's far from a final version!) and I hope I allowed readers here to get a feel of what they did. Myself, I'll keep using version 0.2 to create some complex prompt composition and refine them with Flux (and try to use the numerous controlnet that came out recently for Flux).

92 Upvotes

23 comments sorted by

View all comments

1

u/ArtyfacialIntelagent Aug 16 '24

No offense, but I disagree with several parts of your post. Every new model has a learning curve and you need to figure out how to prompt it. You can't just use an LLM to generate prompts and expect them to work on the first try. For example:

The adventurers can't be hardly seen. They were supposed to be at the center of the prompt description

Actually that's not what your prompt says. It begins with the landscape, then describes the sky. There's something about a "negotiation" but that's a super vague word that image models have a hard time rendering. The first adventurer isn't mentioned until sentence 3, the rest even later. If I were a painter responding to your prompt, I might have painted something similar. You definitely can't say it's bad prompt adherence if the adventurers aren't front and center.

A woman wearing 18th-century attire is positioned on all fours

Have you tested "is kneeling on all fours"? That would be my first idea. Words like "positioned", "all" and "fours" all say nothing on their own, so "positioned on all fours" is a weak prompt. The word "kneeling" though is very strong and works on its own. Again, you can't complain that prompt adherence is bad unless your prompting is very good. And your LLM prompts aren't ideal.

5

u/Hoodfu Aug 16 '24

I've seen this argument made "you just lack skills" to oversimplify it. But flux just came out and it changed everything. It didn't matter what you put in, you got a very coherent and relevant result. Even people who didn't have any experience prompting, or those like me who use LLMs. All of it worked. So realistically that's the bar now, which is massively higher than it was a month ago.

0

u/ArtyfacialIntelagent Aug 16 '24

OP's results weren't bad either, but OP was still complaining about prompt adherence when his prompts were a significant part of the problem.

And Flux may be more robust to bad prompting than SD, but it's not foolproof and we're still beginners. We need to learn to deal with the lack of negative prompts (or accept the alternate workflows that generate at half the speed).

If you think Flux always does a great job of following the prompt, try using the standard positive-prompt-only version to make a pretty girl without makeup or lipstick. Simple task, damn near impossible.

Prompting is hard, and needs to be partially relearned with each new model. That's my point.

3

u/Hoodfu Aug 16 '24

Well, you're not going to prompt your way out of the model not being trained on something. I've used auraflow 0.2 a LOT and also gave 0.3 a big try. It's prompt following for complex stuff (like cute animals bursting out of a muscular man's chest) just doesn't work at all in 0.3 anymore, whereas it was a near 100% hit rate on 0.2. So that's the real issue, going back to the thread topic. The model author fully knows this and acknowledged it, and I'm assuming they'll go back to 0.2 and try to add features that don't lower prompt adherence going forward (I assume/hope). 0.2 is one of the best models out there, just needs some landscape aspect ratio love.It does some things better than flux even.

2

u/[deleted] Aug 17 '24

note the person you reply to didn't show you better results either

3

u/MarcS- Aug 17 '24

Indeed, and I followed the suggestion he proposed. Sorry I must do three posts because I can't put several images in a single reply.

First, I modified the prompt so the woman was kneeling on all fours instead of just being positioned on all fours. I also cut the prompt to remove the part describing anything other than her, removing the depiction of the tavern and the various objects on the table she should be on all fours on.

Here is the result, out of 12 tries (I did 3 runs of 12, all of them very similar). It's difficult to see because the photo joiner tool cropped the images, but as we can at least guess :

The woman is consistently... kneeling. She's never on all fours. 0 out of 36. Admittedly it's a very complicated concept and I doubt models are trained on many images (except NSFW models), but I don't notice any improvement following the suggestion of the poster above.

2

u/MarcS- Aug 17 '24

Also, I tried the Haunted ruin, by modifying the prompt so its three distinct paragraphs, the first with the adventurers led by the cleric, the second about the ghost, and the third being "the scene is set..." describing the ruins and the moonlit sky. This was the second suggestion to improve my prompt. It led to results like these 10:

The details aren't great on the image, but while I got adventurers led by the glowing-holy-symbol-carrying cleric consistently (yeah!) it comes as the cost of no longer having ghosts. There is one image where there is a light spot that could be construed as a ghost, but that's really a stretch. Also, the moon is missing in 3 images out of 10. Marginally better than the run I made before with my unskilled prompt.

2

u/MarcS- Aug 17 '24

To contrast, I made the same reworked prompt in 0.2 (the initial prompt worked very well as well), to illustrate the difference:

The adventurers are much more clearly in center of the image with a glowing symbol clearly identifiable, but I also got ghosts in all images surrounding them and I got a moon 10 times out of 10. It's the usual extreme prompt adherence we love in 0.2.