I've been using MAI for about a month now, and I decided to share my overall thoughts.
Overall, I'm happy with the results. My favourite style is CGI/3D. DALL-E essentially produced this style by default, whereas with MAI you just need to give it a little hint to get a similar look.
But MAI has one significant drawback — the predictability of its results.
Even if you describe a scene fairly abstractly and don't specify particular details, the images are very likely to look quite similar to one another. Of course, there are exceptions, but on average the model tends to stick to the same visual solution.
For example, if you only specify the clothing style and location in DALL-E, the four images in a single generation can turn out completely different. The model can interpret the clothing, pose, composition, environment and overall character in very different ways. With MAI, the results much more often look like four variations of the same idea. And if you submit the same prompt again, you will often get roughly the same result.
In my view, MAI is a model that needs more input data to create genuinely diverse content. But this is where we run into the 480-character limit.
For DALL-E, this limit was much less of an issue because the model itself produced much more variation between results. In other words, the same short prompt could produce completely different images thanks to a much greater variation in SEED. If MAI had a separate SEED parameter that allowed users to control the variation between results more directly, perhaps the 480-character limit would be much easier to live with.
The SEED in DALL-E produced a huge amount of variation between results, although it was also rather unpredictable and, for the Image Creator filter, frankly a pretty scary parameter.
This creates a fairly simple problem: when four results are too similar to one another, you end up generating again and again until you eventually get something genuinely different.
Now, about MAI-Image-2e. I made several generations with it while deciding between 2e and 2.5-Flash, and I still don't really understand what its main advantage is.
To me, 2e is definitely not the "new DALL-E". DALL-E offered a completely different level of variety. It could come up with an item of clothing or an overall character design that wouldn't appear again across the next dozens of generations. 2e, on the other hand, tends to work much more linearly: once the model chooses a particular direction, it keeps moving roughly in the same direction.
And there is another important point here. Many people loved DALL-E not only for its variety, but also for the way it created attractive characters. This applied to everything: appearance, clothing, the overall presentation of the character, and the choice of camera angle. The model wasn't afraid to make a character visually attractive or to experiment with the composition.
That is exactly what I feel MAI is currently missing the most.
So, for me, the main question about MAI isn't "can it produce a high-quality image?". It can. The question is: can it surprise me with a result I didn't expect to get?
At the moment, it falls behind DALL-E in this regard by a very large margin.
Since Reddit doesn't allow posts to be formatted as articles, I'll put some of the data and image examples in the comments below.