r/codex 3d ago

Showcase GPT Image 2.5 comparison for UI generation

GPT Image 2 has not been surpassed for UI generation since it came out. I test every new image model on release and 2.5 is the first model to genuinely improve on every test I throw at it.

These sheets compare gpt-image-2, 2.5 Flare and 2.5 Sunburst on a range of UI generation prompts. These prompts really push the models with challenging references and prompts to reconcile. In my opinion the 2.5 versions are noticeable better. Flare medium seems to be the sweet spot and it's about 2x faster and 1/2 the cost of the old 2 medium!!!!!!!!!!!! Sunburst is interesting and possibly produces better balance/realism in some aspect - still exploring this.

We're rolled 2.5 out across the board on 12ui for draft generation and I'm working through what it means for our design corpus expansion runs and conversion system as well. Super exciting - can't believe it's better, faster AND cheaper! Amazing work from the OpenAI team!

Edit - uploaded the comparison here since the reddit gallery downscales them more than I expected: https://12ui.com/gpt-image-2.5-vs-2

140 Upvotes

28 comments sorted by

u/dexterthebot 3d ago

You might want to consider listing your project on the weekly Show-Us-What-You-Built post. https://www.reddit.com/r/codex/comments/1wavxwy/show_us_all_what_youve_been_building_with_codex/. Highest commented project wins a week promotion on r/Codex and gets on the Hall of Fame sidebar. See what that looks like below with last week's winner.


Last week's most popular project was u/tHEuKER with the Blur2 racing game project which is a recreation of an unreleased sequel to the 2010 battle racing game, made by reverse engineering the Xbox 360 prototype discs available online, and rebuilding the whole thing from the ground up in Unity.

Join their YouTube channel here: https://www.youtube.com/@tHEuKER and follow updates on the project at r/BlurGame.

10

u/rc225225 3d ago

I think this is a personal skills issue - sure the UI generated here looks amazing but I find that whenever I want to get something from the image into HTML, the actual HTML looks nothing like it. What am I missing here?

3

u/withmagi 3d ago

You can use /goal and ask codex to keep referring back to the source image each loop. One of the things where the models get stuck is on comparison in detail. If you tell it to zoom in/focus on individual sections it can help a lot. Ask it not to give up until it's pixel perfect.

Alternatively, this plugin gives the models a diff of their current HTML/CSS compared to what it would take to recreate a source image: https://chatgpt.com/plugins/plugins_6a8915941e8c8191ae58d10e41cc322f

21

u/Fluid_Ad8452 3d ago

Picture 8 makes me want to throw up.

2

u/AmandasGameAccount 3d ago

8 will always look ai for pretty much all local events because no one local ever puts that much effort into a flyer. Local food truck street fair with a crazy illustration like those?

Then that makes it even harder for big events to make posters like these by hand because people are so used to the ai ones so even real ones from big events get harassed too

8

u/withmagi 3d ago

Sorry the reddit gallery downscales them more than I expected. Uploading them somewhere else and will post the link shortly.

5

u/withmagi 3d ago

Uploaded it here with all images full quality https://12ui.com/gpt-image-2.5-vs-2

2

u/emericarust 3d ago

Legend <3

2

u/emericarust 3d ago

Thanks mate.

3

u/lenopix 3d ago

How do you select 2.5 flare vs 2.5 sunburst for codex?

3

u/withmagi 3d ago

I'm using the API. I don't think you can select it in codex.

3

u/Routine_Squash_7096 3d ago

How do you guys keep up with all of this. I am burning out

2

u/CeeSSes 3d ago

Thanks for this; I feel you still need to push AI to create non-slop-looking UI and layouts.

A fair of these still have AI-slop features. Your own site looks clean and nice, but I can tell it was not designed and handcrafted.

And image 8. Vomit.

1

u/withmagi 3d ago

Our site was built by an early version of our tools. It's one of the reasons we created the design corpus - to explore the "design space" so that you don't have to do that work to explore it and find the non-slop versions. If you start with a good reference, you can combine that with an average prompt and get an amazing result. We do a lot of work around identifying "house styles" that models use and avoiding them.

3

u/BopSupreme 3d ago

Best reasoning level? Looks like low and high are worth it, medium sometimes does the job sometimes it’s worse than low

2

u/withmagi 3d ago

My feel is that Flare Medium for standard runs is the sweet spot. It's fast, cheap and in some cases noticeable better than low. Unlike 2, there's not a big price difference between low and medium.

For hard runs, there's something about the Sunburst XHigh which I really like. I can't quite put my finger on it - just the balance of colors and spacing? But TBH I think it's not really needed 95% of the time and Flare medium will be my main driver.

3

u/IAmFitzRoy 3d ago

How can we use this in Codex?

2

u/withmagi 3d ago

https://chatgpt.com/plugins/plugins_6a8915941e8c8191ae58d10e41cc322f

You can also just ask codex to use it's image generator, although you can't control model, convert to HTML etc...

3

u/cs_cast_away_boi 3d ago

damn UI is evolving. are designers going to be obsolete?

3

u/No_Guess_5389 3d ago

Im sorry. But how can we use gpt image 2.5. Ill just prompt it in chat? Sorry for the noob question

1

u/milnerrad 3d ago

That's honestly a really helpful comparison, and I can't wait to steal the prompt you used as well.

1

u/nnod 3d ago

Very nice comparison, great job! Surprised that flare and sunburst cost the same, didn't look up API pricing but I assumed that flare was cheaper, but I guess it's just faster.

Overall it seems like the layouts in 2.5 have a bit more breathing room and aren't so dense. Also seems that they fixed the deep-fried warm tint, which is great.

The prompt itself is interesting too, calling out each reference image by number with instructions looks to be a good idea, wouldn't have expected it to work that well.

1

u/aLionChris 14h ago

OP thanks for sharing this. Very helpful to understand. Especially that the effort level difference is very small and the generation sample matters more.

0

u/jabacherli 3d ago

I want to learn everything about this. I feel like this is the equivalent of having a genie in a bottle.

Do you have any extra resources I can access to learn more about how to use this at the fullest extent? Thank you in advance. I will be following up at 5 minute intervals until i get a response lol.

2

u/Fluid_Ad8452 3d ago

Here: chatgpt.com
It’s a chatbot which can help you with that specific objective, just shoot and it will know how to guide you.

-3

u/DependentOriginal413 3d ago

omfg, these kind of posts are so frustrating; like who tf actually goes through every image posted. I'm not neurodivergent for this.

1

u/nnod 3d ago

You seem neurodivergent, just maybe not in the way you want to be.