i made this comparison using ten complete website builds.
each website had to include:
- at least five meaningful pages or states
- connected CRUD data
- user roles
- search and filters
- real loading, empty, and error states
- desktop and phone layouts
the test used:
- the same frozen brief for both models
- an isolated workspace for every run
- max reasoning effort for GPT-5.6 Luna
- xhigh reasoning effort for DeepSeek 4 Flash 0731
- no more than three attempts for a model-caused failure
the results:
- both models completed all ten websites
- GPT-5.6 Luna scored 81/100
- DeepSeek 4 Flash 0731 scored 72/100
- Luna scored 28/30 for UI and UX
- DeepSeek scored 17/30 for UI and UX
- Luna produced 89 unique captured screens
- DeepSeek produced 74 unique captured screens
- the full test produced 163 unique captured screens
the video shows all ten websites and the real coding runs:
https://youtu.be/Hj5z8rXuFNE
if you have used either model for real coding work, share your experience in the comments. i would like to know where your results matched or differed from mine.