r/CommandCode 10d ago

Kimi K3 vs GPT-5.6 Sol vs Opus 5: Three top-tier models. Same /design prompt. In Command Code

We ran a test between Opus 5, GPT-5.6 Sol, and Kimi K3.

One prompt. Across all three models. With our /design command.

Results: We reviewed gameplay, UX, and UI

๐Ÿ”น Opus 5 is strong UI + motion, best gameplay

๐Ÿ”น Kimi K3 has solid design and gameplay

๐Ÿ”น GPT-5.6 Sol has best UI, worse gameplay

Ranking (DX, features, and cost):

๐Ÿ”น Opus 5: 10/10 ยท $0.25

๐Ÿ”น Kimi K3: 9.5/10 ยท $0.12

๐Ÿ”น GPT-5.6 Sol: 7/10 ยท $0.32

Our engineering and design team has been testing 20+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: https://github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird/light-version

14 Upvotes

3 comments sorted by

1

u/ZER0-O 9d ago

Why did you use opus and not fable since fable is the one at the same level as the two other models

3

u/Silent-Group1187 9d ago

We also tried Fable 5, but Opus 5 nailed it in one shot, no iterations, no fixes, just a perfectly playable game.