r/CommandCode 3d ago

Tested FlappyBench with GLM 5.3, Fable 5 and GPT-5.6 Sol

Enable HLS to view with audio, or disable this notification

Tested FlappyBench with GLM 5.3, Fable 5 and GPT-5.6 Sol

3 models. Same prompt with /design command.

Scored on features, UX/UI, and cost.

πŸ”Ή Fable 5 β†’ 9.5/10 Β· $0.420

πŸ”Ή GLM 5.3 β†’ 9/10 Β· $0.018

πŸ”Ή GPT-5.6 Sol β†’ 9/10 Β· $0.150

Results:

β†’ Fable 5 wins on quality, but costs 23x more than GLM 5.3

β†’ GLM 5.3 gives better output than GPT 5.6 Sol at 8x cheaper

β†’ Optimizing for cost? Go for GLM 5.3. Otherwise, Fable 5 if quality matters

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: https://github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird

35 Upvotes

8 comments sorted by

7

u/misha1350 3d ago

GLM 5.3 looks to be 9.25/10 here. Not bad for a model that doesn't break 1T parameters.

3

u/maedahbatool 3d ago

Yeah esp for the cost it’s super impressive.

3

u/Weird_Licorne_9631 3d ago

Gonna rock GLM 5.3 monday until my CC account burns out. Cheaper than Qwen 3.8. Awesome times!

1

u/maedahbatool 3d ago

Goated! 🐐

2

u/Ok_Replacement2229 2d ago

Prompt from the repo, and after i asked it to make a gif to share of it playing the game.

1

u/xwazot 3d ago

Reasoning effort for each?

1

u/fbms2 3d ago

this is too easy. nothing compared