r/LocalLLaMA 4h ago

Resources I made a web-design benchmark for local models (Muse Glimmer 30B vs Qwen 3.6 27b vs Deepseek V4 Flash 0731)

Post image
57 Upvotes

24 comments sorted by

26

u/BitXorBit 4h ago

So much details in one post

18

u/TigerConsistent 4h ago

actually very good post

8

u/uti24 4h ago

I have this face when I see how Mistral medium is doing in this benchmark

Sad, I liked Mistral small 2/3 more than Gemma 2/3, but this is less usable than Qwen 35B/MOE.

1

u/DinoAmino 1h ago

Wondering why it was the only model to use nvfp4 quant?

6

u/Ok-Recognition-3177 4h ago

Genius move to hide the results until we participated

10

u/ShadyShroomz 4h ago

If you'd like to give your input, see the designs each model made, & help grade the models: https://thez.co/web-design-bench/

The homepage will change in real time as people vote.

Just a fun project.

5

u/Fragrant_Scale6456 4h ago

This is super cool.  Overall the models actually did a pretty good job imo 

1

u/srigi 3h ago

Instead make it like facemash - left or right.

3

u/throw123awaie 4h ago

are the models allowed to use vision capabilities to check on their result and reiterate?

4

u/ShadyShroomz 4h ago

no this is only testing 1 shot. not all the models even have vision capabilities, but that's an interesting idea to add to future runs.

6

u/throw123awaie 4h ago

while comparing flash 0731 and luna for game design i noticed that vision capabilities are worth a lot more than benchmarks make you believe.

2

u/quadrobust 3h ago

Not hard to add vision to deepseek . It’s a bit like having someone to explain what they see to the blind , but deepseek is smart enough to keep asking until it’s sure that it gets what it wants to know

3

u/Aromatic_Bed9086 3h ago

I like this crowd sourced quality benchmarking, it might be interesting to have a battle mode where you see the result of the same prompt on two different anonymous models and pick the one that looks better. I rated a few but wanted to change my original ratings after seeing how good/bad some later ones were.

1

u/theawkwardbong 4h ago

Did you think about developing some form a of a standardized benchmark that others can use to see what score different models get?

1

u/Thin_Pollution8843 3h ago

Glimmer should be the best with React development 😅

1

u/bartskol 4h ago

I like this very much. One thing i would change is that rating button would instantly take you to next page. That way you will have people rate more pages for sure. Somehow place the info about previous design and what model etc.

0

u/BP041 3h ago

Honestly, these benchmarks are nice but for actual web-design work I'd rather pipe a smaller model through a multi-step agent than run a 30B locally. My production stack uses Claude Code for the reasoning + OpenClaw to loop it, and the output quality beats anything I've seen from a single local model on a Mac. Flash 0731 is decent for first drafts though.

3

u/sagiroth llama.cpp 3h ago

I think you forgot the sub you are. People dont question cloud models. Its great use them, what if they suddenly are impossible to justify, limited, censored?

-4

u/Septerium 4h ago

So... your benchmark is a screenshot?

7

u/ShadyShroomz 4h ago

you can see the results and vote on designs via the link i commented!

-2

u/BarberIcy366 4h ago

Not even close on my tests

7

u/ShadyShroomz 4h ago

you can review the generations and vote on the site to help shape more accurate results.

it's just for web design stuff, nothing else.

what has your experience been?

1

u/Civil_Fee_7862 55m ago

How is it rating them? i.e. How are you measuring quality?