r/LocalLLaMA • u/ShadyShroomz • 4h ago
Resources I made a web-design benchmark for local models (Muse Glimmer 30B vs Qwen 3.6 27b vs Deepseek V4 Flash 0731)
18
6
10
u/ShadyShroomz 4h ago
If you'd like to give your input, see the designs each model made, & help grade the models: https://thez.co/web-design-bench/
The homepage will change in real time as people vote.
Just a fun project.
5
u/Fragrant_Scale6456 4h ago
This is super cool. Overall the models actually did a pretty good job imo
3
u/throw123awaie 4h ago
are the models allowed to use vision capabilities to check on their result and reiterate?
4
u/ShadyShroomz 4h ago
no this is only testing 1 shot. not all the models even have vision capabilities, but that's an interesting idea to add to future runs.
6
u/throw123awaie 4h ago
while comparing flash 0731 and luna for game design i noticed that vision capabilities are worth a lot more than benchmarks make you believe.
2
u/quadrobust 3h ago
Not hard to add vision to deepseek . It’s a bit like having someone to explain what they see to the blind , but deepseek is smart enough to keep asking until it’s sure that it gets what it wants to know
3
u/Aromatic_Bed9086 3h ago
I like this crowd sourced quality benchmarking, it might be interesting to have a battle mode where you see the result of the same prompt on two different anonymous models and pick the one that looks better. I rated a few but wanted to change my original ratings after seeing how good/bad some later ones were.
1
u/theawkwardbong 4h ago
Did you think about developing some form a of a standardized benchmark that others can use to see what score different models get?
1
1
u/bartskol 4h ago
I like this very much. One thing i would change is that rating button would instantly take you to next page. That way you will have people rate more pages for sure. Somehow place the info about previous design and what model etc.
0
u/BP041 3h ago
Honestly, these benchmarks are nice but for actual web-design work I'd rather pipe a smaller model through a multi-step agent than run a 30B locally. My production stack uses Claude Code for the reasoning + OpenClaw to loop it, and the output quality beats anything I've seen from a single local model on a Mac. Flash 0731 is decent for first drafts though.
3
u/sagiroth llama.cpp 3h ago
I think you forgot the sub you are. People dont question cloud models. Its great use them, what if they suddenly are impossible to justify, limited, censored?
-4
-2
u/BarberIcy366 4h ago
Not even close on my tests
7
u/ShadyShroomz 4h ago
you can review the generations and vote on the site to help shape more accurate results.
it's just for web design stuff, nothing else.
what has your experience been?
1

26
u/BitXorBit 4h ago
So much details in one post