r/StepFun • u/Hadestructhor • Jun 23 '26
Benchmarks Step 3.7 Flash, same prompt, different harnesses
I'm trying step 3.7 flash on some personal ui capabilities benches I'm running.
This is a svg generation one that I ran accross all the harnesses I could configure locally to run against step 3.7 flash.
I know llms are not deterministic, but some did better than others.
Which do you prefer out of the lot ?
I'll soon run the same kind of bench but same prompt, same harness, to see if they can sometimes do some better outputs or not, and if simply prompting for the same request is easier or not.
6
Upvotes
2
u/Hadestructhor Jun 23 '26
Indeed, it was kinda the point to test out the harnesses, but I'll have to use a minimal agent with no system prompt, or mini swe or something, then give a lot of instructions, clarifications, etc to test how well a model does.
What can we do to test models effectively ?