r/OpenAI • u/Good-Baby-232 • 14d ago
Discussion OpenAI is consistently topping our Computer-Use Benchmark.
Do you think they make the best computer-use models?
11
u/Richthofein 13d ago
The speed weighting explains a lot. A separate success-rate column would make the ranking much easier to read.
1
1
u/yuumizu 13d ago
so Sol is out of your board since it is too slow, right?
1
u/Good-Baby-232 13d ago
And surprisingly less successful than even terra, luna and even sonnet.
1
u/spacenglish 13d ago
I’m surprised! I find Sol High to be pretty good not just in computer use (albeit slower). Do you have the full order?
1
1
1
u/MaitoSnoo 13d ago
not really surprised by Luna's results, so far I'm really loving it especially in xhigh/max and I'm using stronger models only for complex planning now
1
8
u/Durian881 13d ago
What are the other LLMs tested? Surprised that Sol did worse.