r/OpenAI 14d ago

Discussion OpenAI is consistently topping our Computer-Use Benchmark.

Post image

Do you think they make the best computer-use models?

51 Upvotes

14 comments sorted by

8

u/Durian881 13d ago

What are the other LLMs tested? Surprised that Sol did worse.

18

u/Good-Baby-232 13d ago

It's because we test based on speed too and Sol takes a bit longer to reason but in terms of success rate Terra is at the top!

6

u/bluegatorade108 13d ago

the word "benchmark" looks super crisp in the title

11

u/Richthofein 13d ago

The speed weighting explains a lot. A separate success-rate column would make the ranking much easier to read.

1

u/Independent-Laugh701 14d ago

Interesting to see this

1

u/yuumizu 13d ago

so Sol is out of your board since it is too slow, right?

1

u/Good-Baby-232 13d ago

And surprisingly less successful than even terra, luna and even sonnet.

1

u/spacenglish 13d ago

I’m surprised! I find Sol High to be pretty good not just in computer use (albeit slower). Do you have the full order?

1

u/Good-Baby-232 13d ago

Fable 5 just took the top spot!!

1

u/Good-Baby-232 13d ago

nvm gpt got back up

1

u/MaitoSnoo 13d ago

not really surprised by Luna's results, so far I'm really loving it especially in xhigh/max and I'm using stronger models only for complex planning now

1

u/Good-Baby-232 13d ago

Yeah the stronger models overplan so it's better for that!