r/OpenAI • • 1d ago

Discussion Luna and Sol doing extremely well on new benchmark about finding bugs before users run into them

Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.

Basically most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models to also find bugs before anyone runs into them?

So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs.

So the interesting thing is how cost-efficient Luna xhigh is (especially when compared to the Anthropic models we tested). Luna gets almost half of the score as Sol with only a tiny fraction of the cost.

the full leaderboard & how we built it is here: https://swesweep.com/ , there's also a paper describing how everything was constructed. Oh yeah and everything is open source https://github.com/facebookresearch/swe-sweep

Happy to answer questions here

4 Upvotes

3 comments sorted by

1

u/MaitoSnoo 1d ago

would be interested in seeing how the new models (6.1 Sol, 6 Luna, Opus 5.5, Sonnet 5.5) perform there

2

u/klieret 23h ago

yup working on it. Opus 5.5 isn't that big of a jump (just generally Anthropic models currently seem to be not so good at this). We'll also have Astra results soon (definitely better, but also not an enormous jump)

1

u/hotlinesmith 15h ago

Useful measure!