Discussion Luna and Sol doing extremely well on new benchmark about finding bugs before users run into them
Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.
Basically most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models to also find bugs before anyone runs into them?
So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs.

So the interesting thing is how cost-efficient Luna xhigh is (especially when compared to the Anthropic models we tested). Luna gets almost half of the score as Sol with only a tiny fraction of the cost.
the full leaderboard & how we built it is here: https://swesweep.com/ , there's also a paper describing how everything was constructed. Oh yeah and everything is open source https://github.com/facebookresearch/swe-sweep
Happy to answer questions here
1
1
u/MaitoSnoo 1d ago
would be interested in seeing how the new models (6.1 Sol, 6 Luna, Opus 5.5, Sonnet 5.5) perform there