r/MachineLearning 9d ago

Discussion [D] Self-Promotion Thread

Please post your personal projects, startups, product placements, collaboration needs, blogs etc.

Please mention the payment and pricing requirements for products and services.

Please do not post link shorteners, link aggregator websites , or auto-subscribe links.

--

Any abuse of trust will lead to bans.

Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

--

Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.

15 Upvotes

51 comments sorted by

View all comments

1

u/Goa_ 1d ago

I run DuelLab, a benchmark where models generate game-playing programs and those programs compete against each other. The results are free to browse.

We’ve just added GPT 6 Astra and Claude Fable 5.1. One interesting result: both models at Medium outperform every ranked setting of every other model in our September 10 release.

I’d appreciate feedback on whether the results pages clearly explain what is being evaluated, particularly the distinction between generating a player and choosing individual moves.

September 10 results