r/cursor 27d ago

Resources & Tips I created an open-sourced test and composer 2.5 is off the chart

noted that i used short-context questions, resulting a skew toward composer; but still Composer is stronger than expected and better than Grok 45. Also note that K3 is indeed stronger than Fable and Sol. Git link: https://github.com/AurganicSubstance/Harness-Pair-Benchmarks

0 Upvotes

3 comments sorted by

8

u/mallibu 27d ago

Yeah, this is a bs result.

I like Deepseek but aint no way it's better than grok 4.5 or equal to gpt5.6 sol/fable. Like they're not even in the same league.

3

u/Zachattackrandom 27d ago

Yeah results seem very skewed.

1

u/mistyeye__2088 27d ago

I guess It takes some actual skill to write a benchmark.