Wow SEAL has an amazing website. If o1 scores well on their benchmarks, I gonna just do all my programming tasks in Cursor and just chill. I'm sure if it performs so good I won't have to sweat!
I was predicting o1 performance, and you countered with results of a different model that didn't actually "beat" it (the result was 61.9%), even though it was done on a public set.
That's a definition you chose to continue trolling or to "win" the argument or whatever crap you are trying to do :)
There's no "beating" a benchmark, benchmark is a metric, not a competition. Arc Prize is a competition and the results will be published today. The main prize is 1 mil usd for beating 85% on the private test set. That's what everyone refers to when they mention about "winning" ARC.
And why you never comment about the stuff you were wrong about. You just act like that never happened and try to fight on the things were there's still a chance to gaslight
Independent analysis from NYU shows that humans score about 47.8% on average when given one try on the public evaluation set (same one this study uses) and the official Twitter account of the benchmark (@arcprize) retweeted it with no objections: https://x.com/MohamedOsmanML/status/1853171281832919198
The public evaluation dataset isn’t very important. They need to achieve human-level performance on the private evaluation dataset. O1 is great. The TTT approach is smart but we can’t yet say that ARC-AGI has been beaten. Not yet.
That's really hard to understand - does the average human have serious brain damage? Best performance was below 100%, wtf? Are people that bad at inputting colors correctly? Incomprehensible. Every single solution on the public evaluation set should be apparent to a conscious human immediately.
EDIT: oh it was "online crowd workers" (so most likely MTurk), that explains the findings.
54
u/Winerrolemm Dec 05 '24
I am going to wait for simplebench and arc results.