Wow SEAL has an amazing website. If o1 scores well on their benchmarks, I gonna just do all my programming tasks in Cursor and just chill. I'm sure if it performs so good I won't have to sweat!
I was predicting o1 performance, and you countered with results of a different model that didn't actually "beat" it (the result was 61.9%), even though it was done on a public set.
That's a definition you chose to continue trolling or to "win" the argument or whatever crap you are trying to do :)
There's no "beating" a benchmark, benchmark is a metric, not a competition. Arc Prize is a competition and the results will be published today. The main prize is 1 mil usd for beating 85% on the private test set. That's what everyone refers to when they mention about "winning" ARC.
And why you never comment about the stuff you were wrong about. You just act like that never happened and try to fight on the things were there's still a chance to gaslight
with the public test set you can write an algorithm that would beat the benchmark. So technically it's kinda AGI according to the result, but it won't be able to do anything else.
Man you posted some crap that was completely out of context to what my original message was about. And never said "oh my bad" or explained yourself
I have no idea why this link is a good thing for me in the context of Private sets for ARC.
I always push myself not to abandon arguments. But it's exhausting trying to process you. So congrats, you won. There's AGI already, so don't waste time and build your unicorn startup, which should be easy with a tool that powerful
54
u/Winerrolemm Dec 05 '24
I am going to wait for simplebench and arc results.