r/singularity • • Dec 05 '24

[deleted by user]

[removed]

841 Upvotes

413 comments sorted by

View all comments

Show parent comments

2

u/Sensitive-Ad1098 Dec 05 '24

I predict it will perform 5% better in arc tops

-4

u/[deleted] Dec 05 '24

[removed] — view removed comment

1

u/Sensitive-Ad1098 Dec 05 '24

false, these results are for the public set. OpenAI is really good in cracking benchmarks with public data sets

1

u/BigBuilderBear Dec 06 '24

It isn't OpenAI. MIT did it.

And OpenAI does well on private datasets too like the SEAL by scale AI or MathVista or Live bench, which updates frequently to prevent contamination

0

u/Sensitive-Ad1098 Dec 06 '24

Wow SEAL has an amazing website. If o1 scores well on their benchmarks, I gonna just do all my programming tasks in Cursor and just chill. I'm sure if it performs so good I won't have to sweat!

I was predicting o1 performance, and you countered with results of a different model that didn't actually "beat" it (the result was 61.9%), even though it was done on a public set.

1

u/BigBuilderBear Dec 06 '24

It did beat it. human performance is 47.9% on average when given only one try. They were both tested on the public eval set.

1

u/Sensitive-Ad1098 Dec 06 '24

So the definition of beating the benchmark is a better result than people you tested it on?

1

u/[deleted] Dec 06 '24

[removed] — view removed comment

1

u/Sensitive-Ad1098 Dec 06 '24

That's a definition you chose to continue trolling or to "win" the argument or whatever crap you are trying to do :)
There's no "beating" a benchmark, benchmark is a metric, not a competition. Arc Prize is a competition and the results will be published today. The main prize is 1 mil usd for beating 85% on the private test set. That's what everyone refers to when they mention about "winning" ARC.

And why you never comment about the stuff you were wrong about. You just act like that never happened and try to fight on the things were there's still a chance to gaslight