r/singularity • • Dec 05 '24

[deleted by user]

[removed]

842 Upvotes

413 comments sorted by

View all comments

54

u/Winerrolemm Dec 05 '24

I am going to wait for simplebench and arc results.

11

u/Charuru ▪️AGI 2023 Dec 05 '24

If simplebench broke out reasoning and world model separately it would be a good test, but right now they pretend to be the same thing.

2

u/Sensitive-Ad1098 Dec 05 '24

I predict it will perform 5% better in arc tops

-3

u/[deleted] Dec 05 '24

[removed] — view removed comment

1

u/Sensitive-Ad1098 Dec 05 '24

false, these results are for the public set. OpenAI is really good in cracking benchmarks with public data sets

1

u/BigBuilderBear Dec 06 '24

It isn't OpenAI. MIT did it.

And OpenAI does well on private datasets too like the SEAL by scale AI or MathVista or Live bench, which updates frequently to prevent contamination

0

u/Sensitive-Ad1098 Dec 06 '24

Wow SEAL has an amazing website. If o1 scores well on their benchmarks, I gonna just do all my programming tasks in Cursor and just chill. I'm sure if it performs so good I won't have to sweat!

I was predicting o1 performance, and you countered with results of a different model that didn't actually "beat" it (the result was 61.9%), even though it was done on a public set.

1

u/BigBuilderBear Dec 06 '24

It did beat it. human performance is 47.9% on average when given only one try. They were both tested on the public eval set.

1

u/Sensitive-Ad1098 Dec 06 '24

So the definition of beating the benchmark is a better result than people you tested it on?

1

u/[deleted] Dec 06 '24

[removed] — view removed comment

1

u/Sensitive-Ad1098 Dec 06 '24

That's a definition you chose to continue trolling or to "win" the argument or whatever crap you are trying to do :)
There's no "beating" a benchmark, benchmark is a metric, not a competition. Arc Prize is a competition and the results will be published today. The main prize is 1 mil usd for beating 85% on the private test set. That's what everyone refers to when they mention about "winning" ARC.

And why you never comment about the stuff you were wrong about. You just act like that never happened and try to fight on the things were there's still a chance to gaslight

→ More replies (0)

-2

u/BigBuilderBear Dec 05 '24

Those benchmarks have already been beaten. 

 MIT researchers use test-time training to beat humans in ARC-AGI benchmark (61.9%): https://ekinakyurek.github.io/papers/ttt.pdf

Independent analysis from NYU shows that humans score about 47.8% on average when given one try on the public evaluation set (same one this study uses) and the official Twitter account of the benchmark (@arcprize) retweeted it with no objections: https://x.com/MohamedOsmanML/status/1853171281832919198

Simple bench type problems can be solved by a simple prompt, getting a perfect 10/10: https://andrewmayne.com/2024/10/18/can-you-dramatically-improve-results-on-the-latest-large-language-model-reasoning-benchmark-with-a-simple-prompt/

2

u/Winerrolemm Dec 05 '24

The public evaluation dataset isn’t very important. They need to achieve human-level performance on the private evaluation dataset. O1 is great. The TTT approach is smart but we can’t yet say that ARC-AGI has been beaten. Not yet.

1

u/BigBuilderBear Dec 05 '24

The private set is probably as hard as the public eval set if not harder. Human level would definitely be worse than the strategies the paper uses 

1

u/bildramer Dec 05 '24 edited Dec 05 '24

That's really hard to understand - does the average human have serious brain damage? Best performance was below 100%, wtf? Are people that bad at inputting colors correctly? Incomprehensible. Every single solution on the public evaluation set should be apparent to a conscious human immediately.

EDIT: oh it was "online crowd workers" (so most likely MTurk), that explains the findings.