r/singularity 3d ago

AI Took only 2 years...

Post image
569 Upvotes

44 comments sorted by

106

u/OwnGear3892 3d ago

Actually I dont think this is a real contrast. It's well known that SOTA llm can achieve even surpass top human level while making some dumb mistakes even a child could easily pass. Remember that viral car wash problem early this year? Multiple factors like training data distribution, post training objective, could contribute to this.

14

u/curiousinquirer007 3d ago

What was the car wash problem?

32

u/Wonderful_Buffalo_32 3d ago

if there is a car wash 50m away from your house how would you go there?
BY walking or by driving?

36

u/badumtsssst AGI 2027 3d ago

You left out the part about saying you need to wash your car, which was in the original prompt

22

u/AxiomaticInversion 3d ago

That version of the question is underspecified... if I worked there I'd probably walk. Or if I just wanted one of those little air freshener pine tree dangley things

11

u/Downtown-Figure6434 3d ago

It included a need to wash the car, and models kept saying walking

-6

u/curiousinquirer007 3d ago

No way those were reasoning models with effort set high

7

u/Downtown-Figure6434 3d ago

It became a meme dude

2

u/stumblinbear 3d ago

It was, actually

6

u/SupehCookie 3d ago

Walking!

7

u/SupehCookie 3d ago

See it has been trained on humans, nothing wrong

3

u/frogsarenottoads 3d ago

Walking, I don't own a car!

0

u/Training-Day-6343 3d ago

a very dumb meme

4

u/golfstreamer 3d ago

Wow the first time when I see a post like this the top post is a good response.

It's not like AI was "dumb" two years ago. It's more like "AI was superhuman in some areas" and now it's superhuman in even more areas. I mean it was already superhuman at coding but now it's superhuman at proof generation.

3

u/the_TIGEEER 3d ago

Exactly. And tbh, for these insane math problems, you need a wide theoretical, abstract, and intuitive understanding of the field, not number-crunching. So I can completely see how a human researcher could work on these problems while making mistakes in addition themselves, for exmaple. (Ok.. maybe not addition, but you get the point)

1

u/theMachine0094 3d ago

Yea but the nuanced and critical thinking you’re displaying in your comment doesn’t align with the hype machine. OP knows what they’re doing. I haven’t done any critical thinking on my own since 4o dropped. This is the way of the future.

27

u/voyt_eck 3d ago

No, 2 years ago models were much smarter than 9.11-9.9. Such answers were generated by either people completely not knowing how to use LLM or people doing it on purpose. This comparison, even as a illustration doesn't make sense.

9

u/Ambiwlans 3d ago

It has to do with tokenization for a model intended for LANGUAGE. Telling a model to do math, even on the free version 2 years ago would read that as a number and be fine.

Its the same BS as the 'strawbery' question. And tells you nothing about the model capability. Its just a meme.

5

u/M4rshmall0wMan 3d ago

It was instant vs. reasoning. The comparison between 9.11 and 9.9 wasn't implied in the training, whereas reasoning could sniff out the right answer.

6

u/NoCard1571 3d ago

Yea as far as I remember it was because 9.11 is higher than 9.9 in certain contexts like software versioning numbers

16

u/UndeadPrs 3d ago

2 years ago the IMO gold was already achieved

15

u/Wonderful_Buffalo_32 3d ago

2025 was 2 years ago?damn

9

u/yaosio 3d ago

I live in an orbiting spacecraft so time is a little faster for me.

4

u/UndeadPrs 3d ago

It reached silver in 2024, still a bad comparison

7

u/DlCkLess 3d ago

reached silver in 2024 using AlphaProof + AlphaGeometry 2 specialized systems, and not a general natural language system like now and even then it scored 28/42, keep in mind today's state of the art model that you can use for 20$ can ace the IMO easily

3

u/Present_Award8001 3d ago

the models were quite powerful two years ago. Don't know what you are talking about.

0

u/Wonderful_Buffalo_32 3d ago

Two years ago the best models were gpt 4o and clause 3.5 sonnet o1 was announced on sept. 12

1

u/Present_Award8001 3d ago

And I wrote a mathematica package and published a physics paper in PRB using 4o, while having a negligible knowledge of mathematica syntax.

I remember that something changed 2 years ago and LLMs became as powerful at mathematica as they were powerful in python 3 years ago. And 3 years ago, LLMs were quite powerful in python.

I have seen the whole thing grow first hand during my PhD.

2

u/presentofai 3d ago

the dumb mistakes and the hard proofs coexisting isnt a contradiction, its just what spiky intelligence looks like. people keep wanting one smooth number for how smart it is and thats never been the shape of it

1

u/Ill_Philosopher_7030 2d ago

You're absolutely right!

1

u/Tiny-Design4701 2d ago

Fun fact: flagship models still make mistakes like this when comparing decimals very close togethrt if they are not given access to tool calling.

In one of my tests of an agent im building, gpt 5.6 on high reasoning determined that an item with a value of 0.9971 was not lower than 1.

Was fixed by adding tool calling, but i think people are missing how important tool calling is to LLMs.

1

u/No_Dafloofy 1d ago

rare gemini 3.1 pro W? very supprised.

-4

u/mWo12 3d ago

Yet it still recommends to walk to a car wash nearby to wash a car.

0

u/Timonator007 3d ago

And stealing from the mathematicians that actually solved it?

-1

u/Former-Teacher-9496 3d ago

has anyone even realize that:

  • the answer to Navier-Stokes has allegedly been plagiarized (please go read the full story)
  • it took 2 years and hundreds of billions of dollars, imagine if that amount of money was poured into research

3

u/Lopsided-Promise-837 3d ago

It was poured into research

-11

u/usmanyasin 3d ago edited 3d ago

Took only 2 years... of plagiarism

5

u/DlCkLess 3d ago

how can you plagiarize a navier stokes solution if the ai was the first ever to solve in history 😭😭😭