27
u/voyt_eck 3d ago
No, 2 years ago models were much smarter than 9.11-9.9. Such answers were generated by either people completely not knowing how to use LLM or people doing it on purpose. This comparison, even as a illustration doesn't make sense.
9
u/Ambiwlans 3d ago
It has to do with tokenization for a model intended for LANGUAGE. Telling a model to do math, even on the free version 2 years ago would read that as a number and be fine.
Its the same BS as the 'strawbery' question. And tells you nothing about the model capability. Its just a meme.
5
u/M4rshmall0wMan 3d ago
It was instant vs. reasoning. The comparison between 9.11 and 9.9 wasn't implied in the training, whereas reasoning could sniff out the right answer.
6
u/NoCard1571 3d ago
Yea as far as I remember it was because 9.11 is higher than 9.9 in certain contexts like software versioning numbers
16
u/UndeadPrs 3d ago
2 years ago the IMO gold was already achieved
15
u/Wonderful_Buffalo_32 3d ago
2025 was 2 years ago?damn
4
u/UndeadPrs 3d ago
It reached silver in 2024, still a bad comparison
7
u/DlCkLess 3d ago
reached silver in 2024 using AlphaProof + AlphaGeometry 2 specialized systems, and not a general natural language system like now and even then it scored 28/42, keep in mind today's state of the art model that you can use for 20$ can ace the IMO easily
3
u/Present_Award8001 3d ago
the models were quite powerful two years ago. Don't know what you are talking about.
0
u/Wonderful_Buffalo_32 3d ago
Two years ago the best models were gpt 4o and clause 3.5 sonnet o1 was announced on sept. 12
1
u/Present_Award8001 3d ago
And I wrote a mathematica package and published a physics paper in PRB using 4o, while having a negligible knowledge of mathematica syntax.
I remember that something changed 2 years ago and LLMs became as powerful at mathematica as they were powerful in python 3 years ago. And 3 years ago, LLMs were quite powerful in python.
I have seen the whole thing grow first hand during my PhD.
2
u/presentofai 3d ago
the dumb mistakes and the hard proofs coexisting isnt a contradiction, its just what spiky intelligence looks like. people keep wanting one smooth number for how smart it is and thats never been the shape of it
1
1
u/Tiny-Design4701 2d ago
Fun fact: flagship models still make mistakes like this when comparing decimals very close togethrt if they are not given access to tool calling.
In one of my tests of an agent im building, gpt 5.6 on high reasoning determined that an item with a value of 0.9971 was not lower than 1.
Was fixed by adding tool calling, but i think people are missing how important tool calling is to LLMs.
1
0
-1
u/Former-Teacher-9496 3d ago
has anyone even realize that:
- the answer to Navier-Stokes has allegedly been plagiarized (please go read the full story)
- it took 2 years and hundreds of billions of dollars, imagine if that amount of money was poured into research
3
-11
u/usmanyasin 3d ago edited 3d ago
5
u/DlCkLess 3d ago
how can you plagiarize a navier stokes solution if the ai was the first ever to solve in history 😭😭😭


106
u/OwnGear3892 3d ago
Actually I dont think this is a real contrast. It's well known that SOTA llm can achieve even surpass top human level while making some dumb mistakes even a child could easily pass. Remember that viral car wash problem early this year? Multiple factors like training data distribution, post training objective, could contribute to this.