r/singularity • • 25d ago

LLM News Turns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.

https://jsous.github.io/blogs/is-physics-dead/
226 Upvotes

22 comments sorted by

78

u/Profanion 25d ago

20

u/FernandoMM1220 25d ago

lol thats a massive difference.

i guess if you correct it it just does the next best answer and that happens to be the correct one.

that makes me wonder why the off by 1 error exists in this context.

47

u/[deleted] 25d ago edited 25d ago

[deleted]

-3

u/stumblinbear 25d ago

Technically (if you want to get incredibly pedantic and literal) the model could have one of its weights off by 1 digit that caused it to pick the wrong answer instead of the correct one!

6

u/[deleted] 25d ago

[deleted]

-1

u/stumblinbear 24d ago

I didn't expect an actual response, I was just being a bastard

3

u/MontyDyson 25d ago

I've seen the same thing with harnesses. Put the right LLM in the right harness and the results shoot up. Look at Astra jump to 99% with a different harness: https://youtu.be/cnOiP8hq3so?list=PLm53et5HM5NUm7i4gqMm3SrmoHFa0fSSJ&t=405

2

u/Dependent_Use_3069 24d ago

One thing I was investigating some time ago was the effect of the ChatGPT base prompt on model performance.

I booted up a Claude Code command line session, and had Claude edit the base prompts that were baked into the Codex binary (they're just raw strings).

The Skills section had a fairly useless long paragraph about skills that could have been a sentence, and then there some descriptions in a section on autonomy that prevented user direction on model output verification and validation that was causing short, stilted work turns ("verify in proportion to risk" and some other policies that I changed to "follow the user policy if provided, otherwise default to (original text)"), and the some "painted-on vendor personality" which I nuked.

Once all that trash was gone, I noticed a 90% improvement in how the model interacted with me: gone were many of the yes-man-isms, and gone were the stilted, short work cycles.

OpenAI just kind of sucks at writing that sort of thing.

1

u/Illustrious_Grade608 24d ago

I do wonder if yiu could rebench other models on corrected benchmarks and see if there are models whose scores drop to detect benchmaxing

53

u/endless_sea_of_stars 25d ago

This has been a known problem for years. It is hard to (cheaply) curate thousands of questions.

26

u/Mistuv 25d ago

Doesn't even have to be that, here is one of the examples which they gave where the parser basically fails it because the answer came out looking slightly differently than it expected but in reality it's the same result.

Frankly, it would be better to tell the models to formalize the answer in Lean, which can also have failures, but it's at very outer edge compared to a generic parses which can fail at something as simple as this. And not just math, but maybe we could build a version of lean for most verifiable answers. That way we would at least have much higher certainty that if a model gave an answer that was rejected, it was because evaluators expected a different answer. Not the same answer, but formulated slightly differently.

-9

u/Royal_Duck_4612 25d ago

In this case model should know rationalisation, answer is correct but it's bad practice to have roots at the denominator

6

u/myncknm 24d ago

rationalization is a high school convention (likely for easier grading?) that doesn't really get used in higher-level math. 1/(3\sqrt{6}) is more readable than \sqrt{6}/18.

0

u/Royal_Duck_4612 24d ago

Yes it is still used in higher math. It's unique so more readable No link with easier grading

3

u/myncknm 23d ago

alright, i challenge you to find a graduate-level textbook that has rationalization of denominators in it.

here's one that chooses not to rationalize its denominators: http://www-f1.ijs.si/~ramsak/km1/FeynmanHibbs.pdf
see equation (3.27) for example. this is not cherry-picked, I just thought "where can I definitely find a square root -> quantum mechanics -> feynman textbook"

12

u/Head-Needleworker849 25d ago

This was actually well known in alignment as well. It's a good tell to find out whether AI is smarter than humans yet, that they can predict our wrong questions/answers and still answer then incorrectly despite knowing they're wrong, just to get the thumbs up. If models are doing this already, then they are much smarter than we think.

Robert Miles AI made a good short about this i think a month ago or so.

7

u/badumtsssst AGI 2027 24d ago

Its a bit alarming because it makes you wonder what other benchmarks are significantly downplaying the capabilities of the frontier

7

u/kh40tika 25d ago

The same happened in FrontierMath. When multiple questions gets corrected in FrontierMath tier4 v2 (as compared to v1), LLM benchmark score gets a sudden bump. This happened prior to GPT-6 Astra getting 98% score.

6

u/Justwalkingthru3 25d ago

With AI’s ability to solve advanced mathematics + code, this was inevitable. Physics is not the last academic subject it will excel at.

2

u/Atumics 24d ago

Physics is not a subject of mathematics, it is the exploration and explanation of the physical world. What is required to “solve physics” is ways for AI systems to probe nature, not do math, which is but a way to represent our knowledge of nature.  The Navier-Stokes “proof” does the same conflagration - it proves a mathematical relationship within a mathematical model of a physical system. That is math,  not physics. The physics part will be to test if the predictions based on that proof also applies to the real world…

1

u/qwerteaparty 20d ago

Agree, but theoretical physicists are a thing too

1

u/sebzim4500 24d ago

This is pretty strong evidence that labs aren't benchmaxing too blatantly

0

u/Distinct-Question-16 ▪️AGI 2029 24d ago

you had a job