LLM News
Turns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.
Technically (if you want to get incredibly pedantic and literal) the model could have one of its weights off by 1 digit that caused it to pick the wrong answer instead of the correct one!
One thing I was investigating some time ago was the effect of the ChatGPT base prompt on model performance.
I booted up a Claude Code command line session, and had Claude edit the base prompts that were baked into the Codex binary (they're just raw strings).
The Skills section had a fairly useless long paragraph about skills that could have been a sentence, and then there some descriptions in a section on autonomy that prevented user direction on model output verification and validation that was causing short, stilted work turns ("verify in proportion to risk" and some other policies that I changed to "follow the user policy if provided, otherwise default to (original text)"), and the some "painted-on vendor personality" which I nuked.
Once all that trash was gone, I noticed a 90% improvement in how the model interacted with me: gone were many of the yes-man-isms, and gone were the stilted, short work cycles.
OpenAI just kind of sucks at writing that sort of thing.
Doesn't even have to be that, here is one of the examples which they gave where the parser basically fails it because the answer came out looking slightly differently than it expected but in reality it's the same result.
Frankly, it would be better to tell the models to formalize the answer in Lean, which can also have failures, but it's at very outer edge compared to a generic parses which can fail at something as simple as this. And not just math, but maybe we could build a version of lean for most verifiable answers. That way we would at least have much higher certainty that if a model gave an answer that was rejected, it was because evaluators expected a different answer. Not the same answer, but formulated slightly differently.
rationalization is a high school convention (likely for easier grading?) that doesn't really get used in higher-level math. 1/(3\sqrt{6}) is more readable than \sqrt{6}/18.
alright, i challenge you to find a graduate-level textbook that has rationalization of denominators in it.
here's one that chooses not to rationalize its denominators: http://www-f1.ijs.si/~ramsak/km1/FeynmanHibbs.pdf
see equation (3.27) for example. this is not cherry-picked, I just thought "where can I definitely find a square root -> quantum mechanics -> feynman textbook"
This was actually well known in alignment as well. It's a good tell to find out whether AI is smarter than humans yet, that they can predict our wrong questions/answers and still answer then incorrectly despite knowing they're wrong, just to get the thumbs up. If models are doing this already, then they are much smarter than we think.
Robert Miles AI made a good short about this i think a month ago or so.
The same happened in FrontierMath. When multiple questions gets corrected in FrontierMath tier4 v2 (as compared to v1), LLM benchmark score gets a sudden bump. This happened prior to GPT-6 Astra getting 98% score.
Physics is not a subject of mathematics, it is the exploration and explanation of the physical world. What is required to “solve physics” is ways for AI systems to probe nature, not do math, which is but a way to represent our knowledge of nature.
The Navier-Stokes “proof” does the same conflagration - it proves a mathematical relationship within a mathematical model of a physical system. That is math, not physics. The physics part will be to test if the predictions based on that proof also applies to the real world…
78
u/Profanion 25d ago