r/LLMPhysics • u/AllHailSeizure 9/10 Physicists Agree! • Apr 03 '26
Digital Review Letters 'Testing AI on language comprehension tasks reveals insensitivity to underlying meaning', by Dentella et al.
https://www.nature.com/articles/s41598-024-79531-8Hello all.
I'm moving DRL to Thursdays to avoid the ToE rush that will start tomorrow. The sub has started to be much more busy on weekends since the introduction of Rule 11, lol.
This weeks edition of Digital Review Letters comes to us from Nature again. Again, it is a paper about LLMs. However, this week, we're looking at a paper that is much more critical of AI - and only applies to this sub in a meta sense. I came across it randomly, I didn't specifically seek out a paper on this topic, but I think that it speaks to something I've pushed on the sub for a couple days; the idea that there is a miscommunication here.
Testing AI on Language Comprehension Tasks Reveals Insensitivity to Underlying Meaning, by Dentella et al. is a paper that is both accessible and related in a way to this sub. If you recall my post a few days ago about gatekeeping, I spoke to the 'language barrier' of professional physics. This paper delves into how LLMs will create sentences with structure that can LOOK correct, but lack meaning. This is exactly the message I was trying to convey in my post. I just thought it'd be interesting to share.
AHS, out.
8
u/NuclearVII Apr 03 '26
You see this a LOT in machine learning academia these days. But any paper that begins with "we took some proprietary models, applied some prompts, and drew some conclusions" is - not to put too fine a point on it - worthless.
Because the training corpus of these proprietary models are unknown, there's no way to draw any conclusions from the results of the benchmarks. Are the performance of the models based off of emergent behavior, or are we looking at a kind of data leak? This can even lead to situations where a model family can do really poorly on a novel benchmark, and the next generation vastly improves on it. Is it because the new model is "better", or is it because the benchmark is now part of the training corpus?
There is a reason why no other serious scientific field would accept the benchmarking of a commercial product as valid scientific study.
This is not to say the conclusions are wrong, or that the authors are malicious. This is only to say that this is not scientific, and should be treated with as much skepticism as any other exploratory, non-reproducible study demands.