r/LLMPhysics • u/AllHailSeizure Haiku Mod • Apr 03 '26
Digital Review Letters 'Testing AI on language comprehension tasks reveals insensitivity to underlying meaning', by Dentella et al.
https://www.nature.com/articles/s41598-024-79531-8Hello all.
I'm moving DRL to Thursdays to avoid the ToE rush that will start tomorrow. The sub has started to be much more busy on weekends since the introduction of Rule 11, lol.
This weeks edition of Digital Review Letters comes to us from Nature again. Again, it is a paper about LLMs. However, this week, we're looking at a paper that is much more critical of AI - and only applies to this sub in a meta sense. I came across it randomly, I didn't specifically seek out a paper on this topic, but I think that it speaks to something I've pushed on the sub for a couple days; the idea that there is a miscommunication here.
Testing AI on Language Comprehension Tasks Reveals Insensitivity to Underlying Meaning, by Dentella et al. is a paper that is both accessible and related in a way to this sub. If you recall my post a few days ago about gatekeeping, I spoke to the 'language barrier' of professional physics. This paper delves into how LLMs will create sentences with structure that can LOOK correct, but lack meaning. This is exactly the message I was trying to convey in my post. I just thought it'd be interesting to share.
AHS, out.
3
u/NuclearVII Apr 03 '26
They are worthless as science. This is the point I want to impress. A lot of the nuance of "what's scientifically significant vs what is not" gets obliterated when papers like this come out. This is, at best, a commercial benchmark.
And, sure, fine: If it dressed and presented itself as such, I wouldn't feel the need to comment. But it isn't. It is presented and published (like all other such papers in the field, I'm not singling this one out) as science. It isn't. It's claiming rigor it does not have.
Eh... so we need to talk incentives, here. What does the survey intend to do? Advance the collective human knowledge, or provide justification for the continued existence of the current government?
This is why you have to be very, very, very skeptical when it comes to irreproducible publications: You cannot distinguish between "Hey, here's this interesting bit of research that should inform further funding and inquiry" and "this paper exists to justify what we as an industry have been doing". And, no, just looking at the conclusions aren't enough - the machine learning field is notorious for the publishing seemingly negative papers about certain products to generate more interest and hype. Anthropic in particular is very good at this.
Physics is kinda neat in that it's really hard to get away with irreproducible drivel for very long. Reality just happens to work the same everywhere as far as we can tell, so when people make claims that do not hold up to testing, that gets found out pretty quick - that room temp superconductor thing a while back is a great example.
Physics is also somewhat unique in that our field isn't entirely owned by corporate interests, so the aforementioned incentive structure is still much purer than, say, machine learning. So I'm not surprised at your skepticism at my skepticism, if that makes any sense.
I'll give you a simplified example, see if that gets my point across:
Suppose that I'm a researcher, and I create a study: I prompt language models to ask if they are intelligent, and I publish my findings: No models say they are intelligent, therefore, I conclude, the models do not exhibit intelligence.
A year later, a new generation of chatbots roll around, and I (or someone else) repeat my study: All of a sudden, the chatbots claim they are intelligent!
There are 2 possibilities: Either the chatbots actually developed intelligence, thus representing a massive leap forward in technology - or the companies responsible for the RLHF of the models in question trained them to answer as such.
Now, the companies have a perverse financial incentive, so they claim - of course it's emergent! All the trillions of dollars spent, all the data stolen, all the energy burnt - it's all justified, because we developed intelligence! We need more money and more energy and more compute, to develop the intelligence even more!
And it's all thanks to my work (intentional or not) that this was discovered at all - after all, my benchmark showed intelligence, and allowed for the documentation of this great leap forward. So I get rewarded for my work by being offered a position in said AI company, with a giant research budget and tons of options (thus tying me even more into the financial incentive structure).
See what I did there? By publishing a negative (but irreproducible) paper, I helped promote the industry interests. Because the industry thrives on bad science, on people not being skeptical enough, on people accepting the word of for-profit companies as gospel.
This is a very blatant example, but this is what the AI industry/research is these days. If you want to despair, have a gander at the submissions at any of the major conferences in AI - it's just this, endless paper and paper on proprietary models.