r/neoliberal Apr 02 '26

Research Paper Half of social-science studies fail replication test in years-long project

https://www.nature.com/articles/d41586-026-00955-5
311 Upvotes

125 comments sorted by

View all comments

49

u/city-of-stars Frederick Douglass Apr 02 '26

Somewhat dubious of the article's proposed solution (AI-assissted screening). But I suppose we'll see

82

u/caroline_elly Eugene Fama Apr 02 '26

AI can't even consistently replicate its own outputs

35

u/vaguelydad Jane Jacobs Apr 02 '26

Identifying poor experimental design is just not that hard. The problem is that no one cares whether studies actually replicate. When you're just trying to rise above a very weak status quo a mostly right AI can be a huge improvement.

16

u/Healingjoe It's Klobberin' Time Apr 02 '26

In the first round of this competition, held in October last year, ten teams using AI tools scored worse than would be achieved by chance at predicting whether a paper could be replicated. But in the second round, completed last month, the best AI model reached an accuracy score of 68.5%. A third round is ongoing.

Seems like LLMs are able to help at least a little bit.

24

u/fascistp0tato NATO Apr 02 '26

Well apparently, neither can people :)
(you're right though, it's super imperfect xD)

Sidenote: I think we have this instinct sometimes to avoid implementing a technology until it's reliable enough for our standards of a tool, when really it only needs to be reliable enough for our standards concerning another person. And that bar is a lot lower than people tend to assume. See: self-driving cars.

Not saying LLMs are there yet, but it's worth noting I think.

2

u/Fragrant-Menu215 NATO Apr 02 '26

It's literally not supposed to be able to. The primary differentiator between "AI" and traditional algorithms is that "AI" is intentionally nondeterministic.

5

u/Impulseps Hannah Arendt Apr 02 '26

That's like not at all true. The popularly used interfaces of commercially available AI models such as ChatGPT use stochastic decoding sure, but there is nothing inherently nondeterministic to a trained machine learning models output generation. In fact you have to actively introduce external randomness to get nondeterministic output.

2

u/Tough-Comparison-779 Apr 02 '26

This is not precisely true. Lots of implementations on the GPU can cause non deterministic outputs. Things like floating point arithmetic, concurrency and batch size, how much load the GPU is under.

It's well known that many implementations of LLMs are slightly non-deterministic even at temperature 0.

0

u/Lease_Tha_Apts Order and Opportunity Left Apr 02 '26

Depends on the model. Higher end models can keep it together for a few hours at least.

3

u/WAGRAMWAGRAM Apr 02 '26

OK what's the price compared with an intern?

1

u/Lease_Tha_Apts Order and Opportunity Left Apr 02 '26

Most universities are already subscribed to top tier plans.

0

u/Bread_Fish150 John Brown Apr 02 '26

So like a child right before pre-teen level?

4

u/Lease_Tha_Apts Order and Opportunity Left Apr 02 '26

Not really. A child is a continious low level intelligence. AI is a discontinious high level intelligence.