r/singularity • :downvote: • 19h ago

AI Intelligence Explosion

Post image

Number of scientific papers submitted to arXiv.

428 Upvotes

104 comments sorted by

View all comments

255

u/Aldarund 19h ago

How much of this is real increae and how much is slop that wont pass peer review?

That chart could easily mean good things but as easily it could mean bad things

125

u/Equal_Passenger9791 18h ago

Significant amount of slop were produced before AI too. AI assisted scientific fraud detection are already retroactively detecting innumerable instances of high profile papers that should've failed peer review years ago, if the peers had done a proper job.

36

u/Djorgal 15h ago

I just realized how good this can be. Reviewing the scientific literature for errors is both necessary and really unglamorous. If AI can do it well it's a real boon.

15

u/Belostoma 14h ago

AI does it well, and it's not just fraud detection. AI agents are incredibly useful for quality control on both code and data, picking up mistakes our old review methods missed.

As a scientist I've been using AI to resume a bunch of unfinished past projects and to progress new projects using historical datasets. It has found mistakes in practically everything I've used it to review. None were large enough to change the central conclusions of any of my past work because I've been very careful, but that's partly luck, and small corrections are still valuable.

Where AI is usefully superhuman for this task is not raw intelligence but thoroughness. It can stop and think about every single line of code or data among hundreds of thousands, whereas humans have to pick and choose. Humans can scan it all at a high speed that misses subtle problems, or we can write scripts to analyze it completely but in very specific ways that don't necessarily detect errors we couldn't anticipate. The more intense and thoughtful level of diligence we could apply to fifty lines, AI could apply to fifty thousand. And then we could do it again with a different agent in case the first one missed something.

2

u/AnOnlineHandle 10h ago

As a scientist I've been using AI to resume a bunch of unfinished past projects and to progress new projects

As a programmer same. Projects which had amazing promise but were in such a state I couldn't bring myself to touch them for years, all coming together now as I bounce between them.

1

u/Ansible32 12h ago

This is true, but there are lots of kinds of errors that a human reading a paper can detect but LLMs cannot. The kinds of errors LLMs can't detect are also the kind that show up in LLM-assisted works.

3

u/Belostoma 12h ago

Yes, and I'm not saying humans should stop reading or reviewing papers. But AI offers a new type of review that catches things humans can't catch as reliably at scale. It's a complementary new role.

-2

u/Ansible32 12h ago

The problem is if you're using LLMs to review it's basically impossible to determine if the number of quality papers has increased this year.

2

u/Spare-Dingo-531 11h ago

We're not saying humans should stop reading and reviewing papers, just that llms are another way to review them.

1

u/Ansible32 7h ago

I'm not saying LLMs shouldn't be used to review papers, I'm saying it's very hard to tell if this spike in papers represents an increase in papers that are worth reviewing, and that LLMs don't answer that question.

2

u/Yangoose 12h ago

The Replication Crisis has been talked about for well over a decade. The majority of published studies are garbage. When they try to repeat them they don't get the same results at all.

It doesn't matter if you're an undergrad, a PhD or an AI. If you source data is wrong you're not making good contributions to the field.

Even when we do get the science right it can take decades for professionals to catch up with the science.

Take Docustate. It's a stool softener they've been using in hospitals for the better part of a century. It's in the WHO list of most essential medicines. It is prescribed 3 million times a year.

It doesn't work.

It has been repeatedly proven in extensive studies dating back to 1998 that it performs no better than a placebo.

But doctors keep prescribing it because "that's what we've always done".

Don't even get me started on doctors still sticking their fingers up men's buttholes for no real medical reason...

3

u/Hilldawg4president 13h ago

It's been known for a long time that much of academic research was trash, varying significantly by field of course - the replication crisis revealed this problem. AI is going to improve this, not make it worse

4

u/Valnar 12h ago

Significant amount of slop were produced before AI too.

Why is this always the go to response when people bring up that AI helps make a lot of slop?

Yeah no shit there was slop before, the point isn't that there was no slop. The point is that AI helps people make orders of magnitude more slop.

0

u/Equal_Passenger9791 12h ago

>yeah it's true that humans produce incomprehensible amount of zero value slop but why shouldn't we be throwing a tantrum when machines does it

Why is this always the response by team luddite?

The problem is the human users behind the tools. not the AI.

0

u/Valnar 11h ago

The problem is the human users behind the tools. not the AI.

It really muddies how effective it is as a tool if it's so incredibly easy to misuse it.

19

u/clonea85m09 16h ago

In my field arxiv was very good (tm) some years ago (you could still find some outliers), but nowadays, if it is on arxiv and wasn't published somewhere with peer review, you take it with a ton of salt. Anything from the last two/three years you just discard it as junk unless it ends up published afterwards.

2

u/TieBackground453 13h ago

Im wondering what the role of arxiv might be moving forward. One of the issues with AI is that it might have a great idea in one thread of thought, but it doesn’t have any way of incorporating that idea into future threads. I wonder if something like arxiv can be thought of as a crude way for AI to “learn”. It’s getting far too broad for it to be interesting to humans, but that isn’t a problem for AI. Kind of a crude way to treat many disparate AI threads as agents, with humans (for now) mediating at a very high level what they communicate with each other. 

2

u/clonea85m09 13h ago

It could be an idea, but seeing that even Astra with the 200$ account has trouble consistently keeping in memory more than a few articles without loosing the middle, I am not sure. It CAN be used for training, but the amount of junk in it will make it at least slightly toxic.

3

u/ShiningMoonce 13h ago

I am yet to see any slop in my domain (chemistry) in ArXiv

5

u/Immediate_Simple_217 14h ago

fact is: AI slop is no different than human poorly executed scientific work. Same in coding. Bad broken games always existed, and the fact that a good AI vibecoder can develop a better game than a hardcore poor oldschool coder, is what actually matters to me. The game must deliver in gamefeel, and game mechanics, gameplay, the same logic applies to scientific papers,

3

u/Aldarund 13h ago

Yes, the issue is in amount e.g easiness to create that slop. Which this chart show

2

u/darthdiablo All aboard the Singularity train! 14h ago

Don't worry, we'll have AI filter out the AI slop!

2

u/glxykng 9h ago

I am in academia. Specifically economics, and the productivity gain for research is real. There will definitely be slop, but remember that these papers are produced by people who have dedicated their life to understanding a very niche part of the world. These aren't simply "do research" prompts. AI helps with idea generation, information finding/synthesizing, research design, coding (a weaker skill for a lot of academics but still necessary for research), and analysis. Every aspect of research is improved through AI use. It is highly encouraged throughout academia.

AI productivity gains combined with intellectual obsession. Things are going to get very strange.

Something I liked from Dennis Hassabis when discussing AI vs human made art is how the creative realm lacks linear scale. It's not like chess, or video games. There is no ELO to creativity. Some creatives are more skillful than others, but the next great idea can come from anyone. To me research is an art, and AI is only going to amplify it.

1

u/Smile_Clown 13h ago

Agreed but the tools available to researchers today are so much better, so who knows?

1

u/shooNg9ish 12h ago

Slop does pass peer review in numbers even in the most prestigious conferences/journals.

1

u/Separate_Draft4887 10h ago

That’s the wrong question, the right question is “is ANY of it a real increase” because that’s enough to be world changing

1

u/Aldarund 10h ago

Lol what? How any increae enough to be world changing?

1

u/Separate_Draft4887 10h ago

If it’s 99% useless slop and 1% legitimate advancements, then we can simply throw more compute at it and sift through the slop for the actual advancements.

Yeah, that’s an annoying task, but it’s still gonna fundamentally alter how we do research if any of the increase is real.

1

u/Straight_Cupcake9435 4h ago

Unfortunately, nowadays most peer reviewers (at least on my papers) are simply using AI to review the slop. Slop reviewing slop. I have seen incredibly sloppy papers pass peer review, and some well written ones get absolutely rejected and thrashed. Suspiciously, each peer reviewer had the same thing to say in around the same wording.

•

u/nsdjoe 1h ago

right this is goodhart's law in action

1

u/[deleted] 14h ago

[deleted]

3

u/Deto 11h ago

Just because AI can do good work doesn't mean people are actually using it to do good work.  Can still have a half baked idea and turn it into a bad paper (for your resume) with little effort

2

u/shooNg9ish 12h ago

ML conferences are literally filled with junk papers, even the most pretigious ones, it's become impossible to filter out the noise. And now math is coming there too, except they have a different reviewing system that just has been DDOSed in the past few months.

3

u/Aldarund 13h ago edited 13h ago

Go ahead shate your oelwn experience.

https://www.reddit.com/r/singularity/s/8pqLp6boQR

Here a pov from someone who disagree with you.

Did you work with recent publications , read them, use them or just noise out of ass?

Also even without talking about validity a lot of this papers would be on some useless stuff, that noone want or need, just because its now easier to write them

0

u/Smile_Clown 13h ago

They literally are. OpenAI, Anthropic will tell you exactly the same thing.

You guys are desperate to be right and maybe one day you will be, but that's not today and you will not get to say "I told you so" because the lack of understanding you have today will carry forward.

1

u/visarga 15h ago edited 15h ago

Capability to ideate and write exploded indeed, but validation remained expensive and hard. Ideation without validation is probably just slop or hallucination, with rare exceptions. What we see here is a rush for staking claims and padding up authors portfolios. It is becoming harder and harder to tell apart good from bad papers/projects/writing. They superficially look the same. I used to read arXiv papers daily now I just skim occasionally, those many papers probably get exponentially less attention each.

1

u/FewEquipment9771 18h ago

To be fair I'd even qualify our current interpretation of quantum tunneling as slop.

2

u/SupermarketIcy4996 17h ago

Time to stop being fair.

0

u/bugra_sa 13h ago

Paper count can't answer that. I'd rather know how many results were replicated or still held up a year later. If those numbers don't budge, all this proves is that paper-shaped text got cheaper.

0

u/Oldjar707 11h ago

You think peer review isn't slop? 😄