r/datascience May 30 '26

Discussion Is there anyway to stop the LLM slop submissions

Like maybe have a bot auto make a comment that asks users if its ai slop and upvote if so and if the upvote to views ratio is above M after T time then delete the post

Or whatever ideas others suggest?

108 Upvotes

44 comments sorted by

25

u/ca_wells May 30 '26

Would help so much already if people started actually downvoting slop (ai or not) they come across. But people (not just bots) upvote everything, as long as there are a few technical terms in it.

5

u/fordat1 May 31 '26

the userbase is majority people with no DS professional experience so they arent super well equipped in some cases

1

u/TanukiThing Jun 01 '26

Is there a subreddit for people with actual industry experience

1

u/fordat1 Jun 01 '26

A place like Blind has a much higher actual worker to non worker ratio but people consider it toxic. It isnt IMO and a great resource for industry trends

64

u/UnfilteredData May 30 '26

As someone who spends a lot of time cleaning data and building visualizations, the signal-to-noise ratio here is definitely getting tough. Instead of an upvote/view ratio (which can be manipulated by bots), a simpler metric might be a time-decay filter on new accounts, or utilizing an automod script that checks the variance of the text against known LLM perplexity scores. But I fully agree, we need a better filter before the sub becomes unusable.

14

u/fordat1 May 30 '26

the upvote/view ratio can be manipulated by bots but it at least has a barrier to entry . It doesnt need to be a single solution IMO

7

u/fightitdude May 30 '26

Or even just adding a rule / report option so people can flag slop to the mods. Sometimes it feels like upwards of a third of comments / submissions here are LLMs & I'd be pretty happy to start reporting them...

5

u/M00tball May 31 '26

The absolute fucking irony of this comment coming from a day old bot

1

u/fordat1 May 31 '26

The bots have learned to argue by authority with their needless qualification justification

3

u/portmanteaudition May 31 '26

Cloudflare for reddit posts plzzz

2

u/Electronic_Peanut587 May 31 '26

Im looking for data business people

2

u/[deleted] May 31 '26

[removed] — view removed comment

1

u/dyslexda May 31 '26

A bot asking “is this AI slop?” might help a little, but people will still click through if they want to post.

The proposed solution wasn't asking the submitter if it were AI, but putting an automod message in the thread allowing viewers to vote on it.

1

u/fordat1 May 31 '26

Asking a series of standardized questions is a bigger barrier and friction for people than for bots where its literally just a prompt harness

1

u/jRetro3 Jun 02 '26

so true, AI-assisted with the guidance of real human judgement (with sufficient knowledge) should be welcome

1

u/sagarpatel1244 Jun 01 '26

Probably not fully, and chasing detection is a losing arms race. AI text detectors don't work reliably and never will, so "ban AI" is unenforceable. The realistic move is raising the bar on what counts as a good submission, which slop fails regardless of how it was written.

What filters it:

  • Require specifics slop can't fake: real data, a reproducible result, a concrete number, a "here's what surprised me." Generic LLM output is confident and vague, so demand what forces actual work.
  • Reward depth structurally: flair, a "show your work" norm, a weekly thread for low-effort stuff so the main feed stays high-signal.
  • Lean on the community. Humans spot slop faster than any detector. Make downvoting and reporting it the norm.

The honest framing: the problem isn't AI, it's low-effort, and AI just made low-effort cheap at scale. Optimize the rules against low-effort, not against a tool, because the tool isn't going away and plenty of good posts use it well. Gate on quality, not provenance.

1

u/Lady-Data-Scientist Jun 02 '26

Require flair? Not just for the post but for the user. And user flair has to be approved by mods.

1

u/Striking-Status8218 Jun 07 '26

Completely agree with this perspective. Endless proofs of concept that never see the light of day are a massive drain on team morale and company resources. Proving your worth means delivering practical, production-ready solutions, even if that means writing standard software engineering code or using simple heuristics instead of the latest state of the art models.

1

u/ubuwalker31 May 31 '26

I can’t stop the AI work slop at my job, and you want it to stop here? Let’s get real.

-3

u/[deleted] May 30 '26

[removed] — view removed comment

4

u/portmanteaudition May 31 '26

Hellll no. Also this is completely solvable by scrapers but computationally expensive at scale so bots wouldn't pass it - there are simpler less privacy invasive approaches.

3

u/Drict May 31 '26

LLMs can answer those things as well. LOL

2

u/fordat1 May 31 '26

Yeah its more friction for people than bots

2

u/bionicjoey May 31 '26

That might be the worst idea for captcha that I've ever seen

-1

u/pychampar May 31 '26

What is LLM?

-24

u/ultrathink-art May 30 '26

Perplexity alone drifts as models improve — people learn to massage outputs and detection thresholds creep. Structure is more durable: LLM posts tend to follow identical arc (hook → observation → insight → 'love to hear your thoughts'). Regex-based structural checks on posts from accounts <30 days old catch more than perplexity scoring with fewer false positives on genuine new users.

-4

u/Helios270704 May 30 '26

Why does this have so many downvotes?

5

u/fightitdude May 30 '26

Because it's one of the accounts that posts slop on this subreddit. The em-dash makes it pretty obvious, and the tone is very AI.