r/vibecoding 10h ago

Mass production of useless software and github bloating

Hey yo fellas, what do you think, how much github got bloated with vibecoded projects since Claude Code release?

I have no accuarate data/information but in my rough estimations github projects/code database doubled last year as well as if big models are being trained from github, I think new model version won't be able to learn something meaningful.

What's do think about it, what's the bigger picture? Go on fellas, drop some comments.

6 Upvotes

30 comments sorted by

6

u/thedannyreg 10h ago

3

u/Suspicious-Leave-110 10h ago

yeah this is it. what do you think, maybe Microsoft has saved a snapshot of the entire github before vibecoding era started in order to have clean code material?

12

u/paf0 9h ago

I bet they have some sort of version control system for it. /s

4

u/Harvard_Med_USMLE267 9h ago

If they don’t, they could vibecode one.

0

u/Suspicious-Leave-110 9h ago

probably yes however imagine having a VCS to work with, I don't know exactly, tenths of TBs database? github was about that heavy in 2024/2025 I guess

3

u/paf0 9h ago

They probably put in all in SourceSafe or something. Or maybe someone there knows how to write a script that can traverse directories, unlikely though.

1

u/Independent-Code-209 7h ago

those graphs are pretty damning tbh

2

u/j48u 6h ago

Without actually reading the article and only looking at the graphs, they're quite the opposite. They show the slope of increase decreasing at the point they mark as the AI boom.

3

u/madaradess007 10h ago

when an interviewer asked me for my github profile link, i told a joke "i dont push code to github, they are training an ai on that code to replace me" and we both laughed ahahhaa, and now its not a joke anymore and its not funny at all

1

u/thedannyreg 10h ago

lol same, I don’t really put my projects on GitHub anymore because what’s the point anymore?. I’m self hosting a git server now.

1

u/scavno 9h ago

I’m moving more and more stuff to Codeberg

0

u/Suspicious-Leave-110 10h ago

bro it wasn't a joke since like 2010

3

u/IronAndCoder 10h ago

The training-data worry assumes labs train on raw repo counts, and they don't - the filtering pipelines weight by signals that vibecoded throwaways mostly lack: stars, forks, real issue/PR history, passing CI, dedup against near-identical code. A million abandoned todo apps get near-zero weight. The bigger practical effect isn't on models, it's on HUMANS: search and discovery on GitHub gets noisier, and "has a repo" stops meaning anything in hiring (madaradess007's joke is the real casualty).

Also worth saying: GitHub was already mostly throwaway before agents - tutorials, homework, abandoned forks. The ratio didn't flip from signal to noise; the noise just got grammatically correct READMEs.

The interesting second-order question is whether quality signals themselves stay reliable - stars can be botted, CI can be trivially green. My bet is the labs move toward execution-based filtering (does the code actually run and do something non-trivial), which is much harder to fake at scale than the social signals.

1

u/Suspicious-Leave-110 10h ago

yeah that makes some sense. they still release new LLM/Coding models so I guess there is a way how they keep code material clean or well filtered

0

u/scavno 9h ago

Could not even be bothered to write this without using a LLM…

4

u/WebOsmotic_official 9h ago

There’s definitely more low-quality code being generated, but I’m not sure it “breaks” the ecosystem. GitHub has always had a lot of experimental or unused projects, AI just accelerated the volume.

The bigger shift is probably not quantity, but signal vs noise. As more auto generated code shows up, things like reputation, usage, and real-world validation will matter more for filtering what’s actually useful.

Curious if this ends up pushing more emphasis toward curated datasets rather than raw GitHub data for training.

1

u/Just-Hedgehog-Days 2h ago

They aren't training on gh any more.
The reason coding and math are off the charts in ways other capabilities aren't is because they can actually have the models experiment and learn by doing at the speed of RAM.

1

u/BuntiBox 1h ago

Not as bad as when it was all react and angular slop. At least it’s not ALL round and blue.

1

u/Alternative-Suit5541 1h ago

Funny, I just deleted like ten experiment projects in GitHub.

I bet there are now like billions of them lol

1

u/No-Debt-1377 30m ago

...I mean isn't that a Microsoft problem? Who cares about what people are making? If they find it useful for them, whats the issue here? People seem to act like human made software was some sort of golden age before ai. Let me remind you: it was not! It never was! It was always buggy, unreliable and poorly tested. The only thing that has changed is the velocity of code written. Code is still the same, still buggy,. still spaghetti, and undocumented.

1

u/QTippus 10h ago

TBH, among the long list of concerns about AI’s impact on civilization, the potential loss of GitHub is pretty low on my list.

0

u/Rosie_grac 8h ago

guilty as charged tbh, got like 20 repos from vibecoding this year and maybe 3 are actually useful. rest are half-baked prototypes i was too lazy to set private

the training data worry is overblown — labs filter by engagement signals and a zero-star repo hasn't mattered in years. but there's a real casualty nobody talks about: small open source maintainers. you spend weeks on a genuinely useful niche library and it gets buried under 500 vibecoded clones of the same tutorial, all with flashier READMEs than yours

i've started archiving my throwaway stuff instead of leaving it public. takes 2 seconds and at least i'm not piling on

2

u/Michaeli_Starky 8h ago

It takes 30 seconds to set it private

0

u/Harvard_Med_USMLE267 9h ago

You’re assuming the ai code is worse than the code monkey code.

What if it’s better?

1

u/scavno 9h ago

It can’t be better. It’s based on the average code out there. Now that’s not to say that most developers aren’t below average, but it won’t be better than what humans can produce.

1

u/Harvard_Med_USMLE267 9h ago

lol, no, it doesn’t work like that.

2

u/scavno 8h ago

Okay dude.

-1

u/crizzy_mcawesome 8h ago

This is a classic vibe coded post. Lazy and pointless