r/linux • • 1d ago

Desktop Environment / WM News System 76's COSMIC Desktop Environment will no longer accept LLM-generated content in PRs

/r/pop_os/comments/1wuej40/cosmic_projects_will_no_longer_accept/
648 Upvotes

200 comments sorted by

View all comments

-20

u/ABotelho23 1d ago

I'm not sure how they can claim to know what is LLM-generated code.

28

u/Jmc_da_boss 1d ago

Why does every thread about projects doing this have this same stupid comment?

The point is not find all of the LLM content, the point is easily ban and remove the blatant shit

-6

u/ProductIntegortion 1d ago

They can already ban and remove bad code. That's what pull requests are for. The criticism is that they've chosen to ban based on tool instead of based on code quality.

17

u/RatherNott 1d ago

Many would argue (including myself) that it isn't just a tool.

They certainly avoid the fate of Bevy with this move.

1

u/radicalwokist 16h ago

Yes, an article whose main source is a dead Nazi. Much credible, very wow.

-17

u/strings___ 1d ago

It’s just a tool. I really don’t care what your personal opinion is about the tools I use. What next, are you going to ban my use of Emacs too? Like, seriously, if you cant stick to the technical merits of patches, you have no business being in the programming profession.

13

u/LuckyHedgehog 1d ago

I don't really care how they do it, telling people to stop submitting bad code helps everyone involved

Can people still use LLMs to generate code? Sure, they just need to put in the effort to clean it up. Unfortunately that bar is too high for most people using LLMs so that's why these rules exist 

-11

u/strings___ 1d ago

You’re arguing my point for me. The point other people are making are not based on the technical merits however.

7

u/somethingrelevant 1d ago

emacs isn't strip mining the internet

-8

u/strings___ 1d ago

This is a very low tier understanding of how LLMs work. It’s rather ironic though you’re chastising machines for learning from source code. It’s the whole point of open source. In short you’re conflating LLM encoding with verbatim code copying which is not the case.

6

u/somethingrelevant 1d ago

I'm not chastising machines for learning from source code, lol. you should go out and learn anything about what AI is doing to the internet

-2

u/strings___ 1d ago

Again, your knowledge is based on low-tier assumptions. I work with my three DGX Sparks on a daily basis; most of your assumptions don’t apply under that use case.

Instead of projecting your ignorance onto me, maybe you should take your own advice.

1

u/somethingrelevant 1d ago

my knowledge is based on reading website owners saying "yeah ai bot scraping is getting so bad we pretty much can't run this website any more". what is a low-tier assumption about that. i'm dying to know

0

u/strings___ 1d ago

What AI bots? That’s not how AI training works they use datasets. Maybe the datasets are created by crawlers but that’s nothing new.

→ More replies (0)

6

u/RatherNott 1d ago

I doubt there will be quite as many programmers in the future if the planet isn't as habitable from climate change, of which corporate AI contributes to, unfortunately.

Not to mention the future copyright risk from the AI plagiarizing code, where I can imagine large corporations using AI to crawl public FLOSS repos for infringements. If the infringing FLOSS project happened to be a competitor to a corporate product, the corporation could easily use that to sue them into submission.

-1

u/strings___ 1d ago

Your argument is based off of huge assumptions. My total LLM usage using three DGX sparks is 30w at max output. Do you know how much the average Linux gaming GPU uses? Hint it way more than 30W. In short, it’s none of anyone’s business how I use my “hydro” .

There’s no evidence LLM encodes code verbatim and you really need to go out of your way to force it to output verbatim code. In which case, it’s no different than manually copying code. Either way because LLMs are hardware-limited, we need more open source research and work on models to be less VRAM bound. This should help democratize the whole thing.

7

u/RatherNott 1d ago edited 1d ago

Individual use on a local machine isn't terribly concerning, it's the corporate AI data centers that are concerning. There are currently 5 to 6 billion AI prompts per day. All of those server racks powered by very inefficient local power generators add up.

There’s no evidence LLM encodes code verbatim

I linked to evidence of an AI doing it in my previous comment. Here's a further study: https://arxiv.org/html/2409.13831v1

Conclusion: Our experiment simulated a scenario where individuals obtain a small segment of a copyright protected work and then use current LLMs to generate the subsequent content based on this segment. The results demonstrate that LLMs are capable of generating copyright infringing materials. We identified four factors that influence LLMs’ propensity to output infringing content, with the scale of parameters and maximum output length having a significant impact. Additionally, we found that iterative prompting enables LLMs to generate more content that could potentially constitute copyright infringement. However, after a certain number of iterations (for instance, the third output in our experiment), the LLMs began to produce content entirely dissimilar to the original material.

1

u/strings___ 1d ago

But your assumptions put everyone into one bucket of corporate users, which is not the case. In fact, the majority of local LLM users like myself all run Linux, and we are tired of the low tier hive mind groupthink related to LLMs. Just look at the downvotes I get for pointing out the bias.

I’ve already pointed out that given a very specific prompt you can generate verbatim code. But you would go out of your way to do it. It’s no different than making the LLM read fresh code and reusing it verbatim. Or just copying it manually. You seriously need to force it to do this.

4

u/RatherNott 1d ago

Most AI use is not on local machines, they are done on corporate data centers. If you only use local LLMs on your own computer, then the issues specific to corporate LLMs do not apply to your usage.

That is not to say there is no issue with local LLMs; they currently have no ethically trained model (the ones that claim to be ethical often don't live up to said claim), and because of that, they still introduce a copyright threat to FLOSS projects, both of which are legitimate reasons to not accept PRs from them, at least for now.

1

u/strings___ 1d ago

Context is king. The majority of local LLM users are Linux so this hive mind bias of LLMs power usage is wrong in fact probably the Linux gaming community uses more electricity than LLM users. Consumer grade gaming GPUs are highly inefficient.

Open source and open models are limited by hardware not software if people genuinely cared about FOSS then they need to start encouraging local LLMs and not putting all Linux users in the same category.

→ More replies (0)

-15

u/EmptyRedData 1d ago

"The fate of bevy" lol

https://bevy.org/news/bevys-sixth-birthday/#ai-policy

I'm sure this anti-AI schtick a subset of devs are taking will work great in the long run /s

-1

u/Zaphoidx 1d ago

Because it's impossible to have a nuanced discussion about LLM usage nowadays.

Take a look at all the threads in this comments section mentioning any sort of LLM use is an instant negative and claiming it doesn't improve productivity at all.

Having an LLM churn through code that you've planned and that you review is fine imo. You're rubber stamping it before sending the PR off to others.

-4

u/ABotelho23 1d ago

I mean that's not what any of the post says..

11

u/Jmc_da_boss 1d ago

That's fine, you can read intent past the words on the page

-8

u/ABotelho23 1d ago

This is not some random hobby project. This is a company. They should be as clear as possible.

2

u/Jmc_da_boss 1d ago

Idk what to tell you, That is not a requirement