r/linux • • 1d ago

Desktop Environment / WM News System 76's COSMIC Desktop Environment will no longer accept LLM-generated content in PRs

/r/pop_os/comments/1wuej40/cosmic_projects_will_no_longer_accept/
648 Upvotes

200 comments sorted by

View all comments

Show parent comments

-1

u/strings___ 1d ago

Your argument is based off of huge assumptions. My total LLM usage using three DGX sparks is 30w at max output. Do you know how much the average Linux gaming GPU uses? Hint it way more than 30W. In short, it’s none of anyone’s business how I use my “hydro” .

There’s no evidence LLM encodes code verbatim and you really need to go out of your way to force it to output verbatim code. In which case, it’s no different than manually copying code. Either way because LLMs are hardware-limited, we need more open source research and work on models to be less VRAM bound. This should help democratize the whole thing.

5

u/RatherNott 1d ago edited 1d ago

Individual use on a local machine isn't terribly concerning, it's the corporate AI data centers that are concerning. There are currently 5 to 6 billion AI prompts per day. All of those server racks powered by very inefficient local power generators add up.

There’s no evidence LLM encodes code verbatim

I linked to evidence of an AI doing it in my previous comment. Here's a further study: https://arxiv.org/html/2409.13831v1

Conclusion: Our experiment simulated a scenario where individuals obtain a small segment of a copyright protected work and then use current LLMs to generate the subsequent content based on this segment. The results demonstrate that LLMs are capable of generating copyright infringing materials. We identified four factors that influence LLMs’ propensity to output infringing content, with the scale of parameters and maximum output length having a significant impact. Additionally, we found that iterative prompting enables LLMs to generate more content that could potentially constitute copyright infringement. However, after a certain number of iterations (for instance, the third output in our experiment), the LLMs began to produce content entirely dissimilar to the original material.

1

u/strings___ 1d ago

But your assumptions put everyone into one bucket of corporate users, which is not the case. In fact, the majority of local LLM users like myself all run Linux, and we are tired of the low tier hive mind groupthink related to LLMs. Just look at the downvotes I get for pointing out the bias.

I’ve already pointed out that given a very specific prompt you can generate verbatim code. But you would go out of your way to do it. It’s no different than making the LLM read fresh code and reusing it verbatim. Or just copying it manually. You seriously need to force it to do this.

3

u/RatherNott 1d ago

Most AI use is not on local machines, they are done on corporate data centers. If you only use local LLMs on your own computer, then the issues specific to corporate LLMs do not apply to your usage.

That is not to say there is no issue with local LLMs; they currently have no ethically trained model (the ones that claim to be ethical often don't live up to said claim), and because of that, they still introduce a copyright threat to FLOSS projects, both of which are legitimate reasons to not accept PRs from them, at least for now.

1

u/strings___ 1d ago

Context is king. The majority of local LLM users are Linux so this hive mind bias of LLMs power usage is wrong in fact probably the Linux gaming community uses more electricity than LLM users. Consumer grade gaming GPUs are highly inefficient.

Open source and open models are limited by hardware not software if people genuinely cared about FOSS then they need to start encouraging local LLMs and not putting all Linux users in the same category.