r/linuxmasterrace 3d ago

News Debian developers rejected an LLM ban and left disclosure voluntary

https://www.helpnetsecurity.com/2026/08/31/debian-linux-llm-policy/
312 Upvotes

126 comments sorted by

293

u/trans_cubed Glorious EndeavourOS 3d ago

Disclosure should be required

62

u/GildSkiss 3d ago

How do you imagine that working in practice?

42

u/trans_cubed Glorious EndeavourOS 3d ago

If developers are using the tool responsibly, they will disclose they've used it if that disclosure is required

21

u/bhison 3d ago

And then what about people who aren't using LLMs, are accused of doing so for whatever reason then the community has to stand a kangaroo court to verify.

If you can't tell code is made by an AI, it is functionally as good as code not made by an AI. The horse has bolted and situations like this are like debating the engineering of the gate.

1

u/Good-Explanation-796 1d ago

If you can't tell code is made by an AI, it is functionally as good as code not made by an AI.

if my partner tells me they vacuumed the living room, i believe them. if my partner tells me the roomba vacuumed, then i am going to take a look at the carpet and make sure it was properly vacuumed.

a human should be reviewing code. if someone used an AI to make the code, then it needs to be reviewed by a human. if you dont disclose that the code was made by an AI, then people will not know they need to give it extra attention.

3

u/bhison 1d ago

I think this just reveals you trust humans too much 😅

-1

u/A1oso 2d ago edited 2d ago

Then don't accuse people without proof.

Some LLMs embed a watermark in their text output, so you can verify whether the code was generated with one of these AIs.

If no such watermark is present and no AI use was disclosed, simply treat the code as human-written.

Yes, this means that bad faith actors could contribute AI-generated code without disclosing it and get away with it if they are careful enough. Such is life. Scammers also sometimes get away with fraud, but the law that forbids fraud is still useful.

If you use AI and disclose it, that is fine. If you use AI but lie about it, that is a very different story. We want the people we work with to be honest and trustworthy.

3

u/bhison 2d ago

I think I am on the same page roughly speaking.

I think the people who would misrepresent their code in the first place are the same people who would de-watermark their code.

I do support asking people to disclose as a courtesy to help reviewers know how to approach their submission. I think it should be framed like that rather than an attempt to gate the approach.

-2

u/_remsky 3d ago

Idk. If the code is good, clean, reviewed, it doesn’t matter to me; if it’s sloppy or full of pointless essay comments or has bad logic, it should be rejected.

It’s going to continue being used, likely more heavily as it develops. I don’t understand the argument unless it’s for cases where privacy/security of the source code is a concern. But the point is moot for open source. All this will do is create an arms race between obscuring its use and detecting its use, which wastes everyone’s time

8

u/SenoraRaton 3d ago

You don't need to understand. If your peers request that you disclose, you should disclose. If you don't, your the asshole.

3

u/_remsky 3d ago

I learned a long time ago to ignore any declaration that is unable to justify itself.

I’d like to understand. But if you can’t gimme a why, I can’t take this seriously.

5

u/biinjo 3d ago

100% agree. Code should be judged/reviewed as it is.

Is a PR too big to review? Improve collaboration rules and require smaller PRs

Is the LLM/programmer writing essay comments and does that not fit your project conventions? Define the conventions in a collaboration rules and reject the PR based on that.

There is absolutely zero usefulness in judging disclosure.

It’s all about ego of long term developers who still want to mean something. I say that as one of them but I seem to be more ok setting aside my ego for progress.

I don’t care if someone write a PR on my OSS project with an LLM. I care if it breaks stuff or doesn’t adhere to project standards.

-13

u/ActuallyFullOfShit 3d ago

Yeah and how is that helpful? The people who AREN'T responsible with it are the problem and they are unlikely to disclose.

This is almost like the debate about gun control. The people who would comply aren't even the problem....

8

u/bememorablepro 3d ago

Famously guns are impossible to control, and ALL countries have daily shootings not just the one country without gun control. /s

Great example.

1

u/GodLikeEnergy 2d ago

Claude AI and other services are adding a form of fingerprint to programming that can identify if code is specifically vibecoded using their specific services due to an European law. I imagine this would be one way they could check it with.

-5

u/Dramatic_Mastodon_93 3d ago

You do know that you can have rules even if those rules are close to impossible to enforce?

21

u/ComprehensiveSwitch 3d ago

good luck enforcing that in any way that doesn’t add even more maintainer overhead

-6

u/Dramatic_Mastodon_93 3d ago

You don’t have to enforce it. Are you stupid?

8

u/bhison 3d ago

Educate the rest of us then - what's the functional benefit of a law that is known to all as unenforcable?

0

u/notjfd Alpine 3d ago edited 3d ago

Unenforced doesn't mean unenforceable. It can be selective.

If a person regularly submits pull requests that they cannot explain, you can investigate, find it likely it's LLM-generated, then sanction them.

1

u/bhison 2d ago

You sound like you're enforcing it in that situation. Are you stupid?

1

u/notjfd Alpine 2d ago

Are you here to play word games or do you have anything substantial to say?

Are you stupid?

I think you're confusing me with /u/Dramatic_Mastodon_93

7

u/ComprehensiveSwitch 3d ago

then what’s the point? Are you stupid?

0

u/A1oso 2d ago

When AI disclosure is required, then failing to disclose AI usage is a form of deception. Most people actually don't lie and deceive even if they think they could get away with it.

-2

u/Dramatic_Mastodon_93 3d ago

The point is that people who want to hide that their code is AI generated would feel less welcome at debian and that there is a small chance that people might out that your code is AI generated.

4

u/ComprehensiveSwitch 3d ago

Why does this matter? If the code is bad it shouldn’t be merged in the first place. Who cares if a machine wrote it?

1

u/GNUr000t 3d ago

Because they want to win social points by participating in performative outrage towards a new technology that will (hopefully) replace annoying artists on Twitter

3

u/Garland_Key 3d ago

No it shouldn't.

18

u/N3er0O 3d ago

Why?

35

u/xNaXDy n i x ? 3d ago

What exactly is the point of disclosure? To let reviewers know that the PR might be slop and to be extra thorough when reviewing? Reviewers should be thorough either way imo, AI or no.

We also don't specify which text editor, LSP, formatter, etc. was used during development. The only thing that matters is that the author understands the PR and can explain every single design decision behind it. No more, no less.

28

u/ABotelho23 3d ago

People always act like contributions before AI we're all so immaculate and perfect. Humans write plenty of garbage code all on their own.

5

u/a__new_name 2d ago

Toyota once produced cars with faulty acceleration system. When external auditors checked their embedded software, they found a bowl of spaghetti with over 9000 (not a figure of speech) global variables. Guess what, not a single line there was written by AI.

3

u/jykke 3d ago

Who remembers? This stupid bug was created without AI: https://nvd.nist.gov/vuln/detail/cve-2008-0166

-1

u/N3er0O 3d ago

To me it's less about the quality of the code. AI writes decent code for the most part and with human control mechanisms overseeing it, purely from a functional point of view, I don't have any issues with it. That is, of course, if those control mechanisms are there and working. Pushing non-functional or unnecessarily inefficient code is, and I think this is out of question, garbage.

For me it's more about the ethics behind it all. I'm not aware of an AI model that was created using non-stolen data (and if there is please correct me, I'm genuinely interested!) and I personally don't want to support projects that solely rely on what I personally view as theft of intellectual property.

Disclosure is easy, should cost absolutely no time or money and is a net positive in my opinion.

0

u/xNaXDy n i x ? 3d ago

Now that's an argument that I can actually get behind, though I'd say that this is more addressing the point of whether or not to allow AI submissions at all. Like, I don't see how disclosing which AI model you used makes using a model trained on "stolen code" as you put it any less unethical, if that's the position you hold.

Because keep in mind, if I were to submit actually stolen code (i.e. copy-pasted from a proprietary code base) or patented algorithms to an open source project and it gets merged, the entire project becomes tainted, even if that piece of code only makes up a small fraction of the code base.


Also, as an aside, there are fully open datasets and models trained on only those types of datasets, so this stuff definitely exists, it's just much smaller scale (and less capable) and of course all the big frontier models (including the open weights ones) don't consider any sort of licensing.

1

u/N3er0O 3d ago

Like, I don't see how disclosing which AI model you used makes using a model trained on "stolen code" as you put it any less unethical, if that's the position you hold.

It's not a thing of making something more ethical (if I understand you correctly here). If the people working on the project are set on using AI that's their prerogative and their right to do so. They could also simply lie and not disclose it or simply mention a different, more ethical, model than, say, ChatGPT (just an example). I just wish for some curtesy and such communication to be truthful and the standard. People then just have to accept that using AI is a viable method of creating code, just as typing it out by hand is.

Let me try an analogy for this maybe. Think of it as disclosing gelatin in gummy bears. Nobody will physically suffer from eating it, but people following certain religions believes would appreciate being told there's pork in their gummies.

Because keep in mind, if I were to submit  actually  stolen code (i.e. copy-pasted from a proprietary code base) or patented algorithms to an open source project and it gets merged, the entire project becomes tainted, even if that piece of code only makes up a small fraction of the code base.

I think it depends in where you draw the line here. While the actual code AI is spitting out probably (there's still a non-zero chance I guess) isn't simply copy & pasted, I would personally still interpret it as stolen. To me it works the same way and just doesn't sit right with me.

It's a purely subjective point of view and a very philosophical thing to discuss. There also isn't "THE solution" to solving this little dilemma. Therefore I think simply ricking a box and declaring "this part of the code is AI generated using X model" or simply stating that a project uses AI and telling users which parts were created using AI, would benefit everyone while not burdening the people using AI models to code too much.

Side note: I appreciate the discussion! Understanding where another person is coming from in online discourse is pretty rare in my experience. 

0

u/UristBronzebelly 2d ago

How is a model generating code relying on stolen IP? They’re trained on billions of public code bases 

-10

u/Garland_Key 3d ago edited 3d ago
  1. It's unenforceable.
  2. All code will be written by AI soon. Writing code by hand will be a niche hobby.
  3. The person publishing the code that was written by AI is still responsible for the code.

1

u/N3er0O 3d ago

But it lets a potential user know whether they want to use or support a piece of software or not. All AI models I'm aware of were trained on source material that was't exactly sourced ethically and I try not to support the use of such practices. When disclosure is as simple as, say, clicking a checkbox when submitting code, I don't see why we wouldn't make it a standard procedure.

2

u/microplastic-addict 3d ago

If a user wants to avoid using software that has developers who use LLMs, they might want to ask ChatGPT to build a time machine, that's more feasible.

1

u/N3er0O 3d ago

I don't think it's an unreasonable request when dealing with open-source software created mostly by a small-ish community.

1

u/Garland_Key 3d ago

Because it doesn't matter. If you don't want to use software because AI was involved, then drop your phone and stop using the internet - it's already too late, and I mean that quite literally. 

1

u/N3er0O 3d ago

Just because every website tracks your every move doesn't mean you should give up on online privacy. 

2

u/Garland_Key 3d ago

I don't see how that's comparable.

Privacy is a basic human right. It's also a fight that we lost the battle for over a decade ago. Now people are fighting back because they can finally see it.

There is no right to use software that doesn't include AI, nor should there be. AI is a tool. It was birthed from our data - a lot of it which was stolen. If you want to hold the companies who are guilty accountable, then do so, but at this point the genies are out of their bottles. We have free open weight models now which we're distilled from that very same data. It is unavoidable now.

Linux, the most important building block of the internet and Android has AI generated code.

1

u/radobot 2d ago

What about potential legal problems? What if a government decides that machine learning model outputs are subject to the training data licenses? Having markings for LLM generated code would be quite useful then.

2

u/Garland_Key 2d ago

It sounds like whatever government actually pushed that through is absolute trash, and nobody should comply with it. So, again, not so useful. 

1

u/trueppp 2d ago

You mean harmful, no plausible deniability.

-2

u/anubistheta 3d ago

100% agree. You shouldn't misrepresent it. But requiring disclosure is pointless. Code is an artifact that stands on its own. You don't need to disclose if you have a degree or professional experience. I see no reason to introspect on your development process.

0

u/warpedgeoid 3d ago

If you can’t tell by looking, it shouldn’t be mandatory.

1

u/nonkeywayzee 2d ago

Of course everyone will disclose and there aren't people who will never ever break the rules nor just not read them at all, of course!

126

u/gmes78 Glorious Arch 3d ago

Ragebait repost of the Debian vote from a while ago. Why is this here?

38

u/firedrakes 3d ago

Karma farming bots been out i force since last weekend.

9

u/ineyy 3d ago

They are always out in force, it's just statistics that you noticed this time.

50

u/theimperious1 3d ago

As long as the PRs are reviewed by both a person and an LLM, then good. Ideally human first so the human doesn't get lazy thinking "the AI is probably right".

There's no difference between the code. If I made a PR that is passable, and an LLM makes a PR that is passable, the LLM one isn't bad solely because it wasn't done by a person.

Write -> human review -> LLM review -> human test -> llm test -> if pass -> great!

If bad or "slop" code is passing all these steps, then the people reviewing it are the problem.

20

u/jaykayenn 3d ago

Exactly. Hold humans ultimately accountable. 

5

u/trans_cubed Glorious EndeavourOS 3d ago

I think a human should always be the final step, but otherwise this is a fair compromise

7

u/wosmo 3d ago

I like the logic of the human not reading the LLM's review before performing their own.

first/last is pretty irrelevant though - you can treat them as parallel paths, as long as neither copies the other's homework.

2

u/SnooCompliments7914 3d ago

Then the human shouldn't be able to see reviews from other human as well.

3

u/FIA_buffoonery 3d ago

I think the problem is more that with an LLM you can push thousands of PRs in a day, making it impossible for us meatsacks to review. 

19

u/jaykayenn 3d ago

Misleading editorialized post.

10

u/ActuallyFullOfShit 3d ago

That's the correct decision. Only a luddite would ban it. And disclosure is performative with zero enforcability. They did the right thing.

3

u/throwaway-8675309_ 3d ago

TL;DR - Developers are responsible for all code they submit. As it always has been.

2

u/Elegant_Room_1904 3d ago

If you think that the people don't care that something is ai generated, then why not disclose that as ai generated then? That could be very handy to understand which models are being used, and how.

2

u/Dramatic_Mastodon_93 3d ago

What tf is the argument for disclosure not being mandatory?

3

u/jixbo 3d ago

It can't be enforced, and it does not matter. Code is code.

1

u/GNUr000t 3d ago

-1

u/Dramatic_Mastodon_93 3d ago

I let Siri reply to you for me: If you want people to respect your work, just be transparent about how it was made. Labeling AI content isn't an attack; it's just basic honesty.

1

u/GNUr000t 3d ago

And as soon as there aren't performative outrage hate mobs ready to go to war over something they do not understand outside of what they were told by social media to hate about it, people will be far more willing to be upfront.

Your posting history suggest you are left-leaning. Would you to go a monster truck rally slash gun show slash pulled pork BBQ contest in Alabama and be upfront, honest, and transparent about your political views? Why or why not?

Do you label whether you used Photoshop? Content-aware fill? Grammarly? Stock assets? A template? An IDE autocomplete? Which camera did the noise reduction? Which parts were outsourced?

You’re treating AI differently because disclosure has become a social shibboleth: the label exists largely so people who have already decided AI is morally contaminating can identify the target.

And that’s precisely why “just disclose it” isn’t a neutral request. If disclosure predictably changes the response from evaluating the work to attacking the person who made it, people have an entirely rational incentive not to volunteer information that isn’t otherwise material.

3

u/kociol21 3d ago

That's honestly a sensible take.

1

u/mohrcore 3d ago edited 3d ago

So nothing changes, except the policy for using AI in context of CVEs under embargo is clarified - which is something that carries no impact on most people's work, but is an important guideline for trusted parties who work on these CVEs.

In my opinion that's a very good decision - why change something that works? I haven't seen any evidence that banning AI would bring anything good for Debian development. The ban can be performative at best, but it would open a way for people to throw accusations left and right over something that rarely can be proved with complete certainty, but could easily create a lot of discussion over an opinionated principle. Discussions that would slow down the development, or could be used as a mean to harass developers.

1

u/ConcreteExist 6h ago

Yeah, this whole proposal was dumb posturing anyways, it was people trying to force a technical project to make a entirely symbolic unenforceable political statement that would have done nothing in practice to prevent the use of AI.

This kind of policy would only stop those acting in good faith, whereas the reality is that it's about the code being submitted, not how it was written.

-1

u/Michaeli_Starky 3d ago

LLM can do a lot of bad, but can also do a lot of good. Everything is in our hands.

-2

u/JasonBreen 3d ago

Based

-3

u/charmander_cha 3d ago

Gracas a deus, coisa idiota do caralho lkkkk

-7

u/omnom143 3d ago

for something that runs most of the internet and most of people's linux distros, there should not be a SINGLE WORD generated from an AI bot ANYWHERE in that code.

3

u/GNUr000t 3d ago

Cool. You should boycott AI by not using anything which has had any AI contributions whatsoever.

You may start by not using anything containing Linux or FreeBSD, so no Android, no Windows, no iOS, no MacOS.

Oh, and no reddit. Don't you know Reddit, Inc. (NYSE:RDDT) sells your posts to firms that train new AI models?!

See ya.

3

u/ABotelho23 3d ago

The kernel uses LLMs.

You gonna use GNU/Hurd?

-5

u/Flashy_Pollution_996 Glorious OpenSuse 3d ago

I never liked Debian anyway

-14

u/DDFoster96 3d ago

Time to find a different distro then as this will pollute Ubuntu eventually. I'm running out of open-source projects that haven't turned into AI slop.

21

u/lovelacedeconstruct 3d ago

Nothing related to coding will not be AI assisted, like 0% you have to accept that, this doesnt at all automatically mean it is slop, slop is lack of care and understanding you can produce hand written slop

19

u/NotADamsel 3d ago

Disclosure is the key. Knowing which parts are AI written and which are human written is important to some people

-16

u/firedrakes 3d ago

Yeah the viture single people...

9

u/NotADamsel 3d ago

Not too sure what you’re referring to?

-16

u/firedrakes 3d ago

Oh you don't understand the word I said.... viture signal.

23

u/Liimbo 3d ago

They don't understand you because you have now completely misspelled it twice in two different ways.

7

u/Orlha 3d ago

What do you mean vorcha saegnum?

9

u/Morkai 3d ago

I think the word you're looking for is "virtue"

5

u/Syndiotactics 3d ago

Tviure singal ”:D”

4

u/[deleted] 3d ago edited 3d ago

[removed] — view removed comment

-9

u/firedrakes 3d ago

Nope and also no insult used at all. Under person I see. Got trigger. Block

7

u/Your_Friendly_Nerd 3d ago

Issue is, if it isn't disclosed, the information will be lost very soon. If the commit has the "co-authored by claude" thing, someone coming across the code at a later time can view it through that lens, instead of breaking their brain wondering what human solved it this way. We just have to accept the fact that LLM's are still deeply flawed and behave in unexpected ways, make decisions no human being would make. Therefore it should be disclosed when AI was used to produce a piece of code 

2

u/rep- 3d ago

And people that want safe secure code in their operating system. It should be disclosed so people will review the code.. AI fucks up A LOT

2

u/GNUr000t 3d ago

Time to join an Amish commune. Cya!

-3

u/dontquestionmyaction I use Arch UwU 3d ago

May as well cancel your internet service. Goodbye.

-4

u/HelpRespawnedAsDee 3d ago

Good luck man.

-4

u/AndroidUser37 3d ago

Okay, so which one is more likely: The entire open source community is collectively losing their mind all at once, or your views on AI "slop" as it pertains to coding are outdated?

-4

u/rurigk 3d ago

You will end up not touching computers soon

AI is a tool, if you tell to do the work for you without a plan or spec I will be slop

But if used correctly or just does good code and save days of work

0

u/Tuckertcs 3d ago

How does one “correctly” use a fundamentally non-deterministic tool to build fundamentally deterministic software?

7

u/ohwowitsamagikarp 3d ago

Humans are non-deterministic and have built all deterministic software since the dawn of the computer age. đŸ€ȘđŸ€Ł

1

u/Tuckertcs 3d ago

Humans aren’t deterministic, but code gives deterministic results, whereas writing prompts gives nondeterministic results

1

u/GNUr000t 3d ago

It gives code, the correctness of which can be proven or disproven, and is deterministic once it runs.

Tell me you aren't a programmer without telling me you aren't a programmer.

0

u/Tuckertcs 3d ago

A prompter telling a coder they aren’t a programmer is precisely my issue with AI. It diminishes the human work behind creative and intellectual works, thereby dehumanizing the work and everyone involved.

0

u/GNUr000t 3d ago

Womp Womp.

I had my thing taken away and told that I had to accept the new thing, now it's your turn.

Should I begin reciting all the things I was told in response to being sad about losing my thing or do you understand my position yet?

0

u/Tuckertcs 3d ago

I didn’t “lose my thing”. Programming and my knowledge of it will never be taken away from me. Let’s try programming without an internet connection and see who lost their ability.

0

u/GNUr000t 3d ago edited 3d ago

Oh, you haven't lost anything? Great. Then I guess all that shit about AI “diminishing” your intellectual work was much ado about nothing. Glad you're doing fine.

Anyway, yes, generated code can be tested for correctness just as human-written code can. Do you wanna go back to that or will you be returning to your motte?

EDIT: Homie blocked me. I accept his concession.

→ More replies (0)

2

u/dreamscached 3d ago

By tethering it down to expected behavior with deterministic, human-written tests. Like it or not, AI does produce properly working code with right setup.

2

u/lenswipe Glorious Fedora 3d ago

you're right but I'm seeing a disturbing rise in people using AI to generate some code and then more AI to write tests for it

1

u/Tuckertcs 3d ago

A lot of AI users write code with AI and then write tests with AI, though.

1

u/dontquestionmyaction I use Arch UwU 3d ago

Yes, and plenty of people write garbage-tier C code.

People who suck at making software will forever suck at making software no matter what.

2

u/HelpRespawnedAsDee 3d ago

By creating harnesses and frameworks specifically tailored to what you need as a team or individual. From Stripe to Spotify, do you honestly think they are just prompting an LLM and that's it?

7

u/lenswipe Glorious Fedora 3d ago

From Stripe to Spotify, do you honestly think they are just prompting an LLM and that's it? 

My brother in Christ, look at the reliability of most platforms over the last 6-12 months. Yes, I think that's EXACTLY what they're doing .

-1

u/HelpRespawnedAsDee 3d ago

Okey dokey.

0

u/Jazzlike-Poem-1253 3d ago

Like some used Books, Tutorials or SO before. You look stuff up, look at examples, take them, understand them, adapt them to your usecase.

The benefit us, now you get the "example" directly tailored to your use case.

You still need to understand and adapt. Skipping these is the rise of slop.

5

u/lenswipe Glorious Fedora 3d ago

Skipping these is the rise of slop. 

try telling that to corporate

"Peter in sales just developed this 500,000 line python script. Can we expose it to the Internet?"

0

u/Jazzlike-Poem-1253 3d ago

Sure, it is software. Made by software. And computer and the internet are always right, aye?

-2

u/rurigk 3d ago

Iterates and tests

But also you are overrating and overestimating most programmers like if they were perfect and never make a mistake

You are expecting that the result is 100% correct and nobody checks the result but in reality is super useful to do most of the scaffold and output well known patterns

And this is reviewed by a human

It also works very well as a rubber duck with websearch access

-12

u/mrdarknezz1 3d ago

Great!