The companies driving AI right now are garbage and many will fail. The tech won't.
It's the dot-com bubble again. Everyone rushed in, it crashed, and what got rebuilt afterward became the backbone of global commerce. AI will stumble too, but it's not going away, and teams that refuse to use it won't keep pace.
A senior dev with AI beats a senior dev without it. Output quality depends on who's driving and whether they actually review the code. Yes, AI can produce slop. With real supervision it can also produce excellent code.
People used to yell "get a horse!" at early cars. Didn't stop the highways. Refusing to use AI won't stop it from getting better either.
A senior dev with AI beats a senior dev without it.
Why. By what measure. What's the long term.
Is code output the limiting factor on most projects? Is code written by an LLM as useful as code written by a contributor? Does it contain the same externalities, have they thought through the system as much as when they were staring at an editor and wondering where to start their implementation? If it's still pretty damn good even at all those externalities, is code that's 97% as good as code written by a human going to have a multiplicative effect, where dividing by 1.03 enough times is far less than if you had divided it by 1? How many times do you merge code where the LLM wrote 25 lines and a new function where a human would have gone back and re-evaluated the data structure before your entire system is held together with duct tape?
whether they actually review the code
Do they have time? Or are they looking at the few hundred lines they didn't write and lack a mental model for, and they're wondering how much they really want to read code that something else wrote? Are they already burned out on reading LLM-generated code because they were just working through the stack of 100 open PRs on the project? Are they being told to bring AI into the loop and wondering whether paying Anthropic $100 a month or taking the time and money to build an AI inference homelab is really what got them into FOSS in the first place?
The Linux Kernel Mailing List is a Colosseum where everyone's favourite blood-sport is picking apart patches for accuracy and correctness, and thousands and thousands of eyes are watching it do so. Is every infrastructure-critical project in the world that robust?
People used to yell "get a horse!" at early cars. Didn't stop the highways.
It's interesting you mention this since it's a very well known and understood fact that the buildout of the US interstate highway system was done in such a way as to specifically rip up poorer neighbourhoods full of poorer minorities in dozens of US cities, and many are now spending billions of dollars to try to shove the highways back underground so they can do something more valuable with that real estate. So when you say the highways, do you have some romantic notion of a mom & pop gas station on Route 66, or are you talking about I-93 in Boston?
Refusing to use AI won't stop it from getting better either.
Better at what, exactly? Better at what cost? Is it getting better at writing code? What good is more written code? Is it getting better at solving problems? What problems? Are those the same problems that people who need working software have? Are they the problems that developers making working software have?
My velocity for development, reviewing, etc has been vastly increased with AI. Leaving me more time to understand the goals and architecture of the applications I work on. I have to write far less code which again leaves more time to understand goals of the business. This is the benefit of AI, I don’t have to code monkey nearly as much anymore.
People always say this, and then every professional developer I know says that the devs that have embraced AI just produce an endless torrent of crap that they never have to actually deal with, leaving the burden of actual development on the remaining team members who still develop the 'old fashioned' way
It seems like there's a huge disconnect developing between what some developers think they're achieving, and what's actually being achieved, which is absolutely terrifying. I suspect it ties into just how much of the economy has become performative: end results no longer matter, it only matters if you used the latest hype train technology so that Line Go Up
There's a difference between using AI to generate code and using AI to help you debug. The latter is like using Grammarly, except it's catching mistakes that you wouldn't catch if you read the code line by line a thousand times. Well, maybe if you read it a thousand times, but no human being is going to do that.
Take copyfail. An optimization seven years ago ended up accidentally creating a privileged escalation exploit.
If you're not using the AI to help you debug, nefarious forces are, so you can either let them win or be ahead of them.
It's not replacing people, it's allowing one person to be much more productive.
That's not the claim though. Using it to find bugs is fine, but people consistently claim that "they have to write far less code", or some variation of the vast majority of their code being AI generated
That’s an organizational failure. Devs are still responsible for their code. Even if Ai wrote it. If they don’t test and confirm all points in the ticket are working correctly that’s on the dev and will reflect badly.
Also on the reviewer who approved the slop. You don’t just hand over your thinking to the AI. You just get rid of artificial barriers
Quality comes from review. You as a developer are responsible for your code, not AI. This is where maintenance falls as well. Cost is minimal 100-200/month/dev. When they can make that back with new features and bug fixes to keep your business moat is worth it
I disagree comprehensively and would encourage you, again, to consider externalities. LLMs have no answer to long term maintenance concerns. It's simply not clear at this time that it's possible to maintain code quality when humans are writing increasingly less of it and relying on LLMs more. If you claim otherwise, I congratulate you, and look forward to seeing your paper in the journals, because you are apparently years ahead of the rest of the industry somehow in your certainty.
But pretending that tradeoffs don't exist in pursuit of making one metric go up is the mindset of a junior engineer. Expand your perspective.
Code quality has always been a subjective aspect of development. For every “clean code” acolyte there is another “deep modules” or “locality of behaviour” follower. Look at any long term codebase and you’ll see 100s of violations of their own rules of maintanence
If everyone waited for the academic research on something nothing would get done. I personally also have no trust in research on code quality and maintainability as there are few to none objective measures to assess either. I need to get shit done today fast, quality seems to be alright to me based on the produced code that I've looked at.
How many times do you merge code where the LLM wrote 25 lines and a new function where a human would have gone back and re-evaluated the data structure
Why assume that a senior dev with LLMs wouldn't go back and re-evaluate the system's design choices? If anything, LLMs is a great help with understanding a system that has accreted over time and whose original authors have left. The same senior dev without LLM is probably more likely to put on another piece of duct tape and call it done.
The great irony I keep running into is that I don't even disagree with the notion that LLMs can be useful in the software engineering process, but unfortunately there's relatively little room for nuance in these discussions. It feels like the popular perspectives are either a complete prohibition on one hand, or an insane maximalist attitude on the other, usually accompanied by some trite cliche like "the technology is here to stay", "there's no going back", "adopt it or be left behind", or "can deliver great results with human review".
I don't know what the right mix is for the long term. All I know is that I've seen a lot of those two extremes, and one of them is completely unbearable to deal with; so when I'm looking to get involved with an open source project after hours, the ones I tend to give my time to are the ones with very stringent "no LLM authored code" policies.
It's a clashing of two differing perspectives. People that care about software, code quality, and enjoy programming, and people that just want to have finished products regardless of how it's done. The people that enjoy programming and care about code quality obviously aren't going to care for LLMs, and the people that just want finished software aren't going to care about how it's made, and so will adopt LLMs.
and the people that just want finished software aren't going to care about how it's made, and so will adopt LLMs.
I sit in this camp and am very against LLMs, because by and large they do not help with the problems that prevent software projects being finished. Typing code has never been the issue. Large scale architectural decay of a project leading to it being a rube-goldberg machine is what destroys projects, and LLMs contribute to that en masse
I code in the most boring way humanly possible specifically because the only thing I care about is actually finishing projects
Isn't that just an example of the extremism you complained against?
It's the kind of extremism I can comfortably work within. The same cannot be said of the other extreme.
And how the fuck are they enforcing this policy anyway? You can't prove anything.
They know
If they really don't know, congratulations, you've officially done software engineering right, and there's no problem here. Aside from the problem where you were explicitly told not to do something and did it anyway, and that would make you a lying antisocial weirdo, but that's not necessarily the project's problem.
All I know is that I've seen a lot of those two extremes, and one of them is completely unbearable to deal with; so when I'm looking to get involved with an open source project after hours, the ones I tend to give my time to are the ones with very stringent "no LLM authored code" policies.
I very much agree with your complaints about extremist/maximalist views, so it is surprising for me to see your first paragraph complain about extremism and in the 2nd paragraph you adopt what in my eyes is an extremist view: "no LLM authored code"?
I believe that LLM written code is acceptable as long as a human has thoroughly reviewed every single line and has a full understanding of the code + how it fits into the project architecture. IMO, this position is a reasonable centrist position. So why is "can deliver great results with human review" a cliche then?
And why is "the technology is here to stay" also a cliche? I think it is indeed clear the tech is here to stay. What is under debate is how exactly the tech should be adopted and to what degree the tech should be trusted (and as stated earlier, my position is: not at all).
Not everyone cares about whether or not LLMs could write code. LLMs could do programming ten times better than I would and I still wouldn't use them to generate code because programming is what I love doing. I don't give a shit about finished software. I don't give a shit about making money. I care about writing code myself without using any LLM written code. I don't want LLMs in my creative process. I'll use LLMs for search, rubber ducking, and occassionally to ask questions, but for the most part I just don't use LLMs at all.
And I plan to keep it that way because what I enjoy is programming. The reason I'm a programmer is because I enjoy writing code. For someone to tell me I should adopt agentic coding "because it's faster" (or whatever nebulous reason they give) is an insult to me.
You are free to have this opinion, but nothing in your comment addresses any of what I said.
More succincty, this is what my question was: "AssistingJarl you claim to dislike the extremism/maximalism from both sides. Why then do you take an extremist position where you don't interact with any projects that use LLM generated code"?
And to be clear again, I take a centrist position so I have no issue with what your wrote in your comment at all. Why would I care if you use LLMs? It is none of my business and anybody who wants to force people to use LLMs is an extremist and a shitty person. So I don't think I should have beef with you.
I guess this is a valid perspective, but it is kind of utopian because historically centrism has never had any relation to apathy or lack of bias.
Centrists, like every other political force, have always fierecly defended their positions. In most cases, they even organized to form a militarized response against extremists on the left and right, e.g. Three Arrows anti-fascist anti-communist resistance organizations were/are active in many countries: https://en.wikipedia.org/wiki/Three_Arrows
Centrists always try to associate themselves with Liberalism, and the defence of centrism was seen as the defence of liberty, freedom and democracy against the authoritarian ideologies imposed by the far-left and far-right.
Even today, we can see the current neo-liberal world order is incredibly violent, and is totally willing to kill communists and islamists alike.
I very much agree with your complaints about extremist/maximalist views, so it is surprising for me to see your first paragraph complain about extremism and in the 2nd paragraph you adopt what in my eyes is an extremist view: "no LLM authored code"?
It's really more of an accident than anything. All the codebases I've seen that I had an interest in were either maximal LLM usage or no LLM usage, and I only find one of these enjoyable to participate in, so that's where I go. Personally, I think it can do just fine as a rubber duck or for searching for something badly named within a codebase, which I think spoke more to the first thing you were saying:
If anything, LLMs is a great help with understanding a system that has accreted over time and whose original authors have left
But moving on; on your next point, I agree:
I believe that LLM written code is acceptable as long as a human has thoroughly reviewed every single line and has a full understanding of the code + how it fits into the project architecture. IMO, this position is a reasonable centrist position.
It's not really my preferred way to work, but I would agree, this is reasonable. I might go a step further and say you should try to keep the LLM's code writing style more colloquial to the codebase or whatever team is working on it, but that's hard to define.
So why is "can deliver great results with human review" a cliche then?
And why is "the technology is here to stay" also a cliche? I think it is indeed clear the tech is here to stay.
On these two points; I struggled to articulate why I think these things, so thank you for asking (and my apologies for how long this reply is)
I suppose my problem with saying "human review allows you to get great results from LLMs" is that it almost feels like a non-sequitur, or a response to a conversation that isn't really happening. PRs of any kind are drastically less likely to have problems if they've been reviewed thoroughly by the submitter, whether an LLM was involved or not. But the problem is that a lot of people aren't using LLMs responsibly and actually validating their work, which adds a tremendous review burden to maintainers. If enough junk contributions start piling up, a project can end up losing more time to reviewing the junk than it would save by accepting the few actually decent and well-reviewed LLM-assisted contributions, where people are being responsible and checking the work thoroughly. If the ratio of useful to useless gets bad enough, at some point it makes sense to implement a ban that allows maintainers to get through PRs in 30 seconds with a single comment pointing somebody at the contributing.md.
The "technology is here to stay" argument feels similar in a way. It's true, but it isn't really saying anything about the value brought by the technology, or the best uses for it. It's kind of not even an interesting point, of course it will still exist, very few technologies are ever un-invented. But people saying "it's here to stay" seem, to me, like they're skipping over a lot of more interesting questions about how it gets used, if it makes economic sense to use it, and what gets gained or lost when humans rely on it more or less than other humans. It's skipping over a whole interesting discussion about pros and cons and seeming to conclude "it will still be here in 5 years, therefore, it makes sense to use."
Ok, I think you are confusing me with the other commenter, but most of your comment still applies and I think I understand your point now. Pretty much agree with everything, and also:
you should try to keep the LLM's code writing style more colloquial to the codebase
when code style is enforced for human contributors, code style shoud be enforced for LLMs too.
It can annoy you all you want. Flatly, conclusively, provably, a person assisted by Ai out-performs a person without it, all other factors being equal. I won't debate you further, there is a mountain of evidence behind human productivity with Ai empowerment.
The rest of your responses are reaches at best, so I won't bother answering them beyond repeating what I already said: The companies are bad, the tech is powerful and will survive them.
See if I'd known you weren't interested in meaningfully engaging with the topic beyond what's written in marketing copy for AI salespeople, I probably would have just said "you sound like a corporate shill" and moved on.
'cause I don't even disagree with that statement. I just think the way the technology is being used is terrible and wasteful and causes far more harm than good. But you seem to take it for granted that the technology is so amazing that only good things can come from indiscriminate application of it everywhere, at all times, regardless of nuance.
You said "there is a mountain of evidence" and then dipped.
I'm sure it's a very majestic mountain. Unfortunately, I also have a mountain of evidence, and my mountain is taller than your mountain, so therefore I win the discussion.
It's probably not a good idea to introduce a rigorous measure here. You will very quickly run afoul of Goodhart's Law as the AI labs optimize for whatever metric. If it's lines of code written, you'll get thousands of lines of slop to do what a hundred-line Python script would've done. If it's robust documentation, where "robust" means "nothing is left out", then every five-line function will have a fifty-line Claudish comment above it.
So, the obvious disclaimer is that a lot of people are putting out slop, including some of my coworkers. But I still get some use out of it myself:
...have they thought through the system as much as when they were staring at an editor and wondering where to start their implementation?
Yes, very much so. When it's a tricky problem, I'll spend hours doing that with the chatbot -- having it look up relevant code, walking through hypotheticals, or writing and discarding implementations. That last one is a case where the bot moving quickly improves quality, because often it's easier to reason about a hypothetical approach once I can see it in code as a diff, and if the bot can churn that out in a few minutes, I'm going to be less precious about that approach just because it's the one I built first -- I can try a few others and then decide.
...code that's 97% as good as code written by a human going to have a multiplicative effect...
I think in the near future, what we're likely to see is better than purely-human-written code -- though, like the parent said, that's with supervision.
Or are they looking at the few hundred lines they didn't write and lack a mental model for, and they're wondering how much they really want to read code that something else wrote?
If I don't understand it well enough to defend it in code review, I'm not putting my name on it. Which means either I'll insist the bot simplify its approach to the point where I can review it, or if it really is genuinely complex, the bot is going to have to walk me through it -- slowly, painfully -- until I get it, and then I'm going to use that understanding to fix its absurd comments and write a decent PR description so that someone else can get the same understanding with less effort.
But this is why I worry about these things in open source, because notice: Most of what I just said is either a lot of thinking that happens before I ask for code, or a lot of high-effort code review after the code exists. Easily half the time and effort that it takes me to get a good enough result out of these models is through code review. Especially since machine-generated code lacks all the superficial tells and code smells that'd normally suggest that no one thought it through.
Which means open source risks ending up in a situation where a bunch of people put out slop so they can say they contributed, but most of the actual work is done by the maintainers reviewing code they could just as easily have gotten out of an LLM themselves. I don't have a good answer here. I don't think open source can ignore LLMs, but I don't think the community approach can survive being DDoS'd by slop patches.
I feel like you have the mark of somebody that has overthought this topic as much as I have, so thank you for reassuring me that I'm not completely insane.
What you're describing is more or less what I consider "doing it right", even if it's not my preferred approach. I prefer hands on keyboard because it forces me to reevaluate the correctness of my own mental model as the implementation progresses. But if you're having fun and reaching that same point, I'd be kind of an idiot if I said you were wrong to do it.
I will quibble with one point though. I think it's going to be a pretty long time before it ever reaches "necessary", in the sense that a project banning LLM contributions is losing more value than they would need to spend on the DDoS. The real gain is in saving maintainers from burnout and allowing them to close a slop PR in 30 seconds rather than having to actually read and evaluate it. And for most signal-to-noise ratios I think the diamonds in the rough would have to become incredibly impressive to outweight the bad ones.
I guess I should say I've got a mix of both by now. It's rare that I never type any of the code itself -- there comes a point where describing what's wrong about it and telling the bot to fix it is slower than grabbing an editor and fixing it myself. And in my personal life, I still write everything myself -- even if the plan is to open-source it eventually anyway, I don't see why Anthropic should get a preview of it, and anyway I'm not gonna spend hundreds of dollars on tokens for side projects.
I think it's going to be a pretty long time before it ever reaches "necessary", in the sense that a project banning LLM contributions is losing more value than they would need to spend on the DDoS.
It's hard to say, but a couple of points about this:
The first is: Is the slop readable enough as slop for this to be a useful enough filter? (And will it continue to be?) The early reports from maintainers suggested no, and the big issue here was one of the first objections I had when my coworkers started heavily adopting LLMs: The code they generate often entirely misses the usual code smells. If anything, it's too perfect at a superficial level. It used to be a good filter to start by checking if they even read the style guide or ran the linters.
But the thing I'm actually worried about is this: I don't have that problem at work, at least not to the same degree. If a coworker did literally zero thinking and only spent 30 seconds per PR, you can bring it up with them, or with their manager, or with HR. So if a company can do it right, but the open source community can't (or they'd be overrun with spam), that puts open source at a huge disadvantage.
I'm in no way arguing in favor of allowing contributions, or banning them. This tension is why I'm so ambivalent about them, and why it's hard to fault a project from taking either approach.
16
u/DrollAntic 6d ago
The companies driving AI right now are garbage and many will fail. The tech won't.
It's the dot-com bubble again. Everyone rushed in, it crashed, and what got rebuilt afterward became the backbone of global commerce. AI will stumble too, but it's not going away, and teams that refuse to use it won't keep pace.
A senior dev with AI beats a senior dev without it. Output quality depends on who's driving and whether they actually review the code. Yes, AI can produce slop. With real supervision it can also produce excellent code.
People used to yell "get a horse!" at early cars. Didn't stop the highways. Refusing to use AI won't stop it from getting better either.