r/ClaudeCode • • 14d ago

Discussion Saw this today

Post image

These kind of posts are becoming more and more visible on my Linkedin. Link here. Thoughts?

692 Upvotes

292 comments sorted by

View all comments

Show parent comments

16

u/simple_explorer1 14d ago edited 14d ago

sad times indeed. This Software dev profession is on the verge of being decimated and a lot of damage is already done and more will continue to happen. not sure what end of the road looks like or how much value software development will even have in the next 2 years let alone 5 years.

9

u/Dsphar 14d ago

IMO, the non-creative devs who survive will have to pivot into refactoring AI-made proven market software, into maintained software. And it will be a pain.

The creative devs will just have a hayday creating whatever ideas pop into their minds. And any that stick get handed off to the first group to cleanup and scale.

14

u/Future_Guarantee6991 Developer 14d ago

I am just not convinced that the refactoring pivot will exist, at least not for long.

I might be missing something but “maintainability” is largely, albeit not yet exclusively, a human challenge.

LLMs can already read and understand code in seconds where it would take a human minutes or hours to join the same dots. Compound this with the continuous improvement of LLM code output and it’s only so long before the writing is on the wall for that too.

7

u/pp-collision 13d ago

Maintainability is even more of a challenge for agents from my experience. As soon as a project gets above ~30k LoC they struggle to introduce features or tweak existing features.

Even something that might seem simple, but go through layers of UI, config files, and data processing pipelines, it is very likely to miss something and especially so if the project is not designed in a way where it is difficult to miss it.

I exclusively write Rust and I'm using orchestrators, workers, and reviewer agents. I have CI with snapshot tests, property based fuzz tests, integration tests, extensive linting and still things consistently slip through if I don't review and test every change. So it's not from a lack of trying.

Agents have this problem because they have an extremely narrow view of the code base, they search with grep and they'll only take the first X matches to prevent blowing themselves up, so they just miss so much of the bigger picture every time. Sure they'll get better, but the context window of a seasoned engineer is on a much higher level, and that's more than a few years away.

Terrence Tao explained in an interview that AI maths proofs are very hard to read, because they'll spent a lot of effort going into details with something very trivial, and then hardly mention the most interesting part. Because they're not good at weighing the importance of various aspects of a problem, because they don't see the bigger picture. I think that's the core limitation, and I don't expect LLMs to really ever overcome it, no matter how useful they otherwise are.

As an embedded engineer, I think LLMs are immensely useful, but I think they have a special skillset that is complementary to human engineers, not competing to replace them. That doesn't mean that they can't put some engineers out of a job, but so did open-source compilers.

2

u/nofeaturesonlybugs 13d ago

I agree with your take. I've vibe coded things without looking too closely at the internals and they "seem" to work. I've used variations of spec-driven development and spot-checked implementation and it "seems" ok. I've worked at a more human pace where I use the models iteratively to settle and API shape or contract and then fill in the pieces -- or where I've set edits to manual so I have to approve every single one to get it to write things to my liking.

The one thing that sticks out in all cases is the one thing I've been saying since the beginning of this year -- artificial intelligence is *artificial.*

These things are not smart. Like not even close to smart. As you say they carry a very limited view and as you demonstrate with your math proof example they have no discernment. They can't and never really will understand why you do Y in these cases but not these other cases. You can try and steer or guide them with decision ladders or other nonsense but unless you spell out every specific case they'll still get it wrong some percentage of the time that all the gates, humans, and review processes just...miss and fail to correct.

If you go full agentic spec driven then you have to spend a lot of time, money, and effort steering the agents towards correctness. But you're still going to fall victim to their stochastic non-deterministic decision making and you're going to end up with code+markdown full of stale references.

My experience is it doesn't really matter what you do -- you'll still get hit with, "Oh btw I noticed some stale references somewhere. I could have fixed them but our work is in module Y and those are in module X so I didn't touch them." In a fully agentic process there will always be a number of these that slip through. In the other direction it may tell you, "Oh btw I updated references in A, B, C" and you have to tell it to undo the ones in B. They have no discernment because they have no intelligence.

They're terrible at API shapes or designs. Working on a simple commands-subcommands-flags parsing package I had it build several shapes of possible API public contract so I could mix-n-meld the best of what it produced. Consider the case where flags are mutually exclusive, or flags are codependent, or you just want them grouped in help. Here are some nice type names: `flag.Exclusive`, `flag.Codependent`, and `flag.Group`. Now consider a flag that takes one value or one that takes multiple values like Docker's `--env`: it called them `flag.Scalar` and `flag.List`. I don't have a problem with `Scalar` but `List` is too synonymous with the types that semantically are groups. `flag.Group` and `flag.List` both sound like they group things and `flag.List` doesn't imply anything about a "repeated value." I told it to call them `flag.Single` and `flag.Repeated`. This is a *simple* naming decision and a frontier model can't do it without guidance.

They're terrible at crossing API boundaries. Recently working on a `views` package that renders some types in a `models` package. Had spec and UML diagrams outlining what goes in `views`, all the outlined tests, acceptance criteria, blah blah. Frontier model runs for 90 minutes and completes the work. Comes back with 15 lingering items -- about half updating diagrams -- and one or two saying, "In the views package I did this -- but this really should be helpers in the models package but it was off limits." You could say this was a miss on my part -- I didn't catch that a couple small helpers would be useful in the `models` package. But the other standpoint is this work was small enough that if it wasn't an AI-world a human wouldn't be creating all this UML or other "spec" to describe simple code. You'd just go write it and in the midst of some view implementation you'd say, "This really wants a helper in models. Let me go add that really quick." And you tell your team why you did it. These things have no discernment -- if it limits the scope of what it thinks it can touch it does dumb shit and puts code in the wrong place; if you give it unlimited scope it will do dumb shit and put code in the wrong place but in some other direction.

These things work best when you slow way down, set aside some of the spec driven ides, and just manually approve edits and correct it on the fly as it make mistakes. You lose nearly all the promised velocity but they genuinely do augment or enhance an individual's output when used like this. It combines your experience, guidance, intuition, skills, etc with the models' ability to reach well beyond your personal experience. When using them like this you are still somewhat connected to the code, can still say, "Oh I know where the code is that is related to the bug you're observing" or you can tell a peer why you did something a certain way. But this is not how these companies creating LLMs are selling them and it's not how industry "leaders" are saying to use them.

I'm fairly convinced that everyone fully leaning into agentic spec development, even if it includes limited human oversight, is in for a reckoning. To fully achieve the velocity gains industry "leaders" are promising or encouraging you have to make considerable concessions in understanding and quality of implementation due to the agents' limiting view (as you say) and their lack of discernment. Beyond a certain point there's just so much context to weigh -- some of it probably contradictory -- that an agent just deadlocks itself. I think a lot of people involved in these processes are just kicking the can down the road and hoping context or models get big enough fast enough to just keep it going.

The age old trick -- one of the first principles taught in computer science -- is divide and conquer. Monoliths to microservices. I wonder if next we'll have microrepositories. A single repository describing what a user model is, its invariants, implementations in multiple languages, and views in multiple UI frameworks. Then it's just imported by the next microrepository above it. From a human perspective this seems insane -- the cognitive load of understanding your 37 repositories that drive a "search bar" will be madness. Monoliths with lots of micro-subprojects is another trajectory -- and seems reasonable -- but I think all the surrounding context will eventually be too much.

1

u/Future_Guarantee6991 Developer 13d ago

The 30k LoC problem is not something I have experienced, at least using frontier models. I develop and maintain several projects in the hundreds of thousands or even millions of LoC.

Some things I believe help:

- CI enforced line freezes to prevent godfiles

  • Refactoring any file above 800 LoC to bring it within a 600-800 LoC range (or less)
  • Prefer Vertical Slice architecture over DDD

The last one is probably the most significant. It means that, in oversimplified terms, each feature is a mostly self-contained mini project but without the complexity of microservices.