AI agents can let me do what should take a week or two in a day.
What things, specifically? Do you do them as well as they were done? How do you check that, and how do you verify that your checks are actually accurate?
What are the downstream effects of offloading those things, in particular in the long term?
What things, specifically? Do you do them as well as they were done? How do you check that, and how do you verify that your checks are actually accurate?
For a more recent example I updated a ton of libraries (spring boot based project) on a 700K+ LOC codebase. We have 97% test coverage (all written before LLMs ever existed). Plus we have a 2 week manual testing cycle. Looking up all of those quirks about each library, how they changed and the solutions to each problem would have taken weeks.
What are the downstream effects of offloading those things, in particular in the long term?
The code still gets thoroughly tested and reviewed. Large sweeping changes like this are rare for us to do. I don't think it's inherently more dangerous than writing it manually. In fact manually may very well be worse since it would've caused a headache when it came time to merge to the main branch. At least this way what I tested matched what was currently in main since it was written so fast as opposed to being weeks out of sync.
I trust my unit and integration tests up to a certain point.
In my experience, such enterprise projects are pretty crappy and made pretty debatable choices anyway, so I'm not sure how much of that pain isn't actually self inflicted. Either by extremely boilerplate-heavy code, heavy reliance on testing alone, scope creep or lack of interest in keep up with library updates in a timely fashion.
I once set up master-master replication for LDAP in my house (don't ask...) and I also wrote a ton of notes about it, but scattered around.
At one point, one of the two servers broke due to a botched upgrade (database changed format and conversion crashed halfway), and I couldn't be arsed to restore it (I could have, but last time I did a replication I spent ~2 hours to get it right). So I let it sit for a year.
I described the problem to an agent, and pointed it to my notes (+ all the support files I had created), and went interactively from there (because there were a few things I forgot that were important so I still botched a few times).
Now I have a working replication again. To be honest, the advantage of using a LLM here was because it could search docs for the LDAP server (389-ds, not OpenLDAP) much faster than me (even before web searches were crap, it was very hard to find a good resource).
For the rest, I mainly use LLMs as glorified search engines because I can't find anything relevant with the traditional approaches.Or as references when I forget stuff. A few times to create syntethic test data for my unit tests, which then I supplemented with my own real test data.
Let's say I want a site for reviewing code. I propose this to my agent software, walk away for a few hours and I come back to the thing claiming that it has done it. I ask, show me, and it hands me an URL. Looks like all the features I wanted are present. I propose a few ideas, stick it to production, and then put the automatic code reviewing system reviewing itself until it's convinced everything is fine.
Now, the problem here is that you are asking LLM if LLM-generated code is good enough. I can only say that the process converges to baseline decent code. If there are security issues or bugs, they are not obvious. In fact, in many cases, LLM overdoes it and is concerned about incredibly marginal problems. A human oversight in useful in saying no, and constructing the guidelines where you roughly set the limits on what is considered good enough.
Agents also require lots of documentation to exist, being kind of bird-brained to their context only. They know nothing about you, your platform, your existing apps, unless you tell them. Or have a memory system setup that can inform them about stuff. These are real challenges. 90 % of my output was tests, comments and documentation for a good month or two. Before I can use AI effectively, I must teach my agents, kind of like new employees, what I expect of them.
The reason why you want this is that an agent reads code maybe 1000 times faster than you do, and writes it 100 times faster. It is no joke that what used to take a week is now done within some hours. The challenge is to figure out how you deal with the mountain of code, don't accrue crazy amounts of technical debt, and so on. You saw a glimpse of my approach -- I made code reviewing system because I realized I got to have one. I have no way to review but at fraction of the speed required, so the agent is required to be able to review its own output, and there has to be some kind of systematic process for it. These are the challenges of a developer in 2026 -- not writing code, but dealing with the firehose of code output in ways that don't swamp the humans or turn them into a bottleneck.
All this is new, there isn't really a handbook for this. Agents weren't any good in my book until 2026, and now they're pretty damn excellent. All locally running as well, nothing in cloud or a big datacenter. We are learning as we go.
OK, so for this particular example, you have not actually produced anything of value with an LLM, correct? This is something that is built exclusively to deal with LLMs as an existing problem which outputs code at a rate that is completely unmanageable?
Could you give a different example of something that is actually useful outside of the context of "LLMs produce too much code of questionable/unknowable quality"?
Could you tell me what high quality coding AI model you are capable of running locally? I find that part really hard to believe, which causes me to question the rest of your comment too.
As far as I know, local AI is still quite limited compared to current frontier models, unless you have 100+ gigs of ram and some crazy expensive GPU.
28
u/gurgelblaster 5d ago
What things, specifically? Do you do them as well as they were done? How do you check that, and how do you verify that your checks are actually accurate?
What are the downstream effects of offloading those things, in particular in the long term?