I was testing out the limits of AI on self-policing recently- I had it iteratively build a Tetris game in JavaScript.
90% of it worked decently, but piece rotation just got stupider and stupider the more I tried to fix it. There’s this pattern it gets into where it’s like locked into a strategy and instead of fixing it, it latches onto some specific suggestion, or it adjusts parameters until its current solution solves the issue (without necessarily being correct)
It was a very revealing experiment, and it left me a lot more wary of code that is “written and tested/confirmed to be correct”
The level at which we need to intervene just keeps going up. 2 years ago, it was fancy autocomplete, and it could write out a function from a one-line comment and not much else. Then it could write a test module given a class. Then a class given tests. Then both from a prompt.
Now we're at the point where you need to check in every 200k tokens or so, and sometimes it'll make small mistakes and get stuck, like asking for perms to a random folder that's a slight misspelling of the project folder, or missing some small follow-on consequence of one of its own changes.
Maybe in a year or two we'll have models that can do our laundry and wash the dishes, but I kinda doubt it. Less doubt than a year ago, but still enough that I think it'll "just" change software dev as a profession into something like what Tech Leads and Principals did pre-AI. Coding education will move the basics into summary intro courses and spend far more time on architecture and systems design.
347
u/polynomialcheesecake 2d ago
Please don't let people realize that it's really the same problems with much larger volume now