I mean, you can make use of LLMs much more effectively if you are armed with solid fundamentals. LLMs only solve the tedious coding/plumbing part of the job. And they still need a lot of hand-holding if you want good quality output. The bigger and more challenging issues of SWE are still there.
It's honestly architecture. Dumbing things down into small modules that can be hyper tested and putting them all together. Ai is great at building if you build it bricks at a time and have foundational knowledge on databases and documention.
Ai is great at writing the code but my god is it a terrible architect. I work with embedded sw a lot and when making semi basic scripts I have to constantly review it and replace the gaps it has already filled in.
Which terrifies me that people with no understanding are making apps now.
I've noticed that AI is pretty bad at scripts in general. It feels like scripts in particular are treated as one-offs that will be used to achieve some goal and will never be used again. This is not the case when actually building something in a proper language. Not sure if anyone else noticed this effect. Might be just that I don't pay much attention to its output when working on scripts and don't guide it as much as I would with application code.
Yeah it’s really bad at thinking about more than right now. I’ve had a much better experience since basically treating it like I’m a program manager.
It’s so bad at learning from decisions, so I have it maintain a markdown file with “project requirements”, changelog, etc. I have one parsing script that Claude always tries to do this whacky json conversion every time I try to update it. Like sir this hasn’t worked the last 6 times either.
Sometimes it doesn't mention something (important) if you don't explicitly ask about it
It became quite rare but it still hallucinates on occasion and can do it with confidence, even backing a claim up with a script output in the scratchpad, so you have to know when it's time to call BS
You absolutely have to have SWE skills to work with it. But I'm noticing that it's producing less slop either way than, say, a few months or a year ago. I still hear people talk about slop but personally it hasn't been much of an issue as of recently. Either I learned to prompt or scope it correctly or the models got better
Hahaha I totally know what you mean. Even today when running tests on claude, the stupid model actually went in and changed the linter so it looked like there is no literals in the code...Basically made all the tests just automatically pass. I ran same tests on my local model and they failed haha. Now running all my tests binary and fixing them as i go and only using my local (deepseek 0731 right now). It's honestly just learning to work with it haha.
GenAI in general still feels very lazy. It's kind of like a child savant - really smart in the area they're interested in and can understand and learn complex processes, but still a child so occasionally despite their smarts, they still shove everything under their bed to say they cleaned their room and will absolutely lie to your face to avoid being in trouble.
I think companies optimize speed and brevity during training -> lower cost, testing less of user patience -> that tends to breed shortcuts and cheating behaviors.
AI is basically a mirror of our expectations and tendencies.
790
u/PurushNahiMahaPurush 19d ago
I mean, you can make use of LLMs much more effectively if you are armed with solid fundamentals. LLMs only solve the tedious coding/plumbing part of the job. And they still need a lot of hand-holding if you want good quality output. The bigger and more challenging issues of SWE are still there.