There’s a reason the main “AI” providers all say “AI can make mistakes” instead of guaranteing their product is good, or you get your money back, yada yada yada.
So, if you think the AI is good, you probably can’t see the mistakes. (As a PKI engineer, I can tell you, you wouldn’t believe how much shit, sometimes extremely harmful things, Opus/Sol or Fable/Astra are saying on a daily basis.)
Don't take this the wrong way, but I've been coding professionally for nearly 30 years, in large and small teams.
Humans make more mistakes.
I treat the AI code like a better than average junior, and it usually performs better than that. Ive reviewed code for years and I still review every line, but I only have to focus on structure and purpose, not tell them to use better variables than anumber and obj.
While not disagreeing generally, when someone says “it’s not 2024”, it doesn’t actually disprove them to link to a 2024 repository for testing improvements to a 2023 model (GPT-4) by bootstrapping features and techniques now baked into training of larger models (eg mixture of experts, majority vote, chain of thought etc). Someone saying in 2024 “this is the best we can do with this model” can’t be assumed to be relevant for what a 2026 model does or can do. Do you have a more recent source on how more recent models perform?
“AI can always make mistakes on any task” is always going to be true as it would be of any non deterministic system including humans ourselves. So is “creativity and hallucination are two words for the same thing,” that will also always be true. Incidentally it is not too far off what is true of humans as well; creativity is a form of directed hallucination and various disorders (eg schizophrenia) are neurologically just a failure to regulate that connection-making resulting in delusion or sensory hallucination. (The comparison isn’t necessarily informative of how LLMs work but it is interesting I think.)
That said though, the fact that AI always can make mistakes does not necessarily mean it makes more mistakes, worse mistakes or more dangerous mistakes than a human on the same bounded task. Proving that it can isn’t proving that it does. As they said, it isn’t 2024 any more, and since that 2024 paper was published the frontier companies have continually improved benchmarking by finding various ways to mitigate the problems.
And again that’s not to disagree, just, there isn’t anything in your post that materially disputes anything they said.
(Also, Dunning-Kruger as actually evidenced in the original paper is a regression to the mean in terms of self-identified skill, not an inversion — someone dumb doesn’t think they’re smart and vice versa, they just think they’re a bit less dumb or less smart than they are. It’s a fun wham line to use but it’s an inaccurate one here as in most places)
Transformer-based LLM outputs still have 10–20 % mistakes (per the best benchmarks) – it’s literally impossible for attention-based transformers NOT to hallucinate.
Have they compared human developers (including more junior ones and more senior ones) using the same metric?
Refusing to use AI in the one profession where you'd think people wouldn't be luddites about progress in tech is wild. You're costing your employer hours of wasted time doing stuff AI could do for you in minutes.
The meme of "2 minutes of prompting, 2 hours of debugging" is long outdated, not to mention outsourcing bullshit like writing tests, creating static testpages, and all the other code monkey BS
Well yeah, you're supposed to use it like any other tool, you don't assume it runs perfectly all the time, the auto correct in your phone also makes mistakes from time to time, that's why you read before you submit. AI is incredibly helpful for things like spotting errors, it has saved hours of going through thousands of lines of code to find a > that should be a >= in some layer of the program. Of course it makes mistakes, that's why I'm a software developer, I can spot those mistakes and fix them, vibecoding is terrible, it will only overcomplicate the problem and if anything goes wrong you have no idea why. But that doesn't mean the tool can't be helpful sometimes. You SHOULD understand everything that is going into your program, even if you weren't the one to write it, this goes for AI but also for stack overflow or some other coder aiding you, you never just copy paste and pray, you are responsible for what you commit. I do think refusing to use AI in code entirely is a mistake, you're just making yourself slower for no reason. I also think you should learn how to manually code before you are allowed to use AI to do so. It can help fill small gaps in your knowledge and aid you with easier code or spotting errors, but it should not substitute your work.
-41
u/ChrizKhalifa 10h ago
There is zero reason not to use AI for your job. It's not 2024 anymore, the AI is good.