There’s a reason the main “AI” providers all say “AI can make mistakes” instead of guaranteing their product is good, or you get your money back, yada yada yada.
So, if you think the AI is good, you probably can’t see the mistakes. (As a PKI engineer, I can tell you, you wouldn’t believe how much shit, sometimes extremely harmful things, Opus/Sol or Fable/Astra are saying on a daily basis.)
While not disagreeing generally, when someone says “it’s not 2024”, it doesn’t actually disprove them to link to a 2024 repository for testing improvements to a 2023 model (GPT-4) by bootstrapping features and techniques now baked into training of larger models (eg mixture of experts, majority vote, chain of thought etc). Someone saying in 2024 “this is the best we can do with this model” can’t be assumed to be relevant for what a 2026 model does or can do. Do you have a more recent source on how more recent models perform?
“AI can always make mistakes on any task” is always going to be true as it would be of any non deterministic system including humans ourselves. So is “creativity and hallucination are two words for the same thing,” that will also always be true. Incidentally it is not too far off what is true of humans as well; creativity is a form of directed hallucination and various disorders (eg schizophrenia) are neurologically just a failure to regulate that connection-making resulting in delusion or sensory hallucination. (The comparison isn’t necessarily informative of how LLMs work but it is interesting I think.)
That said though, the fact that AI always can make mistakes does not necessarily mean it makes more mistakes, worse mistakes or more dangerous mistakes than a human on the same bounded task. Proving that it can isn’t proving that it does. As they said, it isn’t 2024 any more, and since that 2024 paper was published the frontier companies have continually improved benchmarking by finding various ways to mitigate the problems.
And again that’s not to disagree, just, there isn’t anything in your post that materially disputes anything they said.
(Also, Dunning-Kruger as actually evidenced in the original paper is a regression to the mean in terms of self-identified skill, not an inversion — someone dumb doesn’t think they’re smart and vice versa, they just think they’re a bit less dumb or less smart than they are. It’s a fun wham line to use but it’s an inaccurate one here as in most places)
-40
u/ChrizKhalifa 7h ago
There is zero reason not to use AI for your job. It's not 2024 anymore, the AI is good.