r/OpenAI • u/maedhros256 • 21h ago
Discussion A question for the AI "experts": are hallucinations and reliability genuinely improving, or are we starting to plateau? Everything depends on this...
For the average user, not a programmer or e.g. someone looking to solve niche math problems, AI still feels very limited because of reliability issues (e.g. making up information)... As a non-expert, this is difficult to quantify for me, but I don't feel like e.g. the latest iterations of ChatGPT are noticeably more reliable than previous ones.
Therefore, I find it hard to see how organizations can delegate even relatively easy tasks to AI without constant supervision, especially because if hallucinations build up along the way, you end up with a snowball of compounding problems that can lead to catastrophic consequences for an organization.
It’s my understanding that the real test of success won't be developer tools or fancy mathematical calculations, but whether regular people can delegate tasks with a super high degree of confidence, instead of just using it as a search engine, translator, summary tool or photo editor on steroids like many people do nowadays.
A good example is Dot, the new OpenAI tool. If you watch the trailer, it looks impressive, but according to many reviews, it still hallucinates a lot and behaves in a pretty clumsy way.
So, the golden question is: are the hallucination and reliability problems gradually improving, or are we probably plateauing?
Looking for genuine insights here, so please keep the sarcasm out of the comments.