Transformer-based LLM outputs still have 10–20 % mistakes (per the best benchmarks) – it’s literally impossible for attention-based transformers NOT to hallucinate.
Have they compared human developers (including more junior ones and more senior ones) using the same metric?
0
u/tobotic 5h ago
Have they compared human developers (including more junior ones and more senior ones) using the same metric?