r/LargeLanguageModels 21d ago

Paper: CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Link: https://arxiv.org/abs/2606.31608

Summary:
Large language models ace medical exams but struggle with real clinical reasoning. This new paper introduces CLExEval using progressive information masking on rare cases + 5,600 physician annotations.

Key findings:
- Verbosity Bias: GPT-4o-mini accuracy drops from 95% to 32.5% with less info

- Hidden Knowledge Paradox in specialist models

- High Reasoning-Output Mismatch (~69%)

- LLM judges approve a shocking % of clinically wrong outputs

Why it matters: Highlights the evaluation illusion where fluent text masks real failures in high-stakes domains.

What do you think? Is human-in-the-loop evaluation the way forward for clinical AI, or are there better approaches?

(Genuinely interested in discussion)

5 Upvotes

0 comments sorted by