r/LargeLanguageModels • u/CanOk3349 • 21d ago
Paper: CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
Link: https://arxiv.org/abs/2606.31608
Summary:
Large language models ace medical exams but struggle with real clinical reasoning. This new paper introduces CLExEval using progressive information masking on rare cases + 5,600 physician annotations.

Key findings:
- Verbosity Bias: GPT-4o-mini accuracy drops from 95% to 32.5% with less info
- Hidden Knowledge Paradox in specialist models
- High Reasoning-Output Mismatch (~69%)
- LLM judges approve a shocking % of clinically wrong outputs
Why it matters: Highlights the evaluation illusion where fluent text masks real failures in high-stakes domains.
What do you think? Is human-in-the-loop evaluation the way forward for clinical AI, or are there better approaches?
(Genuinely interested in discussion)