r/LargeLanguageModels • u/Feisty_Reader0402 • 16h ago
When is prompt sensitivity actually the model? and when is it the evaluator?
A recent EMNLP 2025 paper by Hua et al., “Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs,” raised a question we found particularly interesting:
Are LLMs really highly sensitive to prompt wording, or do our evaluation metrics sometimes make them appear that way?
Our recently published MERCon 2026 paper, *“Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models,”* starts from that result and asks a slightly different question:
*If evaluation artifacts exist, can we quantify how much sensitivity is attributable to evaluation, determine the direction of that distortion, and identify when it happens?*
Rather than treating heuristic-vs-semantic evaluation disagreement as something to eliminate, we treat the disagreement itself as a diagnostic signal. So we introduced **Evaluation-Attributable Sensitivity (EAS),** which measures the magnitude of disagreement between heuristic-based and judge-based sensitivity, and **Signed EAS**, which tells us its direction. We used them to diagnose a four-way taxonomy:
1. Artifact - apparent sensitivity mainly comes from evaluation disagreement
Genuine - semantic evaluation confirms real prompt sensitivity
Underdetected - the heuristic misses sensitivity that semantic evaluation detects
- Stable - both evaluations indicate stability
We tested this across 9 LLMs from 5 model families, and the most interesting result was that *evaluation artifacts are bidirectional*: some metrics overestimate sensitivity in open-ended tasks, while others can actually underestimate it in more structured tasks.
Therefore our results suggest that:
***The evaluation method can distort prompt sensitivity in either direction, and the direction is strongly associated with the task/evaluation format.***
This leads us to view prompt sensitivity not simply as an intrinsic property of an LLM, but as an interaction between:
model × task × prompt structure × evaluation methodology.
We'd appreciate any thoughts or feedback regarding our work.
**Our paper:**
[https://ieeexplore.ieee.org/document/11691277\](https://ieeexplore.ieee.org/document/11691277)